An Algorithmic Lucidity

a blog

Tag: politics

Conflict Theory of Bounded Distrust

(originally published at Less Wrong)

Scott Alexander once wrote about the difference between "mistake theorists" who treat politics as an engineering discipline (a symmetrical collaboration in which everyone ultimately just wants the best ideas to win) and "conflict theorists" who treat politics as war (an asymmetrical conflict between sides with fundamentally different interests). Essentially, "[m]istake theorists naturally think conflict theorists are making a mistake"; "[c]onflict theorists naturally think mistake theorists are the enemy in their conflict."

More recently, Alexander considered the phenomenon of "bounded distrust": science and media authorities aren't completely honest, but are only willing to bend the truth so far, and can be trusted on the things they wouldn't lie about. Fox News wants to fuel xenophobia, but they wouldn't make up a terrorist attack out of whole cloth; liberal academics want to combat xenophobia, but they wouldn't outright fabricate crime statistics.

Alexander explains that savvy people who can figure out what kinds of dishonesty an authority will engage in, end up mostly trusting the authority, whereas clueless people become more distrustful. Sufficiently savvy people end up inhabiting a mental universe where the authority is trustworthy, as when Dan Quayle denied that characterizing tax increases as "revenue enhancements" constituted fooling the public—because "no one was fooled".

Alexander concludes with a characteristically mistake-theoretic plea for mutual understanding:

The savvy people need to realize that the clueless people aren't always paranoid, just less experienced than they are at dealing with a hostile environment that lies to them all the time.

And the clueless people need to realize that the savvy people aren't always gullible, just more optimistic about their ability to extract signal from same.

But "a hostile environment that lies to them all the time" is exactly the kind of situation where we would expect a conflict theory to be correct and mistake theories to be wrong!—or at least very incomplete. To speak as if the savvy merely have more skills to extract signal from a "naturally" occurring source of lies, obscures the critical question of what all the lying is for.

In a paper on "the logic of indirect speech", Pinker, Nowak, and Lee give the example of a pulled-over motorist telling a police officer, "Gee, officer, is there some way we could take care of the ticket here?"

This is, of course, a bribery attempt. The reason the driver doesn't just say that ("Can I bribe you into not giving me a ticket?"), is because the driver doesn't know whether this is a corrupt police officer that accepts bribes, or an honest officer who will charge the driver with attempted bribery. The indirect language lets the driver communicate to the corrupt cop (in the possible world where this cop is corrupt), without being arrested by the honest cop who doesn't think he can make an attempted-bribery charge stick in court on the evidence of such vague language (in the possible world where this cop is honest).

We need a conflict theory to understand this type of situation. Someone who assumed that all police officers had the same utility function would be fundamentally out of touch with reality: it's not that the corrupt cops are just "savvier", better able to "extract signal" from the driver's speech. The honest cops can probably do that, too. Rather, corrupt and honest cops are trying to do different things, and the driver's speech is optimized to help the corrupt cops in a way that honest cops can't interfere with (because the honest cops' objective requires working with a court system that is less savvy).

This kind of analysis carries over to Alexander's discussion of government lies—maybe even isomorphically. When a government denies tax increases but announces "revenue enhancements", and supporters of the regime effortlessly know what they mean, while dissidents consider it a lie, it's not that regime supporters are just savvier. The dissidents can probably figure it out, too. Rather, regime supporters and dissidents are trying to do different things. Dissidents want to create common knowledge of the regime's shortcomings: in order to organize a revolt, it's not enough for everyone to hate the government; everyone has to know that everyone else hates the government in order to confidently act in unison, rather than fear being crushed as an individual. The regime's proclamations are optimized to communicate to its supporters in a way that doesn't give moral support to the dissident cause (because the dissidents' objective requires common knowledge, not just savvy individual knowledge, and common knowledge requires unobfuscated language).

This kind of analysis is about behavior, information, and the incentives that shape them. Conscious subjectivity or any awareness of the game dynamics are irrelevant. In the minds of regime supporters, "no one was fooled", because if you were fooled, then you aren't anyone: failing to be complicit with the reigning Power's law would be as insane as trying to defy the law of gravity.

On the other side, if blindness to Power has the same input–output behavior as conscious service to Power, then opponents of the reigning Power have no reason to care about the distinction. In the same way, when a predator firefly sends the mating signal of its prey species, we consider it deception, even if the predator is acting on instinct and can't consciously "intend" to deceive.

Thus, supporters of the regime naturally think dissidents are making a mistake; dissidents naturally think regime supporters are the enemy in their conflict.

Scoring 2020 U.S. Presidential Election Predictions

I was curious to see how various prognosticators—specifically, FiveThirtyEight and The Economist's models, and the PredictIt prediction markets—did on predicting the state-by-state (plus the District of Columbia) results of the recent U.S. presidential election.

Mathematical Sidebar

There are various ways to evaluate probabilistic predictions, but my favorite is to use the logarithmic score: the logarithm of the probability assigned to the right answer. (Logs of probabilities are negative numbers, but negating everything to get positive numbers doesn't change anything interesting—you just minimize instead of maximizing—such that I tend to mentally conflate the log and the negative-log.)

The reason I think the logarithmic score is best is because it has one essential property, one cool property, and a meaningful interpretation.

The essential property is that it incentivizes you to report your actual probabilities: if something actually happens 80% of the time, you get the best expected score by giving it probability 0.8.

The cool property is that it doesn't matter how you chop up your observations: because conjunction of probabilities is multiplication and the logarithm maps multiplication to addition, we get the same score whether we consider "A&B" as one event, or separately add up the scores of "A" and "B, given A".

(Although as David Schneider-Joseph and Oscar Cunningham pointed out in response to the originally published version of this post, my calculations later in this post add up scores of state-by-state predictions as if the state results were independent events, but this independence assumption is not realistic. Sorry.)

The interpretation is that the logarithmic score represents the length of the message you would need to encode the actual outcome using a code optimized for your model. (Er, the negation of the score represents the length—there's that conflation.)

Methodology

So, the election results aren't really final until all the states certify their results—or perhaps, when the Electoral College meets. The Trump campaign has filed lawsuits disputing results in a number of states, and at time of writing, ABC News's map has no call for Alaska, Arizona, North Carolina, and Georgia. But for the purposes of this blog post, I'm just going to go with the current colors on Politico's map, and assume that Biden wins Arizona and Georgia and that Trump wins Alaska and North Carolina.

At around 20:10 Pacific time on 2 November, Election Day Eve, I downloaded the model-output ZIP files from FiveThirtyEight and The Economist. Then today, I looked up 2 November prices for the Democrat-wins contracts in the line-graphs on the pages for the PredictIt "Which party will win this-and-such-state in the 2020 presidential election?" markets. (Both FiveThirtyEight and The Economist would later issue final updates on Election Day, but I'm assuming it couldn't have changed much, and it's a fairer comparison to the PredictIt bucketed-by-day line graphs to use the model outputs from Election Day Eve.)

I got The Economist's per-state Biden-victory probabilities from the projected_win_prob column in /output/site_data//state_averages_and_predictions_topline.csv in The Economist's zipfile, and FiveThirtyEight's per-state Biden-victory probabilities from the winstate_chal column in /election-forecasts-2020/presidential_state_toplines_2020.csv in FiveThirtyEight's zipfile.

I'm only using the states' at-large results and ignoring the thing where Nebraska and Maine do some of their electoral votes by Congressional district.

I'm using the convention of giving probabilities in terms of the event of interest being "Biden/Democrat wins" rather than "Trump/Republican wins", because that's what The Economist seems to be giving me, but I appreciate that the FiveThirtyEight spreadsheet has separate winstate_inc (incumbent), winstate_chal (challenger), and winstate_3rd (thirty-party) columns. (Although the winstate_3rd column is blank!!) Using "incumbent wins" as the event actually seems like a more natural Schelling-point convention to me?—but it doesn't matter.

I wrote a Python script to calculate the scores. The main function looks like this:

def score():
    scores = {"fivethirtyeight": 0, "economist": 0, "predictit": 0}
    swinglike_only_scores = {
        "fivethirtyeight": 0,
        "economist": 0,
        "predictit": 0,
    }
    losses = {}

    swinglikes = [sr for sr in state_results if is_swinglike(sr)]
    print(
        "{} states ({}) are swinglike".format(
            len(swinglikes),
            ','.join(sr.state for sr in swinglikes)
        )
    )

    for state_result in state_results:
        if state_result.actual is None:
            continue

        for predictor in scores.keys():

            if state_result.actual:  # Biden win
                probability = getattr(state_result, predictor)
            else:
                probability = 1 - getattr(state_result, predictor)

            subscore = log2(probability)

            scores[predictor] += subscore

            if is_swinglike(state_result):
                swinglike_only_scores[predictor] += subscore

            losses[(predictor, state_result.state)] = subscore

    return scores, swinglike_only_scores, losses

(Full source code, including inline data.)

Getting the numbers from both spreadsheets and the browser line graphs into my script involved a lot of copy-pasting/Emacs-keyboard-macros/typing, and I didn't exhaustively double-check everything, so it's possible I made some mistakes, but I hope that I didn't, because that would be embarrassing.

Results

Scoring over all the states, The Economist did the best, with a score of about −7.51 bits, followed by FiveThirtyEight at −9.24 bits, followed by PredictIt at −12.14.

But I was worried that this methodology might privilege the statistical models over the markets, because the models (particularly The Economist, which was the most confident across the board) might be eking out more points by assigning very low/high probabilities to "safe" states in ways that the PredictIt markets won't: intuitively, I suspect that the fact that someone is willing to scoop up Biden-takes-Alabama contracts at 2¢ each may not be capturing the full power of what the market can do on harder questions. So I also calculated the scores on just the 12 "swing-like" states (Alaska, Arizona, Florida, Georgia, Iowa, Montana, North Carolina, New Hampshire, Nevada, Ohio, Pennsylvania, and Texas) where at least two of our three predictors gave a probability between 0.1 and 0.9.

On just the "swing-like" states, the results are a lot more even, with The Economist at −7.36 bits, PredictIt at −7.40, and FiveThirtyEight at −7.89.

The five largest predictive "misses" were The Economist and FiveThirtyEight on Florida (the models leaned Biden but the voters said Trump, giving the models log scores of −2.22 and −1.66, respectively), The Economist and FiveThirtyEight on North Carolina (same story for scores of −1.60 and −1.46), and PredictIt on Georgia (the market leaned Trump, but we think the voters are saying Biden for a score of −1.32).

Optimized Propaganda with Bayesian Networks: Comment on “Articulating Lay Theories Through Graphical Models”

(originally published at Less Wrong)

Derek Powell, Kara Weisman, and Ellen M. Markman's "Articulating Lay Theories Through Graphical Models: A Study of Beliefs Surrounding Vaccination Decisions" (a conference paper from CogSci 2018) represents an exciting advance in marketing research, showing how to use causal graphical models to study why ordinary people have the beliefs they do, and how to intervene to make them be less wrong.

The specific case our authors examine is that of childhood vaccination decisions: some parents don't give their babies the recommended vaccines, because they're afraid that vaccines cause autism. (Not true.) This is pretty bad—not only are those unvaccinated kids more likely to get sick themselves, but declining vaccination rates undermine the population's herd immunity, leading to new outbreaks of highly-contagious diseases like the measles in regions where they were once eradicated.

What's wrong with these parents, huh?! But that doesn't have to just be a rhetorical question—Powell et al. show how we can use statistics to make the rhetorical hypophorical and model specifically what's wrong with these people! Realistically, people aren't going to just have a raw, "atomic" dislike of vaccination for no reason: parents who refuse to vaccinate their children do so because they're (irrationally) afraid of giving their kids autism, and not afraid enough of letting their kids get infectious diseases. Nor are beliefs about vaccine effectiveness or side-effects uncaused, but instead depend on other beliefs.

To unravel the structure of the web of beliefs, our authors got Amazon Mechanical Turk participants to take surveys about vaccination-related beliefs, rating statements like "Natural things are always better than synthetic alternatives" or "Parents should trust a doctor's advice even if it goes against their intuitions" on a 7-point Likert-like scale from "Strongly Agree" to "Strongly Disagree".

Throwing some off-the-shelf Bayes-net structure-learning software at a training set from the survey data, plus some ancillary assumptions (more-general "theory" beliefs like "skepticism of medical authorities" can cause more-specific "claim" beliefs like "vaccines have harmful additives", but not vice versa) produces a range of probabilistic models that can be depicted with graphs where nodes representing the different beliefs are connected by arrows that show which beliefs "cause" others: an arrow from a naturalism node (in this context, denoting a worldview that prefers natural over synthetic things) to a parental expertise node means that people think parents know best because they think that nature is good, not the other way around.

Learning these kinds of models is feasible because not all possible causal relationships are consistent with the data: if \(A\) and \(B\) are statistically independent of each other, but each dependent with \(C\) (and are conditionally dependent given the value of \(C\)), it's kind of hard to make sense of this except to posit that \(A\) and \(B\) are causes with the common effect \(C\).

Simpler models with fewer arrows might sacrifice a little bit of predictive accuracy for the benefit of being more intelligible to humans. Powell et al. ended up choosing a model that can predict responses from the test set at r = .825, explaining 68.1% of the variance. Not bad?!—check out the full 14-node graph in Figure 2 on page 4 of the PDF.

Causal graphs are useful as a guide for planning interventions: the graph encodes predictions about what would happen if you changed some of the variables. Our authors point out that since previous work showed that people's beliefs about vaccine dangers were difficult to influence, that suggests trying to intervene on the other parents of the intent-to-vaccinate node in the model: if the hoi polloi won't listen to you when you tell them the costs are minimal (vaccines are safe), instead tell them about the benefits (diseases are really bad and vaccines prevent disease).

To make sure I really understand this, I want to adapt it into a simpler example with made-up numbers where I can do the arithmetic myself. Let me consider a graph with just three nodes—

vaccines are safe → vaccinate against measles ← measles are dangerous

Suppose this represents a structural equation model where an anti-vaxxer-leaning parent-to-be's propensity-to-vaccinate-against-measles \(C\) is expressed in terms of belief-in-vaccine-safety \(A\) and belief-in-measles-danger \(B\) as—

$$C = 0.7 \cdot A + 0.3 \cdot B $$

And suppose that we're a public health authority trying to decide whether to spend our budget (or what's left of it after recent funding cuts) on a public education initiative that will increase \(A\) by 0.1, or one that will increase \(B\) by 0.3.

We should choose the program that intervenes on \(B\), because \((0.3)(0.3) = 0.09\) is bigger than \((0.7)(0.1) = 0.07\). That's actionable advice that we couldn't have derived without a quantitative model of how the lay audience thinks. Exciting!

At this point, some readers may be wondering why I've described this work as "marketing research" about constructing "optimized propaganda." A couple of those words usually have negative connotations, but educating people about the importance of vaccines is a positive thing. What gives?

The thing is, "Learn the causal graph of why they think that and compute how to intervene on it to make them think something else" is a symmetric weapon—a fully general persuasive technique that doesn't depend on whether the thing you're trying to convince them of is true.

In my simplified example, the choice to intervene on \(B\) was based on numerical assumptions that amount to the claim that it's sufficiently easier to change \(B\) than it is to change \(A\), such that intervening on \(B\) is more effective at changing \(C\) than intervening on \(A\) (even though \(C\) depends on \(A\) more than it does on \(B\)). But this methodology is completely indifferent to what \(A\), \(B\), and \(C\) mean. It would have worked just as well, and for the same reasons if the graph had been—

Coca-Cola isn't unhealthy → drink Coca-Cola ← Coca-Cola tastes great

Suppose that we're advertising executives for the Coca-Cola Company trying to decide how to spend our budget (or what's left of it after recent funding cuts). If consumers won't listen to us when we tell them the costs of drinking Coke are minimal (lying that it isn't unhealthy), we should instead tell them about the benefits (Coke tastes good).

Or with different assumptions about the parameters—maybe \(C = 0.8 \cdot A + 0.2 \cdot B\) actually—then intervening to increase belief in "Coca-Cola isn't unhealthy" would be the right move (because \((0.8)(0.1) = 0.08 > 0.06 = (0.2)(0.3)\)). The marketing algorithm that just computes what belief changes will flip the decision node, doesn't have any way to notice or care whether those belief changes are in the direction of more or less accuracy.

To be clear—and I really shouldn't have to say this—this is not a criticism of Powell–Weisman–Markman's research! The "Learn the causal graph of why they think that" methodology is genuinely really cool! It doesn't have to be deployed as a marketing algorithm: the process of figuring out which belief change would flip some downstream node is the same thing as what we call locating a crux.1 The difference is just a matter of forwards or backwards direction: whether you first figure out if the measles vaccine or Coca-Cola are safe and then use whatever answer you come up with to guide your decision, or whether you write the bottom line first.

Of course, most people on most issues don't have the time or expertise to do their own research. For the most part, we can only hope that the sources we trust as authorities are doing their best to use their limited bandwidth to keep us genuinely informed, rather than merely computing what signals to emit in order to control our decisions.

If that's not true, we might be in trouble—perhaps increasingly so, if technological developments grant new advantages to the propagation of disinformation over the discernment of truth. In a possible future world where most words are produced by AIs running a "Learn the causal graph of why they think that and intervene on it to make them think something else" algorithm hooked up to a next-generation GPT, even reading plain text from an untrusted source could be dangerous.


  1. Thanks to Anna Salamon for this observation. 

Comment on “Endogenous Epistemic Factionalization”

(originally published at Less Wrong)

In "Endogenous Epistemic Factionalization" (due in a forthcoming issue of the philosophy-of-science journal Synthese), James Owen Weatherall and Cailin O'Connor propose a possible answer to the question of why people form factions that disagree on multiple subjects.

The existence of persistent disagreements is already kind of a puzzle from a Bayesian perspective. There's only one reality. If everyone is honestly trying to get the right answer and we can all talk to each other, then we should converge on the right answer (or an answer that is less wrong given the evidence we have). The fact that we can't do it is, or should be, an embarrassment to our species. And the existence of correlated persistent disagreements—when not only do I say "top" when you say "bottom" even after we've gone over all the arguments for whether it is in fact the case that top or bottom, but furthermore, the fact that I said "top" lets you predict that I'll probably say "cold" rather than "hot" even before we go over the arguments for that, is an atrocity. (Not hyperbole. Thousands of people are dying horrible suffocation deaths because we can't figure out the optimal response to a new kind of coronavirus.)

Correlations between beliefs are often attributed to ideology or tribalism: if I believe that Markets Are the Answer, I'm likely to propose Market-based solutions to all sorts of seemingly-unrelated social problems, and if I'm loyal to the Green tribe, I'm likely to selectively censor my thoughts in order to fit the Green party line. But ideology can't explain correlated disagreements on unrelated topics that the content of the ideology is silent on, and tribalism can't explain correlated disagreements on narrow, technical topics that aren't tribal shibboleths.

In this paper, Weatherall and O'Connor exhibit a toy model that proposes a simple mechanism that can explain correlated disagreement: if agents disbelieve in evidence presented by those with sufficiently dissimilar beliefs, factions emerge, even though everyone is honestly reporting their observations and updating on what they are told (to the extent that they believe it). The paper didn't seem to provide source code for the simulations it describes, so I followed along in Python. (Replication!)

In each round of the model, our little Bayesian agents choose between repeatedly performing one of two actions, A or B, that can "succeed" or "fail." A is a fair coin: it succeeds exactly half the time. As far as our agents know, B is either slightly better or slightly worse: the per-action probability of success is either 0.5 + ɛ or 0.5 − ɛ, for some ɛ (a parameter to the simulation). But secretly, we the simulation authors know that B is better.

import random

ε = 0.01

def b():
    return random.random() < 0.5 + ε

The agents start out with a uniformly random probability that B is better. The ones who currently believe that A is better, repeatedly do A (and don't learn anything, because they already know that A is exactly a coinflip). The ones who currently believe that B is better, repeatedly do B, but keep track of and publish their results in order to help everyone figure out whether B is slightly better or slightly worse than a coinflip.

class Agent:
    ...

    def experiment(self):
        results = [b() for _ in range(self.trial_count)]
        return results

If \(H_{+}\) represents the hypothesis that B is better than A, and \(H_{-}\) represents the hypothesis that B is worse, then Bayes's theorem says

$$P(H_{+}|E) = \frac{P(E|H_{+})P(H_{+})}{P(E|H_{+})P(H_{+}) + P(E|H_{-})P(H_{-})}$$

where E is the record of how many successes we got in how many times we tried action B. The likelihoods \(P(E|H_{+})\) and \(P(E|H_{-})\) can be calculated from the probability mass function of the binomial distribution, so the agents have all the information they need to update their beliefs based on experiments with B.

from math import factorial

def binomial(p, n, k):
    return (
        factorial(n) / (factorial(k) * factorial(n - k)) *
        p**k * (1 - p)**(n - k)
    )

class Agent:
    ...

    def pure_update(self, credence, hits, trials):
        raw_posterior_good = binomial(0.5 + ε, trials, hits) * credence
        raw_posterior_bad = binomial(0.5 - ε, trials, hits) * (1 - credence)
        normalizing_factor = raw_posterior_good + raw_posterior_bad
        return raw_posterior_good / normalizing_factor

Except in order to study the emergence of clustering among multiple beliefs, we should actually have our agents face multiple "A or B" dilemmas, representing beliefs about unrelated questions. (In each case, B will again be better, but the agents don't start out knowing that.) I chose three questions/beliefs, because that's all I can fit in a pretty 3D scatterplot.

If all the agents update on the experimental results published by the agents who do B, they quickly learn that B is better for all three questions. If we make a pretty 3D scatterplot where each dimension represents the probability that B is better for one of the dilemmas, then the points converge over time to the [1.0, 1.0, 1.0] "corner of Truth", even though they started out uniformly distributed all over the space.

two 3D scatterplots: agents' beliefs starting uniformly scattered on the left, converging tightly to the [1,1,1] "corner of Truth" on the right

But suppose the agents don't trust each other's reports. ("Sure, she says she performed \(B_2\) 50 times and observed 26 successes, but she also believes that \(B_1\) is better than \(A_1\), which is crazy. Are we sure she didn't just make up those 50 trials of \(B_2\)?") Specifically, our agents assign a probability that a report is made-up (and therefore should not be updated on) in proportion to their distance from the reporter in our three-dimensional beliefspace, and a "mistrust factor" (a parameter to the simulation).

from math import sqrt

def euclidean_distance(v, w):
    return sqrt(sum((v[i] - w[i]) ** 2 for i in range(len(v))))

class Agent:
    ...

    def discount_factor(self, reporter_credences):
        return min(
            1, self.mistrust * euclidean_distance(self.credences, reporter_credences)
        )

    def update(self, question, hits, trials, reporter_credences):
        discount = self.discount_factor(reporter_credences)
        posterior = self.pure_update(self.credences[question], hits, trials)
        self.credences[question] = (
            discount * self.credences[question] + (1 - discount) * posterior
        )

(Um, the paper itself actually uses a slightly more complicated mistrust calculation that also takes into account the agent's prior probability of the evidence, but I didn't quite understand the motivation for that, so I'm going with my version. I don't think the grand moral is affected.)

Then we can simulate what happens if the distrustful agents do many rounds of experiments and talk to each other—

def summarize_experiment(results):
    return (len([r for r in results if r]), len(results))

def simulation(
    agent_count,  # number of agents
    question_count,  # number of questions
    round_count,  # number of rounds
    trial_count,  # number of trials per round
    mistrust,  # mistrust factor
):
    agents = [
        Agent(
            [random.random() for _ in range(question_count)],
            trial_count=trial_count,
            mistrust=mistrust,
        )
        for i in range(agent_count)
    ]

    for _ in range(round_count):
        for question in range(question_count):
            experiments = []
            for agent in agents:
                if agent.credences[question] >= 0.5:
                    experiments.append(
                        (summarize_experiment(agent.experiment()), agent.credences)
                    )
            for agent in agents:
                for experiment, reporter_credences in experiments:
                    hits, trials = experiment
                    agent.update(
                        question,
                        hits,
                        trials,
                        reporter_credences,
                    )

    return agents

Depending on the exact parameters, we're likely to get a result that "looks like" this agent_count=200, round_count=20, question_count=3, trial_count=50, mistrust=2 run—

3D scatterplot of agents' beliefs after 20 rounds, split into color-coded clusters instead of all converging on the red "corner of Truth" point

Some of the agents (depicted in red) have successfully converged on the corner of Truth, but the others have polarized into factions that are all wrong about something. (The colors in the pretty 3D scatterplot are a k-means clustering for k := 8.) On average, evidence pushes our agents towards Truth—note the linearity of the blue and purple points, illustrating convergence on two out of the three problems—but agents who erroneously believe that A is better (due to some combination of a bad initial credence and unlucky experimental results that failed to reveal B's ε "edge" in the sample size allotted) can end up too far away to trust those who are gathering evidence for, and correctly converging on, the superiority of B.

Our authors wrap up:

[T]his result is especially notable because there is something reasonable about ignoring evidence generated by those you do not trust—particularly if you do not trust them on account of their past epistemic failures. It would be irresponsible for scientists to update on evidence produced by known quacks. And furthermore, there is something reasonable about deciding who is trustworthy by looking at their beliefs. From my point of view, someone who has regularly come to hold beliefs that diverge from mine looks like an unreliable source of information. In other words, the updating strategy used by our agents is defensible. But, when used on the community level, it seriously undermines the accuracy of beliefs.

I think the moral here is slightly off. The specific something reasonable about ignoring evidence generated by those you do not trust on account of their beliefs, is the assumption that those who have beliefs you disagree with are following a process that produces systematically misleading evidence. In this model, that assumption is just wrong. The problem isn't that the updating strategy used by our agents is individually "defensible" (what does that mean?) but produces inaccuracy "when used on the community level" (what does that mean?); the problem is that you get the wrong answer if your degree of trust doesn't match agents' actual trustworthiness. Still, it's enlighteningly disturbing to see specifically how the "distrust those who disagree" heuristic descends into the madness of factions.

(Full source code.)

Heads I Win, Tails?—Never Heard of Her; Or, Selective Reporting and the Tragedy of the Green Rationalists

(originally published at Less Wrong)

Followup to: What Evidence Filtered Evidence?

In "What Evidence Filtered Evidence?", we are asked to consider a scenario involving a coin that is either biased to land Heads 2/3rds of the time, or Tails 2/3rds of the time. Observing Heads is 1 bit of evidence for the coin being Heads-biased (because the Heads-biased coin lands Heads with probability 2/3, the Tails-biased coin does so with probability 1/3, the likelihood ratio of these is \(\frac{2/3}{1/3} = 2\), and \(\log_{2} 2 = 1\)), and analogously and respectively for Tails.

If such a coin is flipped ten times by someone who doesn't make literally false statements, who then reports that the 4th, 6th, and 9th flips came up Heads, then the update to our beliefs about the coin depends on what algorithm the not-lying1 reporter used to decide to report those flips in particular. If they always report the 4th, 6th, and 9th flips independently of the flip outcomes—if there's no evidential entanglement between the flip outcomes and the choice of which flips get reported—then reported flip-outcomes can be treated the same as flips you observed yourself: three Headses is 3 * 1 = 3 bits of evidence in favor of the hypothesis that the coin is Heads-biased. (So if we were initially 50:50 on the question of which way the coin is biased, our posterior odds after collecting 3 bits of evidence for a Heads-biased coin would be \(2^3:1\) = 8:1, or a probability of 8/(1 + 8) ≈ 0.89 that the coin is Heads-biased.)

On the other hand, if the reporter mentions only and exactly the flips that came out Heads, then we can infer that the other 7 flips came out Tails (if they didn't, the reporter would have mentioned them), giving us posterior odds of \(2^3:2^7\) = 1:16, or a probability of around 0.06 that the coin is Heads-biased.

So far, so standard. (You did read the Sequences, right??) What I'd like to emphasize about this scenario today, however, is that while a Bayesian reasoner who knows the non-lying reporter's algorithm of what flips to report will never be misled by the selective reporting of flips, a Bayesian with mistaken beliefs about the reporter's decision algorithm can be misled quite badly: compare the 0.89 and 0.06 probabilities we just derived given the same reported outcomes, but different assumptions about the reporting algorithm.

If the coin gets flipped a sufficiently large number of times, a reporter whom you trust to be impartial (but isn't), can make you believe anything she wants without ever telling a single lie, just with appropriate selective reporting. Imagine a very biased coin that comes up Heads 99% of the time. If it gets flipped ten thousand times, 100 of those flips will be Tails (in expectation), giving a selective reporter plenty of examples to point to if she wants to convince you that the coin is extremely Tails-biased.

Toy models about biased coins are instructive for constructing examples with explicitly calculable probabilities, but the same structure applies to any real-world situation where you're receiving evidence from other agents, and you have uncertainty about what algorithm is being used to determine what reports get to you. Reality is like the coin's bias; evidence and arguments are like the outcome of a particular flip. Wrong theories will still have some valid arguments and evidence supporting them (as even a very Heads-biased coin will come up Tails sometimes), but theories that are less wrong will have more.

If selective reporting is mostly due to the idiosyncratic bad intent of rare malicious actors, then you might hope for safety in (the law of large) numbers: if Helga in particular is systematically more likely to report Headses than Tailses that she sees, then her flip reports will diverge from everyone else's, and you can take that into account when reading Helga's reports. On the other hand, if selective reporting is mostly due to systemic structural factors that result in correlated selective reporting even among well-intentioned people who are being honest as best they know how,2 then you might have a more serious problem.

"A Fable of Science and Politics" depicts a fictional underground Society polarized between two partisan factions, the Blues and the Greens. "[T]here is a 'Blue' and a 'Green' position on almost every contemporary issue of political or cultural importance." If human brains consistently understood the is/ought distinction, then political or cultural alignment with the Blue or Green agenda wouldn't distort people's beliefs about reality. Unfortunately ... humans. (I'm not even going to finish the sentence.)

Reality itself isn't on anyone's side, but any particular fact, argument, sign, or portent might just so happen to be more easily construed as "supporting" the Blues or the Greens. The Blues want stronger marriage laws; the Greens want no-fault divorce. An evolutionary psychologist investigating effects of kin-recognition mechanisms on child abuse by stepparents might aspire to scientific objectivity, but being objective and staying objective is difficult when you're embedded in an intelligent social web in which in your work is going to be predictably championed by Blues and reviled by Greens.

Let's make another toy model to try to understand the resulting distortions on the Undergrounders' collective epistemology. Suppose Reality is a coin—no, not a coin, a three-sided die,3 with faces colored blue, green, and gray. One-third of the time it comes up blue (representing a fact that is more easily construed as supporting the Blue narrative), one-third of the time it comes up green (representing a fact that is more easily construed as supporting the Green narrative), and one-third of the time it comes up gray (representing a fact that not even the worst ideologues know how to spin as "supporting" their side).

Suppose each faction has social-punishment mechanisms enforcing consensus internally. Without loss of generality, take the Greens (with the understanding that everything that follows goes just the same if you swap "Green" for "Blue" and vice versa).4 People observe rolls of the die of Reality, and can freely choose what rolls to report—except a resident of a Green city who reports more than 1 blue roll for every 3 green rolls is assumed to be a secret Blue Bad Guy, and faces increasing social punishment as their ratio of reported green to blue rolls falls below 3:1. (Reporting gray rolls is always safe.)

The punishment is typically informal: there's no official censorship from Green-controlled local governments, just a visible incentive gradient made out of social-media pile-ons, denied promotions, lost friends and mating opportunities, increased risk of being involuntarily committed to psychiatric prison,5 &c. Even people who privately agree with dissident speech might participate in punishing it, the better to evade punishment themselves.

This scenario presents a problem for people who live in Green cities who want to make and share accurate models of reality. It's impossible to report every die roll (the only 1:1 scale map of the territory, is the territory itself), but it seems clear that the most generally useful models—the ones you would expect arbitrary AIs to come up with—aren't going to be sensitive to which facts are "blue" or "green". The reports of aspiring epistemic rationalists who are just trying to make sense of the world will end up being about one-third blue, one-third green, and one-third gray, matching the distribution of the Reality die.

From the perspective of ordinary nice smart Green citizens who have not been trained in the Way, these reports look unthinkably Blue. Aspiring epistemic rationalists who are actually paying attention can easily distinguish Blue partisans from actual truthseekers,6 but the social-punishment machinery can't process more than five words at a time. The social consequences of being an actual Blue Bad Guy, or just an honest nerd who doesn't know when to keep her stupid trap shut, are the same.

In this scenario,7 public opinion within a subculture or community in a Green area is constrained by the 3:1 (green:blue) "Overton ratio." In particular, under these conditions, it's impossible to have a rationalist community—at least the most naïve conception of such. If your marketing literature says, "Speak the truth, even if your voice trembles," but all the savvy high-status people's actual reporting algorithm is, "Speak the truth, except when that would cause the local social-punishment machinery to mark me as a Blue Bad Guy and hurt me and any people or institutions I'm associated with—in which case, tell the most convenient lie-of-omission", then smart sincere idealists who have internalized your marketing literature as a moral ideal and trust the community to implement that ideal, are going to be misled by the community's stated beliefs—and confused at some of the pushback they get when submitting reports with a 1:1:1 blue:green:gray ratio.

Well, misled to some extent—maybe not much! In the absence of an Oracle AI (or a competing rationalist community in Blue territory) to compare notes with, then it's not clear how one could get a better map than trusting what the "green rationalists" say. With a few more made-up modeling assumptions, we can quantify the distortion introduced by the Overton-ratio constraint, which will hopefully help develop an intuition for how large of a problem this sort of thing might be in real life.

Imagine that Society needs to make a decision about an Issue (like a question about divorce law or merchant taxes). Suppose that the facts relevant to making optimal decisions about an Issue are represented by nine rolls of the Reality die, and that the quality (utility) of Society's decision is proportional to the (base-two logarithm) entropy of the distribution of what facts get heard and discussed.8

The maximum achievable decision quality is \(\log_{2} 9\) ≈ 3.17.

On average, Green partisans will find 3 "green" facts9 and 3 "gray" facts to report, and mercilessly stonewall anyone who tries to report any "blue" facts, for a decision quality of \(\log_{2} 6\) ≈ 2.58.

On average, the Overton-constrained rationalists will report the same 3 "green" and 3 "gray" facts, but something interesting happens with "blue" facts: each individual can only afford to report one "blue" fact without blowing their Overton budget—but it doesn't have to be the same fact for each person. Reports of all 3 (on average) blue rolls get to enter the public discussion, but get mentioned (cited, retweeted, &c.) 1/3 as often as green or gray rolls, in accordance with the Overton ratio. So it turns out that the constrained rationalists end up with a decision quality of \(\frac{6}{7} \log_{2} 7 + \frac{1}{7} \log_{2} 21\) ≈ 3.03,10 significantly better than the Green partisans—but still falling short of the theoretical ideal where all the relevant facts get their due attention.

If it's just not pragmatic to expect people to defy their incentives, is this the best we can do? Accept a somewhat distorted state of discourse, forever?

At least one partial remedy seems apparent. Recall from our original coin-flipping example that a Bayesian who knows what the filtering process looks like, can take it into account and make the correct update. If you're filtering your evidence to avoid social punishment, but it's possible to clue in your fellow rationalists to your filtering algorithm without triggering the social-punishment machinery—you mustn't assume that everyone already knows!—that's potentially a big win. In other words, blatant cherry-picking is the best kind!


  1. I don't quite want to use the word honest here. 

  2. And it turns out that knowing how to be honest is much more work than one might initially think. You have read the Sequences, right?! 

  3. For lack of an appropriate Platonic solid in three-dimensional space, maybe imagine tossing a triangle in two-dimensional space?? 

  4. As an author, I'm facing some conflicting desiderata in my color choices here. I want to say "Blues and Greens" in that order for consistency with "A Fable of Science and Politics" (and other classics from the Sequences). Then when making an arbitrary choice to talk in terms of one of the factions in order to avoid cluttering the exposition, you might have expected me to say "Without loss of generality, take the Blues," because the first item in a sequence ("Blues" in "Blues and Greens") is a more of a Schelling point than the second, or last, item. But I don't want to take the Blues, because that color choice has other associations that I'm trying to avoid right now: if I said "take the Blues", I fear many readers would assume that I'm trying to directly push a partisan point about soft censorship and preference-falsification social pressures in liberal/left-leaning subcultures in the contemporary United States. To be fair, it's true that soft censorship and preference-falsification social pressures in liberal/left-leaning subcultures in the contemporary United States are, historically, what inspired me, personally, to write this post. It's okay for you to notice that! But I'm trying to talk about the general mechanisms that generate this class of distortions on a Society's collective epistemology, independently of which faction or which ideology happens to be "on top" in a particular place and time. If I'm doing my job right, then my analogue in a "nearby" Everett branch whose local subculture was as "right-polarized" as my Berkeley environment is "left-polarized", would have written a post making the same arguments. 

  5. Okay, they market themselves as psychiatric "hospitals", but let's not be confused by misleading labels

  6. Or rather, aspiring epistemic rationalists can do a decent job of assessing the extent to which someone is exhibiting truth-tracking behavior, or Blue-partisan behavior. Obviously, people who are consciously trying to seek truth, are not necessarily going to succeed at overcoming bias, and attempts to correct for the "pro-Green" distortionary forces being discussed in this parable could easily veer into "pro-Blue" over-correction. 

  7. Please be appropriately skeptical about the real-world relevance of my made-up modeling assumptions! If it turned out that my choice of assumptions were (subconsciously) selected for the resulting conclusions about how bad evidence-filtering is, that would be really bad for the same reason that I'm claiming that evidence-filtering is really bad! 

  8. The entropy of a discrete probability distribution is maximized by the uniform distribution, in which all outcomes receive equal probability-mass. I only chose these "exactly nine equally-relevant facts/rolls" and "entropic utility" assumptions to make the arithmetic easy on me; a more realistic model might admit arbitrarily many facts into discussion of the Issue, but posit a distribution of facts/rolls with diminishing marginal relevance to Society's decision quality. 

  9. The scare quotes around the adjective "'green'" (&c.) when applied to the word "fact" (as opposed to a die roll outcome representing a fact in our toy model) are significant! The facts aren't actually on anyone's side! We're trying to model the distortions that arise from stupid humans thinking that the facts are on someone's side! This is sufficiently important—and difficult to remember—that I should probably repeat it until it becomes obnoxious! 

  10. You have three green slots, three gray slots, and three blue slots. You put three counters each on each of the green and gray slots, and one counter each on each of the blue slots. The frequencies of counters per slot is [3, 3, 3, 3, 3, 3, 1, 1, 1]. The total number of counters you put down is 3*6 + 3 = 18 + 3 = 21. To turn the frequencies into a probability distribution, you divide everything by 21, to get [1/7, 1/7, 1/7, 1/7, 1/7, 1/7, 1/21, 1/21, 1/21]. Then the entropy is \(6\cdot-\frac{1}{7}\log_{2}\frac{1}{7}+3\cdot-\frac{1}{21}\log_{2}\frac{1}{21}\), which simplifies to \(\frac{6}{7}\log_{2}7+\frac{1}{7}\log_{2}21\)

Brand Rust

2007–2016: "Of course I'm still fundamentally part of the Blue Team, like all non-evil people, but I genuinely think there are some decision-relevant facts about biology, economics, and statistics that folks may not have adequately taken into account!"

2017: "You know, maybe I'm just ... not part of the Blue Team? Maybe I can live with that?"

An Intuition on the Bayes-Structural Justification for Free Speech Norms

We can metaphorically (but like, hopefully it's a good metaphor) think of speech as being the sum of a positive-sum information-conveying component and a zero-sum social-control/memetic-warfare component. Coalitions of agents that allow their members to convey information amongst themselves will tend to outcompete coalitions that don't, because it's better for the coalition to be able to use all of the information it has.

Therefore, if we want the human species to better approximate a coalition of agents who act in accordance with the game-theoretic Bayes-structure of the universe, we want social norms that reward or at least not-punish information-conveying speech (so that other members of the coalition can learn from it if it's useful to them, and otherwise ignore it).

It's tempting to think that we should want social norms that punish the social-control/memetic-warfare component of speech, thereby reducing internal conflict within the coalition and forcing people's speech to mostly consist of information. This might be a good idea if the rules for punishing the social-control/memetic-warfare component are very clear and specific (e.g., no personal insults during a discussion about something that's not the person you want to insult), but it's alarmingly easy to get this wrong: you think you can punish generalized hate speech without any negative consequences, but you probably won't notice when members of the coalition begin to slowly gerrymander the hate speech category boundary in the service of their own values. Whoops!

Dreaming of Political Bayescraft

My old political philosophy: "Socially liberal, fiscally confused; I don't know how to run a goddamned country (and neither do you)."

Commentary: Pretty good, but not quite meta enough.

My new political philosophy: "Being smart is more important than being good (for humans). All ideologies are false; some are useful."

Commentary: Social design space is very large and very high-dimensional; the forces of memetic evolution are somewhat benevolent (all ideas that you've heard of have to be genuinely appealing to some feature of human psychology, or no one would have an incentive to tell you about them), but really smart people who know lots of science and lots of probability and game theory might be able to do better for themselves! Any time you find yourself being tempted to be loyal to an idea, it turns out that what you should actually be loyal to is whatever underlying feature of human psychology makes the idea look like a good idea; that way, you'll find it easier to fucking update when it turns out that the implementation of your favorite idea isn't as fun as you expected! This stance is itself, technically, loyalty to an idea, but hopefully it's a sufficiently meta idea to avoid running into the standard traps while also being sufficiently object-level to have easily-discoverable decision-relevant implications and not run afoul of the principle of ultrafinite recursion ("all infinite recursions are at most three levels deep").

Prescription II

that feel eighteen months post-Obergefell when you realize you missed your chance to be pro-civil-unions-with-all-the-same-legal-privileges but anti-calling-it-marriage while that position was still in the Overton window

(in keeping with the principle that it shouldn't be so exotic to want to protect people's freedom to do beautiful new things without necessarily thereby insisting on redefining existing words that already mean something else)

Missing Refutations

It looks like the opposing all-human team is winning the exhibition game of me and my it's-not-chess engine (as White) versus everyone in the office who (unlike me) actually knows something about chess (as Black). I mean, naïvely, my team is up a bishop right now, but our king is pretty exposed, and the principal variation that generated one of our recent moves (16. Bxb4 Bf5 17. Kd1 Qxd4+ 18. Kc1 Ng3 19. Qxc7 Nxh1) looks dreadful.

Real chess aficionados (chessters? chessies?) will laugh at me, but it actually took me a while to understand why Ng3 was in that principal variation (I might even have invoked the engine again to help). The position after Ng3 looks like

    a b c d e f g h
 8 ♜       ♜   ♚   
 7 ♟ ♟ ♟     ♟ ♟ ♟ 
 6                 
 5           ♝     
 4   ♗   ♛         
 3 ♙           ♞   
 2   ♙ ♕     ♙ ♙ ♙ 
 1 ♖ ♘ ♔     ♗   ♖

and—forgive me—I didn't understand why that wasn't refuted by fxg3 or hxg3; in my novice's utter blindness, I somehow failed to see the discovered attack on the white queen, the necessity of evading which allows the black knight to capture the white rook, and preparation for which was clearly the purpose of 16. ..Bf5 (insofar as we—anthropomorphically?—attribute purpose to a sequence of moves discovered by a minimax search algorithm which doesn't represent concepts like discovered attack anywhere).

It's the little things like this that reinforce my suspicion that human cognition doesn't really work most of the time, that to the extent that we have a technological civilization with nice things in it, it's a matter of a bunch of apes, some of whom are just barely generally-intelligent, happening to stumble into patterns of coöperation that happened to scale, and not a matter of anyone within the system actually understanding much. It's the current year, and the current year turns out to be not only congruent to 0 mod 2, but furthermore to 0 mod 4, which means outrage season is brewing again. I'm mostly pretty good at ignoring it, except for the undercurrent of contempt that I can't help but feel at how almost everyone seems to think it's okay for important decisions to be made in this way.

The thing about chess is that it's a very simple game. There are these thirty-two figurines on an eight-by-eight grid, and you take turns moving the figurines subject to a few rules with a well-specified objective in mind. Simple. If you're lazy like me, you can even write a computer program to do it for you. And if you hadn't already spent many, many hours of study and practice mastering the game yourself, what such a program will surely show you (as it showed me, though I expected as much) is that your unaided attempts to make good moves will get it wrong. The novice thinks: checkmate is the goal, giving check is good, capturing material is good; I should look for moves that do those things. And the novice will lose, badly. What you actually need to do is look for your best move given your opponent's best reply given your best counterreply given your opponent's best countercounterreply, and so on as far forward as you can afford to compute. It's not magic and it's not impossible, but it's subtle and it takes work.

And the thing about the real world is that it is unimaginably complicated. There are seven billion humans on this radius-6.37·\(10^6\) m planet, all haphazardly pursuing their own objectives subject to no rules except whatever high-level generalizations about the underlying fundamental physics strike you as sufficiently robust to be worth reifying as a rule. Unimaginable. And I feel like those appealing to the masses seeking leadership positions in the reigning institutions of governance don't properly appreciate this. You hear them or their boosters talk about how "we" obviously need to make college or healthcare free, or crush our terrorist foes, and to me it just sounds like so many shouts of "fxg3!" or "hxg3!". The goals are noble: of course we want America to be great, and for its people to be healthy and wealthy and knowledgeable and safe, just as a chess player wants to checkmate the opposing king. But the interventions that sound intuitively appealing aren't necessarily the same as the interventions that will actually work! "We" should hope for leaders with the courage and skepticism and competence to say, "No, wait, that won't work because of ...Bxc2," and continue the search for better alternatives for our nation and its children.

But if I can predict that that's not what "we" are going to choose, maybe I should continue searching for better alternative ways to spend my time than complaining about it.

An Education News Bulletin

Apparently a gang of extortionists calling themselves the "California state Bureau for Private Postsecondary Education" are threatening to shut down a number of organizations that provide assistance in learning to program, including App Academy, which I recently benefitted from attending. I could explain why the behavior of the BPPE is an outrage that must be opposed by anyone with a scrap of decency in their heart, but I'm too busy coding and counting my money.

Counterfactual Social Thought

I keep feeling like I need to study Bayes nets in order to clarify my thinking about society. (This is probably not standard advice given to aspiring young sociologists, but I'm trying not to care about that.) Ordinary political speech is full of claims about causality ("Policy X causes Y, which is bad!" "Of course Y is bad, but don't you see?—the real cause of Y is Z, and if you hadn't been brainwashed by the System, you'd see that!"), but human intuitions about causality are probably confused (and would be clarified by Pearl) much like our intuitions about evidence are confused (and are clarified by Bayes).

Almost every policy proposal is, implicitly, a counterfactual conditional. "We need to implement Policy A in order to protect B" means that if Policy A were implemented, then it would have beneficial effects on B. But most people with policy opinions aren't actually in a position to implement the changes they talk about. Insofar as you construe the function of thought as to select actions in order to optimize the world with respect to some preference ordering, having passionate opinions about issues you can't affect is kind of puzzling. In a small group, an individual voice can change the outcome: if I argue that our party of five should dine at this restaurant rather than that one, then my voice may well carry the day. But people often argue about priorities for an entire country of millions of people, vast and diverse beyond any individual's comprehension! What's that about?

(To be sure, you can come up with reasonable arguments why someone should concern themselves with large-scale politics: every collective effort requires the actions of many, so one might cooperate with a group rather than defect, precisely because bad things would happen if everyone defected; or, maybe some particular individual is exceptionally well-positioned to make a difference through their own actions; or, a tiny probability of having a large effect might be worthwhile in expectation; or, ... &c. Whether or not these are good arguments, I don't think they're an adequate explanation of what's actually going on inside most people's heads.)

I often find myself feeling angry and upset that the mainstream society around me doesn't reflect my values; I spend hours composing rhetoric and slogans about how our dominant forms of social organization are systematically flawed in knowable ways. I imagine the world being different—and only in brief moments of lucidity do I realize that what I'm doing is daydreaming, fantasizing. Thinking about social change doesn't feel like a mere fantasy in the way that thinking about how great it would be to have magical superpowers is obviously fantasy, but it is: thinking about good outcomes in the absence of actual planning about how to achieve those outcomes from the present state is wasted cognition except insofar as the thinking-about-good-outcomes is valuable for its own sake. Fantasy is a fine thing in moderation; it's not being able to reliably tell the difference between fantasy and reality that's dangerous. In the case of magical superpowers, the difference is obvious. In the case of the mainstream magically adopting my priorities, it's somehow not obvious; somehow I find it hard to stop thinking about worlds that are not my own. Why?

It's easy to tell an evolutionary psychology just-so story: precisely because arguing about politics actually is important in small groups like the ones our ancestors lived in during the environment of evolutionary adaptedness, I can't bring my brain to notice that things don't work the same way when you're one voice among three-times-ten-to-the-eighth. Whereas obviously-fantastical fantasy is just wireheading in the sense that it's a byproduct of imagination and preferring-certain-experiences, both of which are adaptive in themselves, but which together result in non-adaptive daydreaming; since this has a different etiology than political daydreaming, it's not surprising that it would have a different character ...

But that's just a story I made up; I'm not claiming it's actually true; most of the stories people make up aren't actually true.

Egoism as Defense Against a Life of Unending Heartbreak

Then the Dean understood what had puzzled him in Roark's manner.

"You know," he said, "You would sound much more convincing if you spoke as if you cared whether I agreed with you or not."

"That's true," said Roark. "I don't care whether you agree with me or not." He said it so simply that it did not sound defensive, it sounded like the statement of a fact which he noticed, puzzled, for the first time.

"You don't care what others think—which might be understandable. But you don't care even to make them think as you do?"

"No."

"But that's ... that's monstrous."

"Is it? Probably. I couldn't say."

In this passage from Ayn Rand's The Fountainhead, fictional character Howard Roark demonstrates a very important skill that I really need to learn—that of emotional indifference to arbitrary people's opinions: not the mere immunity of "It's okay that people now disagree with the manifest rightness of my Cause, because I know the forces of Good will win in the end," but the kind of outright indifference that I feel about, let's say, the amount of precipitation in Copenhagen in March 1957. Someone disagrees with the manifest rightness of my Cause? Sure, whatever—hey, did you see the latest Questionable Content?

I say this purely for pragmatic reasons. There's nothing philosophically noble about being narrowly selfish, about devoting the full force of one's attention to questions like "What do I want to study?" or "How am I going to make money?" rather than "Why are my ideological enemies so evil, and what can be done to stop them?" So if there's no inherent reason why scholarship or business are more worthy than activism, why explicitly renounce the activist frame of mind?

Because activism is painful and it doesn't work very well. I'm tired of hating Society (whatever that means) for not being what I wish it were. What good has it done me or anyone else, all the hours I've spent over the past five years, fuming and raving over how I have been wronged?

Better by far to focus on the tiny scrap of the world that I actually have control over: to bury my head in a subculture of my kind of people, to master valuable skills, to make piles of money and donate a fraction of it to organizations that are doing valuable work, rather than waste any more precious moments of thinking time being personally offended by all the evil in the world. If someone's interested in what I think, then I graciously welcome them to read this blog. If I see a cheap opportunity to possibly change someone's behavior for the better—to offer them a rationality tip, or glance at them disapprovingly when they violate a social norm that really ought to be kept intact—then I'll take it. But if not, then not.

Some might object that if everyone thought this way, then all social progress would halt: the moral progress of humankind depends on the selfless activists who sacrifice so much to change so little. To which I reply: sure. But not everyone thinks this way, and my relinquishing the cloud of moral outrage that's been making me so unhappy isn't going to change that. We all have our roles to play in this Great Romance of Determinism, and it's long past time for me to finally choose a role that works, rather than one that doesn't.