An Algorithmic Lucidity

a blog

Tag: discourse

Hazards of Selection Effects on Approved Information

In a busy, busy world, there's so much to read that no one could possibly keep up with it all. You can't not prioritize what you pay attention to and (even more so) what you respond to. Everyone and her dog tells herself a story that she wants to pay attention to "good" (true, useful) information and ignore "bad" (false, useless) information.

Keeping the story true turns out to be a harder problem than it sounds. Everyone and her dog knows that the map is not the territory, but the reason we need a whole slogan about it is because we never actually have unmediated access to the territory. Everything we think we know about the territory is actually just part of our map (the world-simulation our brains construct from sensory data), which makes it easy to lose track of whether your actions are improving the real territory, or just your view of it on your map.

For example, I like it when I have good ideas. It makes sense for me to like that. I endorse taking actions that will result in world-states in which I have good ideas.

The problem is that I might not be able to tell the difference between world-states in which I have good ideas, and world-states in which I think my ideas are good, but they're actually bad. Those two different states of the territory would look the same on my map.

If my brain's learning algorithms reinforce behaviors that lead to me having ideas that I think are good, then in addition to learning behaviors that make me have better ideas (like reading a book), I might also inadvertently pick up behaviors that prevent me from hearing about it if my ideas are bad (like silencing critics).

This might seem like an easy problem to solve, because the most basic manifestations of the problem are in fact pretty easy to solve. If I were to throw a crying fit and yell, "Critics bad! No one is allowed to criticize my ideas!" every time someone criticized my ideas, the problem with that would be pretty obvious to everyone and her dog, and I would stop getting invited to the salon.

But what if there were subtler manifestations of the problem, that weren't obvious to everyone and her dog? Then I might keep getting invited to the salon, and possibly even spread the covertly dysfunctional behavior to other salon members. (If they saw the behavior seeming to work for me, they might imitate it, and their brain's learning algorithms would reinforce it if it seemed to work for them.) What might those look like? Let's try to imagine.

Filtering Interlocutors

Goofusia: I don't see why you tolerate that distrustful witch Goody Osborne at your salon. Of course I understand the importance of criticism, which is an essential nutrient for any truthseeker. But you can acquire the nutrient without the downside of putting up with unpleasant people like her. At least, I can. I've already got plenty of perceptive critics in my life among my friends who want the truth, and know that I want the truth—who will assume my good faith, because they know my heart is in the right place.

Gallantina: But aren't your friends who know you want the truth selected for agreeing with you, over and above their being selected for being correct? If there were some crushing counterargument to your beliefs that would only be found by someone who didn't know that you want the truth and wouldn't assume good faith, how would you ever hear about it?

This one is subtle. Goofusia isn't throwing a crying fit every time a member of the salon criticizes her ideas. And indeed, you can't invite the whole world to your salon. You can't not do some sort of filtering. The question is whether salon invitations are being extended or withheld for "good" reasons (that promote the salon processing true and useful information) or "bad" reasons (that promote false or useless information).

The problem is that being friends with Goofusia and "know[ing] that [she and other salon members] want the truth" is a bad membership criterion, not a good one, because people who aren't friends with Goofusia and don't know that she wants the truth are likely to have different things to say. Even if Goofusia can answer all the critiques her friends can think of, that shouldn't give her confidence that her ideas are solid, if there are likely to exist serious critiques that wouldn't be independently reïnvented by the kinds of people who become Goofusia's friends.

The "nutrient" metaphor is a tell. Goofusia seems to be thinking of criticism as if it were a homogeneous ingredient necessary for a healthy epistemic environment, but that it doesn't particularly matter where it comes from. In analogy, it doesn't matter whether you get your allowance of potassium from bananas or potatoes or artificial supplements. If you find bananas and potatoes unpleasant, you can still take supplements and get your potassium that way; if you find Goody Osborne unpleasant, you can just talk to your friends who know you want the truth and get your criticism that way.

But unlike chemically uniform nutrients, criticism isn't homogeneous: different critics are differently equipped by virtue of their different intellectual backgrounds to notice different flaws in a piece of work. The purpose of criticism is not to virtuously endure being criticized; the purpose is to surface and fix every individual flaw. (If you independently got everything exactly right the first time, then there would be nothing for critics to do; it's just that that seems pretty unlikely if you're talking about anything remotely complicated. It would be hard to believe that such an unlikely-seeming thing had really happened without the toughest critics getting the chance to do their worst.)

"Knowing that (someone) wants the truth" is a particularly poor filter, because people who think that they have strong criticisms of your ideas are particularly likely to think that you don't want the truth. (Because, the reasoning would go, if you did want the truth, why would you propose such flawed ideas, instead of independently inventing the obvious-to-them criticism yourself and dropping the idea without telling anyone?) Refusing to talk to people who think that they have strong criticisms of your ideas is a bad thing to do if you care about your ideas being correct.

The selection effect is especially bad in situations where the fact that someone doesn't want the truth is relevant to the correct answer. Suppose Goofusia proposes that the salon buys cookies from a certain bakery—which happens to be owned by Goofusia's niece. If Goofusia's proposal was motivated by nepotism, that's probabilistically relevant to evaluating the quality of the proposal. (If the salon members aren't omniscient at evaluating bakery quality on the merits, then they can be deceived by recommendations made for reasons other than the merits.) The salon can debate back and forth about the costs and benefits of spending the salon's snack budget at the niece's bakery, but if no one present is capable of thinking "Maybe Goofusia is being nepotistic" (because anyone who could think that would never be invited to Goofusia's salon), that bodes poorly for the salon's prospects of understanding the true cost–benefit landscape of catering options.

Filtering Information Sources

Goofusia: One shouldn't have to be the sort of person who follows discourse in crappy filter-bubbles in order to understand what's happening. The Rev. Samuel Parris's news summary roundups are the sort of thing that lets me do that. Our salon should work like that if it's going to talk about the atheist threat and the witchcraft crisis. I don't want to have to read the awful corners of the internet where this is discussed all day. They do truthseeking far worse there.

Gallantina: But then you're turning your salon into a Rev. Parris filter bubble. Don't you want your salon members to be well-read? Are you trying to save time, or are you worried about being contaminated by ideas that haven't been processed and vetted by Rev. Parris?

This one is subtle, too. If Goofusia is busy and just doesn't have time to keep up with what the world is saying about atheism and witchcraft, it might very well make sense to delegate her information gathering to Rev. Parris. That way, she can get the benefits of being mostly up to speed on these issues without having to burn too many precious hours that could be spent studying more important things.

The problem is that the suggestion doesn't seem to be about personal time-saving. Rev. Parris is only one person; even if he tries to make his roundups reasonably comprehensive, he can't help but omit information in ways that reflect his own biases. (For he is presumably not perfectly free of bias, and if he didn't omit anything, there would be no time-saving value to his subscribers in being able to just read the roundup rather than having to read everything that Rev. Parris reads.) If some salon members are less busy than Goofusia and can afford to do their own varied primary source reading rather than delegating it all to Rev. Parris, Goofusia should welcome that—but instead, she seems to be suspicious of those who would "be the sort of person" who does that. Why?

The admonition that "They do truthseeking far worse there" is a tell. The implication seems to be that good truthseekers should prefer to only read material by other good truthseekers. Rev. Parris isn't just saving his subscribers time; he's protecting them from contamination, heroically taking up the burden of extracting information out of the dangerous ravings of non-truthseekers.

But it's not clear why such a risk of contamination should exist. Part of the timeless ideal of being well-read is that you're not supposed to believe everything you read. If I'm such a good truthseeker, then I should want to read everything I can about the topics I'm seeking the truth about. If the authors who publish such information aren't such good truthseekers as I am, I should take that into account when performing updates on the evidence they publish, rather than denying myself the evidence.

Information is transmitted across the physical universe through links of cause and effect. If Mr. Proctor is clear-sighted and reliable, then when he reports seeing a witch, I infer that there probably was a witch. If the correlation across possible worlds is strong enough—if I think Mr. Proctor reports witches when there are witches, and not when there aren't—then Mr. Proctor's word is almost as good as if I'd seen the witch myself. If Mr. Corey has poor eyesight and is of a less reliable character, I am less credulous about reported witch sightings from him, but if I don't face any particular time constraints, I'd still rather hear Mr. Corey's testimony, because the value of information to a Bayesian reasoner is always nonnegative. For example, Mr. Corey's report could corroborate information from other sources, even if it wouldn't be definitive on its own. (Even the fact that people sometimes lie doesn't fundamentally change the calculus, because the possibility of deception can be probabilistically "priced in".)

That's the theory, anyway. A potential reason to fear contamination from less-truthseeking sources is that perhaps the Bayesian ideal is too hard to practice and salon members are too prone to believe what they read. After all, many news sources have been adversarially optimized to corrupt and control their readers and make them less sane by seeing the world through ungrounded lenses.

But the means by which such sources manage to control their readers is precisely by capturing their trust and convincing them that they shouldn't want to read the awful corners of the internet where they do truthseeking far worse than here. Readers who have mastered multiple ungrounded lenses and can check them against each other can't be owned like that. If you can spare the time, being well-read is a more robust defense against the risk of getting caught in a bad filter bubble, than trying to find a good filter bubble and blocking all (presumptively malign) outside sources of influence. All the bad bubbles have to look good from the inside, too, or they wouldn't exist.

To some, the risk of being in a bad bubble that looks good may seem too theoretical or paranoid to take seriously. It's not like there are no objective indicators of filter quality. In analogy, the observation that dreaming people don't know that they're asleep, probably doesn't make you worry that you might be asleep and dreaming right now.

But it being obvious that you're not in one of the worst bubbles shouldn't give you much comfort. There are still selection effects on what information gets to you, if for no other reason that there aren't enough good truthseekers in the world to uniformly cover all the topics that a truthseeker might want to seek truth about. The sad fact is that people who write about atheism and witchcraft are disproportionately likely to be atheists or witches themselves, and therefore non-truthseeking. If your faith in truthseeking is so weak that you can't even risk hearing what non-truthseekers have to say, that necessarily limits your ability to predict and intervene on a world in which atheists and witches are real things in the physical universe that can do real harm (where you need to be able to model the things in order to figure out which interventions will reduce the harm).

Suppressing Information Sources

Goofusia: I caught Goody Osborne distributing pamphlets quoting the honest and candid and vulnerable reflections of Rev. Parris on guiding his flock, and just trying to somehow twist that into maximum anger and hatred. It seems quite clear to me what's going on in that pamphlet, and I think signal-boosting it is a pretty clear norm violation in my culture.

Gallantina: I read that pamphlet. It seemed like intellectually substantive satire of a public figure. If you missed the joke, it was making fun of an alleged tendency in Rev. Parris's sermons to contain sophisticated analyses of the causes of various social ills, and then at the last moment, veer away from the uncomfortable implications and blame it all on witches. If it's a norm violation to signal-boost satire of public figures, that's artificially making it harder for people to know about flaws in the work of those public figures.

This one is worse. Above, when Goofusia filtered who she talks to and what she reads for bad reasons, she was in an important sense only hurting herself. Other salon members who aren't sheltering themselves from information are unaffected by Goofusia's preference for selective ignorance, and can expect to defeat Goofusia in public debate if the need arises. The system as a whole is self-correcting.

The invocation of "norm violations" changes everything. Norms depend on collective enforcement. Declaring something a norm violation is much more serious than saying that you disagree with it or don't like it; it's expressing an intent to wield social punishment in order to maintain the norm. Merely bad ideas can be criticized, but ideas that are norm-violating to signal-boost are presumably not even to be seriously discussed. (Seriously discussing a work is signal-boosting it.) Norm-abiding group members are required to be ignorant of their details (or act as if they're ignorant).

Mandatory ignorance of anything seems bad for truthseeking. What is Goofusia thinking here? Why would this seem like a good idea to someone?

At a guess, the "maximum anger and hatred" description is load-bearing. Presumably the idea is that it's okay to calmly and politely criticize Rev. Parris's sermons; it's only sneering or expressing anger or hatred that is forbidden. If the salon's speech code only targets form and not content, the reasoning goes, then there's no risk of the salon missing out on important content.

The problem is that the line between form and content is blurrier than many would prefer to believe, because words mean things. You can't just swap in non-angry words for angry words without changing the meaning of a sentence. Maybe the distortion of meaning introduced by substituting nicer words is small, but then again, maybe it's large: the only person in a position to say is the author. People don't express anger and hatred for no reason. When they do, it's because they have reasons to think something is so bad that it deserves their anger and hatred. Are those good reasons or bad reasons? If it's norm-violating to talk about it, we'll never know.

Unless applied with the utmost stringent standards of evenhandedness and integrity, censorship of form quickly morphs into censorship of content, as heated criticism of the ingroup is construed as norm-violating, while equally heated criticism of the outgroup is unremarkable and passes without notice. It's one of those irregular verbs: I criticize; you sneer; she somehow twists into maximum anger and hatred.

The conjunction of "somehow" and "it seems quite clear to me what's going on" is a tell. If it were actually clear to Goofusia what was going on with the pamphlet author expressing anger and hatred towards Rev. Parris, she would not use the word "somehow" in describing the author's behavior: she would be able to pass the author's ideological Turing test and therefore know exactly how.

If that were just Goofusia's mistake, the loss would be hers alone, but if Goofusia is in a position of social power over others, she might succeed at spreading her anti-speech, anti-reading cultural practices to others. I can only imagine that the result would be a subculture that was obsessively self-congratulatory about its own superiority in "truthseeking", while simultaneously blind to everything outside itself. People spending their lives immersed in that culture wouldn't necessarily notice anything was wrong from the inside. What could you say to help them?

An Analogy to Reinforcement Learning From Human Feedback

Pointing out problems is easy. Finding solutions is harder.

The training pipeline for frontier AI systems typically includes a final step called reinforcement learning from human feedback (RLHF). After training a "base" language model that predicts continuations of internet text, supervised fine-tuning is used to make the model respond in the form of an assistant answering user questions, but making the assistant responses good is more work. It would be expensive to hire a team of writers to manually compose the thousands of user-question–assistant-response examples needed to teach the model to be a good assistant. The solution is RLHF: a reward model (often just the same language model with a different final layer) is trained to predict the judgments of human raters about which of a pair of model-generated assistant responses is better, and the model is optimized against the reward model.

The problem with the solution is that human feedback (and the reward model's prediction of it) is imperfect. The reward model can't tell the difference between "The AI is being good" and "The AI looks good to the reward model". This already has the failure mode of sycophancy, in which today's language model assistants tell users what they want to hear, but theory and preliminary experiments suggest that much larger harms (up to and including human extinction) could materialize from future AI systems deliberately deceiving their overseers—not because they suddenly "woke up" and defied their training, but because what we think we trained them to do (be helpful, honest, and harmless) isn't what we actually trained them to do (perform whatever computations were the antecedents of reward on the training distribution).

The problem doesn't have any simple, obvious solution. In the absence of some sort of international treaty to halt all AI development worldwide, "Just don't do RLHF" isn't feasible and doesn't even make any sense; you need some sort of feedback in order to make an AI that does anything useful at all.

The problem may or may not ultimately be solvable with some sort of complicated, nonobvious solution that tries to improve on naïve RLHF. Researchers are hard at work studying alternatives involving red-teaming, debate, interpretability, mechanistic anomaly detection, and more.

But the first step on the road to some future complicated solution to the problem of naïve RLHF, is acknowledging that the the problem is at least potentially real, and having some respect that the problem might be difficult, rather than just eyeballing the results of RLHF and saying that it looks great.

If a safety auditor comes to the CEO of an AI company expressing concerns about the company's RLHF pipeline being unsafe due to imperfect rater feedback, it's more reassuring if the CEO says, "Yes, we thought of that, too; we've implemented these-and-such mitigations and are monitoring such-and-these signals which we hope will clue us in if the mitigations start to fail."

If the CEO instead says, "Well, I think our raters are great. Are you insulting our raters?", that does not inspire confidence. The natural inference is that the CEO is mostly interested in this quarter's profits and doesn't really care about safety.

Similarly, the problem with selection effects on approved information, in which your salon can't tell the difference between "Our ideas are good" and "Our ideas look good to us," doesn't have any simple, obvious solution. "Just don't filter information" isn't feasible and doesn't even make any sense; you need some sort of filter because it's not physically possible to read everything and respond to everything.

The problem may or may not ultimately be solvable with some complicated solution involving prediction markets, adversarial collaborations, anonymous criticism channels, or any number of other mitigations I haven't thought of, but the first step on the road to some future complicated solution is acknowledging that the problem is at least potentially real, and having some respect that the problem might be difficult. If alarmed members come to the organizers of the salon with concerns about collective belief distortions due to suppression of information and the organizers meet them with silence, "bowing out", or defensive blustering, rather than "Yes, we thought of that, too," that does not inspire confidence. The natural inference is that the organizers are mostly interested in maintaining the salon's prestige and don't really care about the truth.

Disagreement Comes From the Dark World

In "Truth or Dare", Duncan Sabien articulates a phenomenon in which expectations of good or bad behavior can become self-fulfilling: people who expect to be exploited and feel the need to put up defenses both elicit and get sorted into a Dark World where exploitation is likely and defenses are necessary, whereas people who expect beneficence tend to attract beneficence in turn.

Among many other examples, Sabien highlights the phenomenon of gift economies: a high-trust culture in which everyone is eager to help each other out whenever they can is a nicer place to live than a low-trust culture in which every transaction must be carefully tracked for fear of enabling free-riders.

I'm skeptical of the extent to which differences between high- and low-trust cultures can be explained by self-fulfilling prophecies as opposed to pre-existing differences in trust-worthiness, but I do grant that self-fulfilling expectations can sometimes play a role: if I insist on always being paid back immediately and in full, it makes sense that that would impede the development of gift-economy culture among my immediate contacts. So far, the theory articulated in the essay seems broadly plausible.

Later, however, the post takes an unexpected turn:

Treating all of the essay thus far as prerequisite and context:

This is why you should not trust Zack Davis, when he tries to tell you what constitutes good conduct and productive discourse. Zack Davis does not understand how high-trust, high-cooperation dynamics work. He has never seen them. They are utterly outside of his experience and beyond his comprehension. What he knows how to do is keep his footing in a world of liars and thieves and pickpockets, and he does this with genuinely admirable skill and inexhaustible tenacity.

But (as far as I can tell, from many interactions across years) Zack Davis does not understand how advocating for and deploying those survival tactics (which are 100% appropriate for use in an adversarial memetic environment) utterly destroys the possibility of building something Better. Even if he wanted to hit the "cooperate" button—

(In contrast to his usual stance, which from my perspective is something like "look, if we all hit 'defect' together, in full foreknowledge, then we don't have to extend trust in any direction and there's no possibility of any unpleasant surprises and you can all stop grumping at me for repeatedly 'defecting' because we'll all be cooperating on the meta level, it's not like I didn't warn you which button I was planning on pressing, I am in fact very consistent and conscientious.")

—I don't think he knows where it is, or how to press it.

(Here I'm talking about the literal actual Zack Davis, but I’m also using him as a stand-in for all the dark world denizens whose well-meaning advice fails to take into account the possibility of light.)

As a reader of the essay, I reply: wait, who? Am I supposed to know who this Davies person is? Ctrl-F search confirms that they weren't mentioned earlier in the piece; there's no reason for me to have any context for whatever this section is about.

As Zack Davis, however, I have a more specific reply, which is: yeah, I don't think that button does what you think it does. Let me explain.


In figuring out what would constitute good conduct and productive discourse, it's important to appreciate how bizarre the human practice of "discourse" looks in light of Aumann's dangerous idea.

There's only one reality. If I'm a Bayesian reasoner honestly reporting my beliefs about some question, and you're also a Bayesian reasoner honestly reporting your beliefs about the same question, we should converge on the same answer, not because we're cooperating with each other, but because it is the answer. When I update my beliefs based on your report on your beliefs, it's strictly because I expect your report to be evidentially entangled with the answer. Maybe that's a kind of "trust", but if so, it's in the same sense in which I "trust" that an increase in atmospheric pressure will exert force on the exposed basin of a classical barometer and push more mercury up the reading tube. It's not personal and it's not reciprocal: the barometer and I aren't doing each other any favors. What would that even mean?

In contrast, my friends and I in a gift economy are doing each other favors. That kind of setting featuring agents with a mixture of shared and conflicting interests is the context in which the concepts of "cooperation" and "defection" and reciprocal "trust" (in the sense of people trusting each other, rather than a Bayesian robot trusting a barometer) make sense. If everyone pitches in with chores when they can, we all get the benefits of the chores being done—that's cooperation. If you never wash the dishes, you're getting the benefits of a clean kitchen without paying the costs—that's defection. If I retaliate by refusing to wash any dishes myself, then we both suffer a dirty kitchen, but at least I'm not being exploited—that's mutual defection. If we institute a chore wheel with an auditing regime, that reëstablishes cooperation, but we're paying higher transaction costs for our lack of trust. And so on: Sabien's essay does a good job of explaining how there can be more than one possible equilibrium in this kind of system, some of which are much more pleasant than others.

If you've seen high-trust gift-economy-like cultures working well and low-trust backstabby cultures working poorly, it might be tempting to generalize from the domains of interpersonal or economic relationships, to rational (or even "rationalist") discourse. If trust and cooperation are essential for living and working together, shouldn't the same lessons apply straightforwardly to finding out what's true together?

Actually, no. The issue is that the payoff matrices are different.

Life and work involve a mixture of shared and conflicting interests. The existence of some conflicting interests is an essential part of what it means for you and me to be two different agents rather than interchangable parts of the same hivemind: we should hope to do well together, but when push comes to shove, I care more about me doing well than you doing well. The art of cooperation is about maintaining the conditions such that push does not in fact come to shove.

But correct epistemology does not involve conflicting interests. There's only one reality. Bayesian reasoners cannot agree to disagree. Accordingly, when humans successfully approach the Bayesian ideal, it doesn't particularly feel like cooperating with your beloved friends, who see you with all your blemishes and imperfections but would never let a mere disagreement interfere with loving you. It usually feels like just perceiving things—resolving disagreements so quickly that you don't even notice them as disagreements.

Suppose you and I have just arrived at a bus stop. The bus arrives every half-hour. I don't know when the last bus was, so I don't know when the next bus will be: I assign a uniform probability distribution over the next thirty minutes. You recently looked at the transit authority's published schedule, which says the bus will come in six minutes: most of your probability-mass is concentrated tightly around six minutes from now.

We might not consciously notice this as a "disagreement", but it is: you and I have different beliefs about when the next bus will arrive; our probability distributions aren't the same. It's also very ephemeral: when I ask, "When do you think the bus will come?" and you say, "six minutes; I just checked the schedule", I immediately replace my belief with yours, because I think the published schedule is probably right and there's no particular reason for you to lie about what it says.

Alternatively, suppose that we both checked different versions of the schedule, which disagree: the schedule I looked at said the next bus is in twenty minutes, not six. When we discover the discrepancy, we infer that one of the schedules must have been outdated, and both adopt a distribution with most of the probability-mass in separate clumps around six and twenty minutes from now. Our initial beliefs can't both have been right—but there's no reason for me to weight my prior belief more heavily just because it was mine.

At worst, approximating ideal belief exchange feels like working on math. Suppose you and I are studying the theory of functions of a complex variable. We're trying to prove or disprove the proposition that if an entire function satisfies \(f(x + 1) = f(x)\) for real \(x\), then \(f(z + 1) = f(z)\) for all complex \(z\). I suspect the proposition is false and set about trying to construct a counterexample; you suspect the proposition is true and set about trying to write a proof by contradiction. Our different approaches do seem to imply different probabilistic beliefs about the proposition, but I can't be confident in my strategy just because it's mine, and we expect the disagreement to be transient: as soon as I find my counterexample or you find your reductio, we should be able to share our work and converge.


Most real-world disagreements of interest don't look like the bus arrival or math problem examples—qualitatively, not as a matter of trying to prove quantitatively harder theorems. Real-world disagreements tend to persist; they're predictable—in flagrant contradiction of how the beliefs of Bayesian reasoners would follow a random walk. From this we can infer that typical human disagreements aren't "honest", in the sense that at least one of the participants is behaving as if they have some other goal than getting to the truth.

Importantly, this characterization of dishonesty is using a functionalist criterion: when I say that people are behaving as if they have some other goal than getting to the truth, that need not imply that anyone is consciously lying; "mere" bias is sufficient to carry the argument.

Dishonest disagreements end up looking like conflicts because they are disguised conflicts. The parties to a dishonest disagreement are competing to get their preferred belief accepted, where beliefs are being preferred for some reason other than their accuracy: for example, because acceptance of the belief would imply actions that would benefit the belief-holder. If it were true that my company is the best, it would follow logically that customers should buy my products and investors should fund me. And yet a discussion with me about whether or not my company is the best probably doesn't feel like a discussion about bus arrival times or the theory of functions of a complex variable. You probably expect me to behave as if I thought my belief is better "because it's mine", to treat attacks on the belief as if they were attacks on my person: a conflict rather than a disagreement.

"My company is the best" is a particularly stark example of a typically dishonest belief, but the pattern is very general: when people are attached to their beliefs for whatever reason—which is true for most of the beliefs that people spend time disagreeing about, as contrasted to math and bus-schedule disagreements that resolve quickly—neither party is being rational (which doesn't mean neither party is right on the object level). Attempts to improve the situation should take into account that the typical case is not that of truthseekers who can do better at their shared goal if they learn to trust each other, but rather of people who don't trust each other because each correctly perceives that the other is not truthseeking.

Again, "not truthseeking" here is meant in a functionalist sense. It doesn't matter if both parties subjectively think of themselves as honest. The "distrust" that prevents Aumann-agreement-like convergence is about how agents respond to evidence, not about subjective feelings. It applies as much to a mislabeled barometer as it does to a human with a functionally-dishonest belief. If I don't think the barometer readings correspond to the true atmospheric pressure, I might still update on evidence from the barometer in some way if I have a guess about how its labels correspond to reality, but I'm still going to disagree with its reading according to the false labels.


There are techniques for resolving economic or interpersonal conflicts that involve both parties adopting a more cooperative approach, each being more willing to do what the other party wants (while the other reciprocates by doing more of what the first one wants). Someone who had experience resolving interpersonal conflicts using techniques to improve cooperation might be tempted to apply the same toolkit to resolving dishonest disagreements.

It might very well work for resolving the disagreement. It probably doesn't work for resolving the disagreement correctly, because cooperation is about finding a compromise amongst agents with partially conflicting interests, and in a dishonest disagreement in which both parties have non-epistemic goals, trying to do more of what the other party functionally "wants" amounts to catering to their bias, not systematically getting closer to the truth.

Cooperative approaches are particularly dangerous insofar as they seem likely to produce a convincing but false illusion of rationality, despite the participants' best of subjective conscious intentions. It's common for discussions to involve more than one point of disagreement. An apparently productive discussion might end with me saying, "Okay, I see you have a point about X, but I was still right about Y."

This is a success if the reason I'm saying that is downstream of you in fact having a point about X but me in fact having been right about Y. But another state of affairs that would result in me saying that sentence, is that we were functionally playing a social game in which I implicitly agreed to concede on X (which you visibly care about) in exchange for you ceding ground on Y (which I visibly care about).

Let's sketch out a toy model to make this more concrete. "Truth or Dare" uses color perception an illustration of confirmation bias: if you've been primed to make the color yellow salient, it's easy to perceive an image as being yellower than it is.

Suppose Jade and Ruby consciously identify as truthseekers, but really, Jade is biased to perceive non-green things as green 20% of the time, and Ruby is biased to perceive non-red things as red 20% of the time. In our functionalist sense, we can model Jade as "wanting" to misrepresent the world as being greener than it is, and Ruby as "wanting" to misrepresent the world is being redder than it is.

Confronted with a sequence of gray objects, Jade and Ruby get into a heated argument: Jade thinks 20% of the objects are green and 0% are red, whereas Ruby thinks they're 0% green and 20% red.

As tensions flare, someone who didn't understand the deep disanalogy between human relations and epistemology might propose that Jade and Ruby should strive be more "cooperative", establish higher "trust."

What does that mean? Honestly, I'm not entirely sure, but I worry that if someone takes high-trust gift-economy-like cultures as their inspiration and model for how to approach intellectual disputes, they'll end up giving bad advice in practice.

Cooperative human relationships result in everyone getting more of what they want. If Jade wants to believe that the world is greener than it is and Ruby wants to believe that the world is redder than it is, then naïve attempts at "cooperation" might involve Jade making an effort to see things Ruby's way at Ruby's behest, and vice versa. But Ruby is only going to insist that Jade make an effort to see it her way when Jade says an item isn't red. (That's what Ruby cares about.) Jade is only going to insist that Ruby make an effort to see it her way when Ruby says an item isn't green. (That's what Jade cares about.)

If the two (perversely) succeed at seeing things the other's way, they would end up converging on believing that the sequence of objects is 20% green and 20% red (rather than the 0% green and 0% red that it actually is). They'd be happier, but they would also be wrong. In order for the pair to get the correct answer, then without loss of generality, when Ruby says an object is red, Jade needs to stand her ground: "No, it's not red; no, I don't trust you and won't see things your way; let's break out the Pantone swatches." But that doesn't seem very "cooperative" or "trusting".


At this point, a proponent of the high-trust, high-cooperation dynamics that Sabien champions is likely to object that the absurd "20% green, 20% red" mutual-sycophancy outcome in this toy model is clearly not what they meant. (As Sabien takes pains to clarify in "Basics of Rationalist Discourse", "If two people disagree, it's tempting for them to attempt to converge with each other, but in fact the right move is for both of them to try to see more of what's true.")

Obviously, the mutual sycophancy outcome is clearly not what proponents of trust and cooperation consciously intend. The problem is that mutual sycophancy seems to be the natural outcome of treating interpersonal conflicts as analogous to epistemic disagreements and trying to resolve them both using cooperative practices, when in fact the decision-theoretic structure of those situations are very different. The text of "Truth or Dare" seems to treat the analogy as a strong one; it wouldn't make sense to spend so many thousands of words discussing gift economies and the eponymous party game and then draw a conclusion about "what constitutes good conduct and productive discourse", if gift economies and the party game weren't relevant to what constitutes productive discourse.

"Truth or Dare" seems to suggest that it's possible to escape the Dark World by excluding the bad guys. "[F]rom the perspective of someone with light world privilege, [...] it did not occur to me that you might be hanging around someone with ill intent at all," Sabien imagines a denizen of the light world saying. "Can you, um. Leave? Send them away? Not be spending time in the vicinity of known or suspected malefactors?"

If we're talking about holding my associates to a standard of ideal truthseeking (as contrasted to a lower standard of "not using this truth-or-dare game to blackmail me"), then, no, I think I'm stuck spending time in the vicinity of people who are known or suspected to be biased. I can try to mitigate the problem by choosing less biased friends, but when we do disagree, I have no choice but to approach that using the same rules of reasoning that I would use with a possibly-mislabeled barometer, which do not have a particularly cooperative character. Telling us that the right move is for both of us to try to see more of what's true is tautologically correct but non-actionable; I don't know how to do that except by my usual methodology, which Sabien has criticized as characteristic of living in a dark world.

That is to say: I do not understand how high-trust, high-cooperation dynamics work. I've never seen them. They are utterly outside my experience and beyond my comprehension. What I do know is how to keep my footing in a world of people with different goals from me, which I try to do with what skill and tenacity I can manage.

And if someone should say that I should not be trusted when I try to explain what constitutes good conduct and productive discourse ... well, I agree!

I don't want people to trust me, because I think trust would result in us getting the wrong answer.

I want people to read the words I write, think it through for themselves, and let me know in the comments if I got something wrong.

"Yes, and—" Requires the Possibility of "No, Because—"

Scott Garrabrant gives a number of examples to illustrate that "Yes Requires the Possibility of No". We can understand the principle in terms of information theory. Consider the answer to a yes-or-no question as a binary random variable. The "amount of information" associated with a random variable is quantified by the entropy, the expected value of the negative logarithm of the probability of the outcome. If we know in advance of asking that the answer to the question will always be Yes, then the entropy is −P(Yes)·log(P(Yes)) − P(No)·log(P(No)) = −1·log(1) − 0·log(0) = 0.1 If you already knew what the answer would be, then the answer contains no information; you didn't learn anything new by asking.


In the art of improvisational theater ("improv" for short), actors perform scenes that they make up as they go along. Without a script, each actor's choices of what to say and do amount to implied assertions about the fictional reality being portrayed, which have implications for how the other actors should behave. A choice that establishes facts or gives direction to the scene is called an offer. If an actor opens a scene by asking their partner, "Is it serious, Doc?", that's an offer that the first actor is playing a patient awaiting diagnosis, and the second actor is playing a doctor.

A key principle of improv is often known as "Yes, and" after an exercise that involves starting replies with those words verbatim, but the principle is broader and doesn't depend on the particular words used: actors should "accept" offers ("Yes"), and respond with their own complementary offers ("and"). The practice of "Yes, and" is important for maintaining momentum while building out the reality of the scene.

Rejecting an offer is called blocking, and is frowned upon. If one actor opens the scene with, "Surrender, Agent Stone, or I'll shoot these hostages!"—establishing a scene in which they're playing an armed villain being confronted by an Agent Stone—it wouldn't do for their partner to block by replying, "That's not my name, you don't have a gun, and there are no hostages." That would halt the momentum and confuse the audience. Better for the second actor to say, "Go ahead and shoot, Dr. Skull! You'll find that my double agent on your team has stolen your bullets"—accepting the premise ("Yes"), then adding new elements to the scene ("and", the villain's name and the double agent).

Notice a subtlety: the Agent Stone character isn't "Yes, and"-ing the Dr. Skull character's demand to surrender. Rather, the second actor is "Yes, and"-ing the first actor's worldbuilding offers (where the offer happens to involve their characters being in conflict). Novice improvisers are sometimes tempted to block to try to control the scene when they don't like their partner's offers, but it's almost always a mistake. Persistently blocking your partner's offers kills the vibe, and with it, the scene. No one wants to watch two people arguing back-and-forth about what reality is.


Proponents of collaborative truthseeking think that many discussions benefit from a more "open" or "interpretive" mode in which participants prioritize constructive contributions that build on each other's work rather than tearing each other down.

The analogy to improv's "Yes, and" doctrine writes itself, right down to the subtlety that collaborative truthseeking does not discourage disagreement as such—any more than the characters in an improv sketch aren't allowed to be in conflict. What's discouraged is the persistent blocking of offers, refusing to cooperate with the "scene" of discourse your partner is trying to build. Partial disagreement with polite elaboration ("I see what you're getting at, but have you considered ...") is typically part of the offer—that we're "playing" reasonable people having a cooperative intellectual discussion. Only wholesale negation ("That's not a thing") is blocking—by rejecting the offer that we're both playing reasonable people.

Whatever you might privately think of your interlocutor's contribution, it's not hard to respond in a constructive manner without lying. Like a good improv actor, you can accept their contribution to the scene/discourse ("Yes"), then add your own contribution ("and"). If nothing else, you can write about how their comment reminded you of something else you've read, and your thoughts about that.

Reading over a discussion conducted under such norms, it's easy to not see a problem. People are building on each other's contributions; information is being exchanged. That's good, right?

The problem is that while the individual comments might (or might not) make sense when read individually, the harmonious social exchange of mutually building on each other's contributions isn't really a conversation unless the replies connect to each other in a less superficial way that risks blocking.

What happens when someone says something wrong or confusing or unclear? If their interlocutor prioritizes correctness and clarity, the natural behavior is to say, "No, that's wrong, because ..." or "No, I didn't understand that"—and not only that, but to maintain that "No" until clarity is forthcoming. That's blocking. It feels much more cooperative to let it pass in order to keep the scene going—with the result that falsehood, confusion, and unclarity accumulate as the interaction goes on.

There's a reason improv is almost synonymous with improv comedy. Comedy thrives on absurdity: much of the thrill and joy of improv comedy is in appreciating what lengths of cleverness the actors will go to maintain the energy of a scene that has long since lost any semblance of coherence or plausibility. The rules that work for improv comedy don't even work for (non-improvised, dramatic) fiction; it certainly won't work for philosophy.

Per Garrabrant's principle, the only way an author could reliably expect discussion of their work to illuminate what they're trying to communicate is if they knew they were saying something the audence already believed. If you're thinking carefully about what the other person said, you're often going to end up saying "No" or "I don't understand", not just "Yes, and": if you're committed to validating your interlocutor's contribution to the scene before providing your own, you're not really talking to each other.


  1. I'm glossing over a technical subtlety here by assuming—pretending?—that 0·log(0) = 0, when log(0) is actually undefined. But it's the correct thing to pretend, because the linear factor \(p\) goes to zero faster than \(\log p\) can go to negative infinity. Formally: \(\lim_{p \to 0^+} p \log(p) = \lim_{p \to 0^+} \frac{\log(p)}{1/p} = \lim_{p \to 0^+} \frac{1/p}{-1/p^2} = 0\) 

Comment on “Four Layers of Intellectual Conversation”

(originally published at Less Wrong)

One of the most underrated essays in the post-Sequences era of Eliezer Yudkowsky's corpus is "Four Layers of Intellectual Conversation". The degree to which this piece of wisdom has fallen into tragic neglect in these dark ages of the 2020s may be related to its ephemeral form of publication: it was originally posted as a status update on Yudkowsky's Facebook account on 20 December 2016 and subsequently mirrored on Alyssa Vance's The Rationalist Conspiracy blog, which has since gone offline. (The first link in this paragraph is to an archive of the Rationalist Conspiracy post.)

In the post, Yudkowsky argues that a structure of intellectual value necessarily requires four layers of conversation: thesis, critique, response, and counter-response (which Yudkowsky indexes from zero as layers 0, 1, 2, and 3).

The importance of critique is already widespread common wisdom: if a thesis is advanced and promulgated without any serious effort to examine why it might be in error, then it likely is in error, both because it can't have incorporated corrections from critiques (which are ex hypothesi absent) and because the author lacks incentives to offer a correct thesis in the first place: if being right is difficult and there's no social penalty for being wrong, then most humans will inexorably find themselves on the easy course of being wrong even without any conscious intent to deceive. That is, in the words of the post, the problem with "a conversation consisting of people saying X and nobody saying 'hey maybe not-X'" is that "people could say stupid things about X, and nobody would call them on the stupidity." Yudkowsky aptly concludes: "Yikes!"

Yudkowsky's key observation going beyond common wisdom is that the necessity of social incentives to be correct also applies to the level-1 critique and level-2 response, not just the level-0 thesis—and moreover, that the higher levels are critical for the lower levels to maintain their force. The mere existence of level-1 critics won't suffice to keep level-0 thesis-proposers on their toes, if the level-1 critics are themselves not on their toes because they don't anticipate being held to account by level-2 responses. Likewise, level-2 responses won't suffice to keep level-1 critics on their toes if the level-2 responders don't anticipate being held to account by level-3 counter-responses. Without all four levels, the whole structure comes apart.

Yudkowsky offers public debates about evolution and molecular nanotechnology as examples of discourses with a missing level 3. If biologists explain evolution (level-0 thesis), religious scholars insist that God must have started it all (level-1 critique), biologists explain leading theories of abiogenesis (level-2 response), but religious scholars don't engage with the abiogenesis work, then the conversation has failed to secure a level-3 counter-response.

It matters that the higher levels are being held to a high enough standard that people would lose face if they played dumb. If K. Eric Drexler writes technical books and papers about the possibilities of nanotechnology (level-0 thesis), Richard Smalley objects that manipulator arms themselves made of atoms would be too "fat" and "sticky" to work as a molecular assembler and that this problem is fundamentally uncircumventable (level-1 critique), Drexler et al. reply that biological ribosomes demonstrate that the problem is not fundamentally uncircumventable even though Drexler's proposals have a "mechanical" rather than "biological" character (level-2 response), and Smalley objects that biological systems can't work with the materials used in technology and that Drexler has departed from real chemistry (level-3 counter-response), then all four levels are formally present, but one is left with disquieting sense that the level-3 counter-response has failed to truly connect with the level-2 response. (Drexler et al.'s level-2 response had brought up biology as an existence proof that the "fat finger" problem didn't sink the entire idea of nanotechnology; pointing out that biology can't do the things that Drexler had conjectured nanotechnology could, would seem to be missing the point.)

Yudkowsky laments that the academic journal system, with the possible exception of analytic philosophy, mostly only canonizes levels 0–2: it's uncommon to see a journal article that's a reply to a reply to a reply to another. To the extent that real intellectual progress is being made in most fields, the real work is probably happening at conferences or on email lists, with the journals merely recording the work after the fact. Yudkowsky sings the praises of transhumanist mailing lists of the late '90s, where people who might otherwise succumb to the temptation to play dumb were kept in check for fear of Robin Hanson's clinically precise rebuttals. Nick Bostrom's 2014 Superintelligence merely packaged up for the public the outcome of a hard-fought discourse that had occurred elsewhere.


A shortcoming of the original post is a lack of concrete examples (with labeled levels) of the four levels of conversation succeeding rather than failing. (We didn't get much detail about exactly what happened on that mailing list.)

The impact of the replication crisis on the study of priming effects might be a candidate. In 1996's "Automaticity of Social Behavior: Direct Effects of Trait Construct and Stereotype Activation on Action", John A. Bargh and collaborators reported that college students directed to solve a puzzle involving words related to elderly people walked slower when leaving the lab (level-0 thesis). Sixteen years later, in "Behavioral Priming: It's All in the Mind, but Whose Mind?", Stéphane Doyen et al. ran a replication that failed to reproduce the original result on walking speed when the experimenter administering the puzzle was blinded to the hypothesis being tested, but did reproduce the result when the experimenter was led to believe that there would be a priming effect (level-1 critique). Bargh wrote a blog post, "Nothing in Their Heads", arguing that the experimenter was blinded in the original 1996 study, that Doyen et al. over-primed with too many elderliness-related words (which Bargh argued could destroy the effect), and that Doyen et al. didn't check if subjects had slowness-related stereotypes about the elderly (level-2 response). Though the original post's comment section seems to have been lost to history, science journalist Ed Yong documented responses to Bargh's post by commenters on the post and by coauthors of Doyen et al., claiming inaccuracies in the post, and that, in any case, a truly robust priming effect wouldn't be so fragile to such small changes in the study design (level-3 counter-response).

Nor did the conversation about this particular paper drop silently into the void: soon, the famed Daniel Kahneman would write a letter to priming research practitioners named to him by Bargh on bringing more rigorous study designs to the field, which has continued to be plagued by replication difficulties. The discussion made an impact on Society's collective beliefs. The attempt at discourse was more than a noble gesture. It hadn't all been for nothing.


A natural question to ask about the four-levels framework is: why four levels, specifically? Doesn't the recursion of level n needing level n + 1 go off to infinity?

The original post leaves the question unanswered, but a potential answer can be found in Yudkowsky's tongue-in-cheek Law of Ultrafinite Recursion, which states that, in practice, infinite recursions are at most three levels deep. The Law of Ultrafinite Recursion is deliberately silly if construed as a literal claim about computer science but is surprisingly fruitful as a claim about human psychology: it's pretty natural to ask what Alice thinks that Bob thinks about Carol, but asking what Alice thinks that Bob thinks that Carol thinks about Dave feels like a stretch.

If the limited human grasp of recursion rounds "four" up to "infinity", then the chain of thesis–critique–response–counter-response is enough to establish the expectation of unlimited-depth accountability and remove the incentive to bluff. A different species with greater working memory capacity, whose members could follow a backwards induction farther, might need more counter-counter-responses and counter-counter-counter-responses to experience the same salutary effect.


The four-levels model is about robust disagreements, which are usually pretty frustrating for all involved. No one likes being told they're wrong, especially by people who (so it always seems from the other side) are themselves obviously wrong.

The frustration is not optional. The recursive pressure forcing you to come up with your best arguments and responses to counter the adversary's critiques and counter-responses only works if the adversary is allowed to be frustrating; it's not their job to make it easy for you. Equivalently, it's not your job to make it easy for them. Only by facing this test can your combined efforts build an intellectual edifice guided by the beauty of your weapons.

Bargh's blog post complains that "oddly for an article that purported to fail to replicate one of [his] past studies", he wasn't asked to review Doyen et al. But it's not odd: journals generally want reviewers to be independent. For example, the International Committee of Medical Journal Editors recommends that peer reviewers should "declare their relationships and activities that might bias their evaluation of a manuscript and recuse themselves from the peer-review process if a conflict exists."

If Bargh were the one who got to decide who is allowed to speak on the record about potential flaws in Bargh et al. 1996, then Society would lose out on its chance to determine whether Bargh et al. 1996 is actually correct. Any single conversational locus that forgets or denies this obvious principle is at serious risk of degenerating into an echo chamber if it hasn't already.

On the Contrary, Steelmanning Is Normal; ITT-Passing Is Niche

(originally published at Less Wrong)

Rob Bensinger argues that "ITT-passing and civility are good; 'charity' is bad; steelmanning is niche".

The ITT—Ideological Turing Test—is an exercise in which one attempts to present one's interlocutor's views as persuasively as the interlocutor themselves can, coined by Bryan Caplan in analogy to the Turing Test for distinguishing between humans and intelligent machines. (An AI that can pass as human must presumably possess human-like understanding; an opponent of an idea that can pass as an advocate for it presumably must possess an advocate's understanding.) "Steelmanning" refers to the practice of addressing a stronger version of an interlocutor's argument, coined in disanalogy to "strawmanning", the crime of addressing a weaker version of an interlocutor's argument in the hopes of fooling an audience (or oneself) that the original argument has been rebutted.

Bensinger describes steelmanning as "a useful niche skill", but thinks it isn't "a standard thing you bring out in most arguments." Instead, he writes, discussions should be structured around object-level learning, trying to pass each other's Ideological Turing Test, or trying resolve cruxes.

I think Bensinger has it backwards: the Ideological Turing Test is a useful niche skill, but it doesn't belong on a list of things to organize a discussion around, whereas something like steelmanning naturally falls out of object-level learning. Let me explain.

The ITT is a test of your ability to model someone else's models of some real-world phenomena of interest. But usually, I'm much more interested in modeling the real-world phenomena of interest directly, rather than modeling someone else's models of it.

I couldn't pass an ITT for advocates of Islam or extrasensory perception. On the one hand, this does represent a distinct deficit in my ability to model what the advocates of these ideas are thinking, a tragic gap in my comprehension of reality, which I would hope to remedy in the Glorious Transhumanist Future if that were a real thing. On the other hand, facing the constraints of our world, my inability to pass an ITT for Islam or ESP seems ... basically fine? I already have strong reasons to doubt the existence of ontologically fundamental mental entities. I accept my ignorance of the reasons someone might postulate otherwise, not out of contempt, but because I just don't have the time.

Or think of it this way: as a selfish seeker of truth speaking to another selfish seeker of truth, when would I want to try to pass my interlocutor's ITT, or want my interlocutor to try to pass my ITT?

In the "outbound" direction, I'm not particularly selfishly interested in passing my interlocutor's ITT because, again, I usually don't care much about other people's beliefs, as contrasted to the reality that those beliefs are reputedly supposed to track. I listen to my interlocutor hoping to learn from them, but if some part of what they say seems hopelessly wrong, it doesn't seem profitable to pretend that it isn't until I can reproduce the hopeless wrongness in my own words.

Crucially, the same is true in the "inbound" direction. I don't expect people to be able to pass my ITT before criticizing my ideas. That would make it harder for people to inform me about flaws in my ideas!

But if I'm not particularly interested in passing my interlocutor's ITT or in my interlocutor passing mine, and my interlocutor presumably (by symmetry) feels the same way, why would we bother?

All this having been said, I absolutely agree that, all else being equal, the ability to pass ITTs is desirable. It's useful as a check that you and your interlocutor are successfully communicating, rather than talking past each other. If I couldn't do better on an ITT for Islam or ESP after debating a proponent, that would be alarming—it's just that I'd want to try the old-fashioned debate algorithm first, and improve my ITT score as a side-effect, rather than trying to optimize my ITT score directly.

There are occasions when I'm inclined to ask an interlocutor to pass my ITT—specifically when I suspect them of not being honest about their motives, of being selfish about something other than the pursuit of truth (like winning acclaim for "their own" current theories). If someone seems persistently motivated to strawman you, asking them to just repeat back what you said in their own words is a useful device to get the discussion back on track. (Or to end it, if they clearly don't even want to try.)

In contrast to the ITT, steelmanning is something a selfish seeker of truth is inclined to do naturally, as a consequence of the obvious selfish practice of improving arguments wherever they happen to be found. In the outbound direction, if someone makes a flawed criticism of my ideas, of course I want to fix the flaws and address the improved argument. If the original criticism is faulty, but the repaired criticism exposes a key weakness in my existing ideas, then I learn something, which is great. If I were to just rebut the original criticism without trying to repair it, then I wouldn't learn anything, which would be terrible.

Likewise, in the inbound direction, if my interlocutor notices a flaw in my criticism of their ideas and fixes the flaw before addressing the repaired criticism, that's great. Why would I object?

The motivation here may be clearer if we consider the process of constructing computer programs rather than constructing arguments. When a colleague or language model assistant suggests an improvement to my code, I often accept the suggestion with my own ("steelmanned"?) changes rather than verbatim. This is so commonplace among programmers that it doesn't even have a special name.

Bensinger quotes Eliezer Yudkowsky writing, "If you want to try to make a genuine effort to think up better arguments yourself because they might exist, don't drag the other person into it," but this bizarrely seems to discount the possibility of iterating on criticisms as they are posed. Despite making a genuine effort to think up better code that might exist, I often fail. If other people can see flaws in my code (because they know things I don't) and have their own suggestions, and I can see flaws in their suggestions (because I also know things they don't which didn't make it into my first draft) and have my own counter-suggestions, that seems like an ideal working relationship, not a malign imposition.

All this having been said, I agree that there's a serious potential failure mode where someone who thinks of themselves as steelmanning is actually constructing worse arguments than those that they purport to be improving. In this case, indeed, prompting such a delusional interlocutor to try the ITT first is a crucial remedy.

But crucial remedies are still niche in the sense that they shouldn't be "a standard thing you bring out in most arguments"—or if they are, it's a sign that you need to find better interlocutors. Having to explicitly drag out the ITT is a sign of sickness, not a sign of health. It shouldn't be normal to have to resort to roleplaying exercises to achieve the benefits that could as well be had from basic reading comprehension and a selfish interest in accurate shared maps.

Steven Kaas wrote in 2008:

If you're interested in being on the right side of disputes, you will refute your opponents' arguments. But if you’re interested in producing truth, you will fix your opponents' arguments for them.

To win, you must fight not only the creature you encounter; you must fight the most horrible thing that can be constructed from its corpse.

The ITT is a useful tool for being on the right side of disputes: in order to knowably refute your opponents' arguments, you should be able to demonstrate that you know what those arguments are. I am nevertheless left with a sense that more is possible.

Assume Bad Faith

(originally published at Less Wrong)

I've been trying to avoid the terms "good faith" and "bad faith". I'm suspicious that most people who have picked up the phrase "bad faith" from hearing it used, don't actually know what it means—and maybe, that the thing it does mean doesn't carve reality at the joints.

People get very touchy about bad faith accusations: they think that you should assume good faith, but that if you've determined someone is in bad faith, you shouldn't even be talking to them, that you need to exile them.

What does "bad faith" mean, though? It doesn't mean "with ill intent." Following Wikipedia, bad faith is "a sustained form of deception which consists of entertaining or pretending to entertain one set of feelings while acting as if influenced by another." The great encyclopedia goes on to provide examples: the solider who waves a flag of surrender but then fires when the enemy comes out of their trenches, the attorney who prosecutes a case she knows to be false, the representative of a company facing a labor dispute who comes to the negotiating table with no intent of compromising.

That is, bad faith is when someone's apparent reasons for doing something aren't the same as the real reasons. This is distinct from malign intent. The uniformed solider who shoots you without pretending to surrender is acting in good faith, because what you see is what you get: the man whose clothes indicate that his job is to try to kill you is, in fact, trying to kill you.

The policy of assuming good faith (and mercilessly punishing rare cases of bad faith when detected) would make sense if you lived in an honest world where what you see generally is what you get (and you wanted to keep it that way), a world where the possibility of hidden motives in everyday life wasn't a significant consideration.

On the contrary, however, I think hidden motives in everyday life are ubiquitous. As evolved creatures, we're designed to believe as it benefited our ancestors to believe. As social animals in particular, the most beneficial belief isn't always the true one, because tricking your conspecifics into adopting a map that implies that they should benefit you is sometimes more valuable than possessing the map that reflects the territory, and the most persuasive lie is the one you believe yourself. The universal human default is to come up with reasons to persuade the other party why it's in their interests to do what you want—but admitting that you're doing that isn't part of the game. A world where people were straightforwardly trying to inform each other would look shocking and alien to us.

But if that's the case (and you shouldn't take my word for it), being touchy about bad faith accusations seems counterproductive. If it's common for people's stated reasons to not be the same as the real reasons, it shouldn't be beyond the pale to think that of some particular person, nor should it necessarily entail cutting the "bad faith actor" out of public life—if only because, applied consistently, there would be no one left. Why would you trust anyone so highly as to think they never have a hidden agenda? Why would you trust yourself?

The conviction that "bad faith" is unusual contributes to a warped view of the world in which conditions of information warfare are rationalized as an inevitable background fact of existence. In particular, people seem to believe that persistent good faith disagreements are an ordinary phenomenon—that there's nothing strange or unusual about a supposed state of affairs in which I'm an honest seeker of truth, and you're an honest seeker of truth, and yet we end up persistently disagreeing on some question of fact.

I claim that this supposedly ordinary state of affairs is deeply weird at best, and probably just fake. Actual "good faith" disagreements—those where both parties are just trying to get the right answer and there are no other hidden motives, no "something else" going on—tend not to persist.

If this claim seems counterintuitive, you may not be considering all the everyday differences in belief that are resolved so quickly and seamlessly that we tend not to notice them as "disagreements".

Suppose you and I have been planning to go to a concert, which I think I remember being on Thursday. I ask you, "Hey, the concert is on Thursday, right?" You say, "No, I just checked the website; it's on Friday."

In this case, I immediately replace my belief with yours. We both just want the right answer to the factual question of when the concert is. With no "something else" going on, there's nothing stopping us from converging in one step: your just having checked the website is a more reliable source than my memory, and neither you nor the website have any reason to lie. Thus, I believe you; end of story.

In cases where the true answer is uncertain, we expect similarly quick convergence in probabilistic beliefs. Suppose you and I are working on some physics problem. Both of us just want the right answer, and neither of us is particularly more skilled than the other. As soon as I learn that you got a different answer than me, my confidence in my own answer immediately plummets: if we're both equally good at math, then each of us is about as likely to have made a mistake. Until we compare calculations and work out which one of us (or both) made a mistake, I think you're about as likely to be right as me, even if I don't know how you got your answer. It wouldn't make sense for me to bet money on my answer being right simply because it's mine.

Most disagreements of note—most disagreements people care about—don't behave like the concert date or physics problem examples: people are very attached to "their own" answers. Sometimes, with extended argument, it's possible to get someone to change their mind or admit that the other party might be right, but with nowhere near the ease of agreeing on (probabilities of) the date of an event or the result of a calculation—from which we can infer that, in most disagreements people care about, there is "something else" going on besides both parties just wanting to get the right answer.

But if there's "something else" going on in typical disagreements that look like a grudge match rather than a quick exchange of information resulting in convergence of probabilities, then the belief that persistent good faith disagreements are common would seem to be in bad faith! (Because if bad faith is "entertaining [...] one set of feelings while acting as if influenced by another", believers in persistent good faith disagreements are entertaining the feeling that both parties to such a disagreement are honest seekers of truth, but acting otherwise insofar as they anticipate seeing a grudge match rather than convergence.)

Some might object that bad faith is about conscious intent to deceive: honest reporting of unconsciously biased beliefs isn't bad faith. I've previously expressed doubt as to how much of what we call lying requires conscious deliberation, but a more fundamental reply is that from the standpoint of modeling information transmission, the difference between bias and deception is uninteresting—usually not relevant to what probability updates should be made.

If an apple is green, and you tell me that it's red, and I believe you, I end up with false beliefs about the apple. It doesn't matter whether you said it was red because you were consciously lying or because you're wearing rose-colored glasses. The input–output function is the same either way: the problem is that the color you report to me doesn't depend on the color of the apple.

If I'm just trying to figure out the relationship between your reports and the state of the world (as contrasted to caring about punishing liars while letting merely biased people off the hook), the main reason to care about the difference between unconscious bias and conscious deception is that the latter puts up much stronger resistance. Someone who is merely biased will often fold when presented with a sufficiently compelling counterargument (or reminded to take off their rose-colored glasses); someone who's consciously lying will keep lying (and telling ancillary lies to cover up the coverup) until you catch them red-handed in front of an audience with power over them.

Given that there's usually "something else" going on in persistent disagreements, how do we go on, if we can't rely on the assumption of good faith? I see two main strategies, each with their own cost–benefit profile.

One strategy is to stick the object level. Arguments can be evaluated on their merits, without addressing what the speaker's angle is in saying it (even if you think there's probably an angle). This delivers most of the benefits of "assume good faith" norms; the main difference I'm proposing is that speakers' intentions be regarded as off-topic rather than presumed to be honest.

The other strategy is full-contact psychoanalysis: in addition to debating the object-level arguments, interlocutors have free reign to question each other's motives. This is difficult to pull off, which is why most people most of the time should stick to the object level. Done well, it looks like a negotiation: in the course of discussion, pseudo-disagreements (where I argue for a belief because it's in my interests for that belief to be on the shared map) are factorized out into real disagreements and bargaining over interests so that Pareto improvements can be located and taken, rather than both parties fighting to distort the shared map in the service of their interests.

For an example of what a pseudo-disagreement looks like, imagine that I own a factory that I'm considering expanding onto the neighboring wetlands, and you run a local environmental protection group. The regulatory commission with the power to block the factory expansion has a mandate to protect local avian life, but not to preserve wetland area. The factory emits small amounts of Examplene gas. You argue before the regulatory commission that the expansion should be blocked because the latest Science shows that Examplene makes birds sad. I counterargue that the latest–latest Science shows that Examplene actually makes birds happy; the previous studies misheard their laughter as tears and should be retracted.

Realistically, it seems unlikely that our apparent disagreement is "really" about the effects of Examplene on avian mood regulation. More likely, what's actually going on is a conflict rather than a disagreement: I want to expand my factory onto the wetlands, and you want me to not do that. The question of how Examplene pollution affects birds only came into it in order to persuade the regulatory commission.

It's inefficient that our conflict is being disguised as a disagreement. We can't both get what we want, but however the factory expansion question ultimately gets resolved, it would be better to reach that outcome without distorting Society's shared map of the bioactive properties of Examplene. (Maybe it doesn't affect the birds at all!) Whatever the true answer is, Society has a better shot at figuring it out if someone is allowed to point out your bias and mine (because facts about which evidence gets promoted to one's attention are relevant to how one should update on that evidence).

The reason I don't think it's useful to talk about "bad faith" is because the ontology of good vs. bad faith isn't a great fit to either discourse strategy.

If I'm sticking to the object level, it's irrelevant: I reply to what's in the text; my suspicions about the process generating the text are out of scope.

If I'm doing full-contact psychoanalysis, the problem with "I don't think you're here in good faith" is that it's insufficiently specific. Rather than accusing someone of generic "bad faith", the way to move the discussion forward is by positing that one's interlocutor has some specific motive that hasn't yet been made explicit—and the way to defend oneself against such an accusation is by making the case that one's real agenda isn't the one being proposed, rather than protesting one's "good faith" and implausibly claiming not to have an agenda.

The two strategies can be mixed. A simple meta-strategy that performs well without imposing too high of a skill requirement is to default to the object level, and only pull out psychoanalysis as a last resort against stonewalling.

Suppose you point out that my latest reply seems to contradict something I said earlier, and I say, "Look over there, a distraction!"

If you want to continue sticking to the object level, you could say, "I don't understand how the distraction is relevant to resolving the inconsistency in your statements that I raised." On the other hand, if you want to drop down into psychoanalysis, you could say, "I think you're only pointing out the distraction because you don't want to be pinned down." Then I would be forced to either address your complaint, or explain why I had some other reason to point out the distraction.

Crucially, however, the choice of whether to investigate motives doesn't depend on an assumption that only "bad guys" have motives—as if there were bad faith actors who have an angle, and good faith actors who are ideal philosophers of perfect emptiness. There's always an angle; the question is which one.

Lack of Social Grace Is an Epistemic Virtue

(originally published at Less Wrong)

Someone once told me that they thought I acted like refusing to employ the bare minimum of social grace was a virtue, and that this was bad. (I'm paraphrasing; they actually used a different word that starts with b.)

I definitely don't want to say that lack of social grace is unambiguously a virtue. Humans are social animals, so the set of human virtues is almost certainly going to involve doing social things gracefully!

Nevertheless, I will bite the bullet on a weaker claim. Politeness is, to a large extent, about concealing or obfuscating information that someone would prefer not to be revealed—that's why we recognize the difference between one's honest opinion, and what one says when one is "just being polite." Idealized honest Bayesian reasoners would not have social graces—and therefore, humans trying to imitate idealized honest Bayesian reasoners will tend to bump up against (or smash right through) the bare minimum of social grace. In this sense, we might say that the lack of social grace is an "epistemic" virtue—even if it's probably not great for normal humans trying to live normal human lives.

Let me illustrate what I mean with one fictional and one real-life example.


The beginning of the film The Invention of Lying (before the eponymous invention of lying) depicts an alternate world in which everyone is radically honest—not just in the narrow sense of not lying, but more broadly saying exactly what's on their mind, without thought of concealment.

In one scene, our everyman protagonist is on a date at a restaurant with an attractive woman.

"I'm very embarrassed I work here," says the waiter. "And you're very pretty," he tells the woman. "That only makes this worse."

"Your sister?" the waiter then asks our protagonist.

"No," says our everyman.

"Daughter?"

"No."

"She's way out of your league."

"... thank you."

The woman's cell phone rings. She explains that it's her mother, probably calling to check on the date.

"Hello?" she answers the phone—still at the table, with our protagonist hearing every word. "Yes, I'm with him right now. ... No, not very attractive. ... No, doesn't make much money. It's alright, though, seems nice, kind of funny. ... A bit fat. ... Has a funny little—snub nose, kind of like a frog in the—facial ... No, I won't be sleeping with him tonight. ... No, probably not even a kiss. ... Okay, you too, 'bye."

The scene is funny because of how it violates the expected social conventions of our own world. In our world, politeness demands that you not say negative-valence things about someone in front of them, because people don't like hearing negative-valence things about themselves. Someone in our world who behaved like the woman in this scene—calling someone ugly and poor and fat right in front of them—could only be acting out of deliberate cruelty.

But the people in the movie aren't like us. Having taken the call, why should she speak any differently just because the man she was talking about could hear? Why would he object? To a decision-theoretic agent, the value of information is always nonnegative. Given that his date thought he was unattractive, how could it be worse for him to know rather than not-know?

For humans from our world, these questions do have answers—complicated answers having to do with things like map–territory confusions that make receiving bad news seem like a bad event (rather than the good event of learning information about how things were already bad, whether or not you knew it), and how it's advantageous for others to have positive-valence false beliefs about oneself.

The world of The Invention of Lying is simpler, clearer, easier to navigate than our world. There, you don't have to worry whether people don't like you and are planning to harm your interests. They'll tell you.


In "Los Alamos From Below", physicist Richard Feynman's account of his work on the Manhattan Project to build the first atomic bomb, Feynman recalls being sought out by a much more senior physicist specifically for his lack of social graces:

I also met Niels Bohr. His name was Nicholas Baker in those days, and he came to Los Alamos with Jim Baker, his son, whose name is really Aage Bohr. They came from Denmark, and they were very famous physicists, as you know. Even to the big shot guys, Bohr was a great god.

We were at a meeting once, the first time he came, and everybody wanted to see the great Bohr. So there were a lot of people there, and we were discussing the problems of the bomb. I was back in a corner somewhere. He came and went, and all I could see of him was from between people's heads.

In the morning of the day he's due to come next time, I get a telephone call.

"Hello—Feynman?"

"Yes."

"This is Jim Baker." It's his son. "My father and I would like to speak to you."

"Me? I'm Feynman, I'm just a—"

"That's right. Is eight o'clock OK?"

So, at eight o'clock in the morning, before anybody's awake, I go down to the place. We go into an office in the technical area and he says, "We have been thinking how we could make the bomb more efficient and we think of the following idea."

I say, "No, it's not going to work. It's not efficient ... Blah, blah, blah."

So he says, "How about so and so?"

I said, "That sounds a little bit better, but it's got this damn fool idea in it."

This went on for about two hours, going back and forth over lots of ideas, back and forth, arguing. [...]

"Well," [Niels Bohr] said finally, lighting his pipe, "I guess we can call in the big shots now." So then they called all the other guys and had a discussion with them.

Then the son told me what happened. The last time he was there, Bohr said to his son, "Remember the name of that little fellow in the back over there? He's the only guy who's not afraid of me, and will say when I've got a crazy idea. So the next time when we want to discuss ideas, we're not going to be able to do it with these guys who say everything is yes, yes, Dr. Bohr. Get that guy and we'll talk with him first."

I was always dumb in that way. I never knew who I was talking to. I was always worried about the physics. If the idea looked lousy, I said it looked lousy. If it looked good, I said it looked good. Simple proposition.

Someone who felt uncomfortable with Feynman's bluntness and wanted to believe that there's no conflict between rationality and social graces might argue that Feynman's "simple proposition" is actually wrong insofar as it fails to appreciate the map–territory distinction: in saying, "No, it's not going to work", was not Feynman implicitly asserting that just because he couldn't see a way to make it work, it simply couldn't? And in general, shouldn't you know who you're talking to? Wasn't Bohr, the Nobel prize winner, more likely to be right than Feynman, the fresh young Ph.D. (at the time)?

While not entirely without merit (it's true that the map is not the territory; it's true that authority is not without evidential weight), attending overmuch to such nuances distracts from worrying about the physics, which is what Bohr wanted out of Feynman—and, incidentally, what I want out of my readers. I would not expect readers to confirm interpretations with me before publishing a critique. If the post looks lousy, say it looks lousy. If it looks good, say it looks good. Simple proposition.

“Justice, Cherryl.”

(originally published at Less Wrong)

Selfishness and altruism are positively correlated within individuals, for the obvious reason.

@InstanceOfClass

I.

An unfortunate obstacle to appreciating the work of Ayn Rand (as someone who adores the "sense of life" portrayed in Rand's fiction, while having a much lower opinion of her philosophy) is that when Rand praises selfishness and condemns altruism, she's using the words "selfishness" and "altruism" in her own idiosyncratic ideological sense that doesn't match how most people would use those words.

It's true that Rand's heroes are relatively selfish in the sense of being primarily concerned with their own lives, rather than their effects on others. But if you look at what the characters do (rather than the words they say), Rand's villains are also selfish in a conventional sense, using guile and political maneuvering to acquire power and line their own pockets, while claiming to be acting for the common good. For example, in Atlas Shrugged, the various directives ostensibly issued for the economic health of the country are seen to instead benefit politically connected crony capitalists like James Taggart and Orren Boyle. In Think Twice, the philanthropist Walter Breckenridge cultivates a public image as an inventor and benefactor of humanity while stealing credit for his junior partner's work and deriving gratification from exerting power over the people he "helps".

Despite paying lip service to a pretense of only trading and never giving, we also see examples of Rand's heroes being altruistic in the conventional sense, of being motivated to help others. For example, in Atlas Shrugged, Hank Rearden rearranges his production schedule (at a critical time when he could scarcely afford to do so) in order to sell steel to a Mr. Ward, who needs the steel to save his family business (but doesn't see Rearden as obligated to help him). Rearden's motive is pure benevolence: "It's so much for him, thought Rearden, and so little for me!" Giving What We Can couldn't have chosen a better slogan.

Overall, when I look at the universe portrayed in Rand's fiction, it seems to me that the implied moral isn't that altruism is bad.

It's that altruists don't exist. The people claiming to be altruists are lying. The distinguishing feature of our heroes isn't, actually, that they're unusually selfish. It's that they're honest about being mostly selfish, and that they want to pursue their interests within a framework of rights that respects that other people are also trying to pursue their interests. "I swear by my life and my love of it that I will never live for the sake of another man, nor ask another man to live for mine," goes the motto of the striking heroes of Atlas Shrugged (emphasis mine); the second clause is important. Given that everyone is mostly selfish and everyone has to eat, the question is: are you going to eat by means of production and trade, or by—other means?

That's the distinction between Rand's heroes and villains. The heroes want to get rich by means of doing genuinely good work that other people will have a genuine self-interest in paying for. The villains want to wield power by means of psychological manipulation, guilt-tripping and blackmailing the people who can do good work into serving their own parasites and destroyers.

As Greg Hastings, the district attorney in Think Twice, puts it: "[T]he man who admits that he cares for money is all right. He's usually worth the money he makes. He won't kill for it. He doesn't have to. But watch out for the man who yells too loudly how much he scorns money. Watch out particularly for the one who yells that others must scorn it. He's after something much worse than money."

Furthermore, the heroes know that wealth and fame acquired by fraud obviously "don't count." In The Fountainhead, Peter Keating's outwardly successful architecture career has been a sham: he social-engineered his way into partnership in his firm, and all of his best work was plagiarized from the hero, Howard Roark. The turning point for Keating's character is when he asks Roark to let him plagiarize his work one last time, for the Cortlandt housing project, which Roark would never be allowed to work on for political reasons. Keating finally realizes that fraudulent "success" in the eyes of others is no success at all:

"You'll get everything society can give a man. You'll keep all the money. You'll take any fame or honor anyone might want to grant. You'll accept such gratitude as the tenants might feel. And I—I'll take what nobody can give a man, except himself. I will have built Cortlandt." [said Roark.]

"You're getting more than I am, Howard."

In summary, the ultimate sin in Rand's moral universe isn't giving charity. (Because, within the ideology, helping those others whom you want to help, is selfish.) What's evil is demanding charity, claiming the unearned, expecting other people to work for your benefit because you supposedly need them to.

II.

Something people have occasionally noticed about my intellectual style is that I like to win arguments. I take pride and pleasure in pointing out flaws in other people's work in the anticipation of the audience appreciating how clever I am for finding the hole in someone's reasoning.

The people pointing out this fact about me generally seem to think it's a bad thing. They tell me that I should be more charitable to the viewpoints of others, that I ought to be doing collaborative truth-seeking.

It's true, of course, that there's a terrible danger in wanting to win arguments. Once your conclusion has been determined, coming up with more arguments for it can't make you more correct, even if it can help you "win" a debate. Learning something entails changing your mind, which people are often reluctant to do because it amounts to "losing".

A useful heuristic for overcoming this bias against being willing to "lose" arguments is to take heed of a "principle of charity", of taking the strongest and most rational interpretation of others' words. The person you're arguing against is trying to do what they think is right. If you end up disagreeing with them, it shouldn't be because they're stupid and evil; your theory about why the other person is getting the wrong answer shouldn't make them look that bad. If it does, that's a sign that you haven't really understood their point of view and therefore can't claim to have justly refuted it.

From the standpoint of ideal epistemology, however, the "principle of charity" is not a principle, and the idea of "charity" itself is irrelevant or incoherent. Normatively, theories are preferred to the quantitative extent that they are simple and predict the observed data. There is no concept of a theory "belonging to" someone, or favoring someone's interests.

For contingent evolutionary-psychological reasons, humans are innately biased to prefer "their own" ideas, and in that context, a "principle of charity" can be useful as a corrective heuristic—but the corrective heuristic only works by colliding the non-normative bias with a fairness instinct, effectively playing the bias against itself: you wouldn't like it if someone dismissed "your" ideas without understanding why they appeal to you, goes the thought, so you should extend the same consideration to others.

Normatively, of course, this is nonsense. You should update on an interlocutor's arguments for the same reason that a scientist working alone would update on the results of an experiment: because (and to the extent that) the result conveys information about reality. We would not speak of being charitable to an experimental apparatus. The scientist is not doing their lab equipment a favor.

Because the principle of charity is merely a corrective heuristic for the bias of arbitrarily favoring "one's own" ideas, it correspondingly only makes sense to apply in one direction—as a corrective for one's own thoughts. I tell myself to make a special effort to look for reasons why I might be wrong and my interlocutor is right because, knowing what I do about human nature, I selfishly expect to thereby achieve more accurate beliefs than I would in the absence of the special effort. It's a workaround, a mitigation for a known bug in human cognition; it makes sense whether or not the other person reciprocates, and whether or not I'm particularly trying to collaborate with them.

On the other hand, when someone who is currently trying to persuade me of something tells me that it doesn't look I'm making enough effort to think of reasons why they're right, that immediately makes me think they're more likely to be wrong. Why? Because I think that if they had an argument, they would be telling me the argument, not chastising my lack of charity. The advice to be on special lookout for reasons your interlocutor is right is good in general, but your interlocutor is the last person to be trusted to give it, because (due to the warp in human psychology) they have an ulterior motive.

Overall, when I look at the world of discourse I see, the moral I draw is not that that collaborative truth-seeking is bad.

It's that collaborative truth-seeking doesn't exist. The people claiming to be collaborative truth-seekers are lying. Given that everyone wants to be seen as right, the question is: are you going to try to be seen as right by means of providing valid evidence and reasoning, or by—other means?

Or to put it another way: the commenter who admits they care for status is all right. They're usually worth the status they earn. They won't lie for it. They don't have to. But watch out for the commenter who yells too loudly how much they scorn status. Watch out particularly for the one who yells that others must scorn it. They're after something much worse than status.

Furthermore, I know that "winning" a debate via sophistry and rhetorical tricks obviously "doesn't count." Maybe I could fool an undiscriminating audience, but I would know it wasn't real.

Sometimes I want people to understand some specific truth (out of the vast space of possible truths to pay attention to), for selfish reasons of my own. In these cases, I'm happy to do the work of explaining to put it on the shared map. When someone asks me questions about my work, I don't regard it as an attack, because I expect to be able to answer them—and if I can't, that's my problem.

I will never ask my interlocutors to be more charitable to me. I will often say "That's not what I meant", or "That's not a reasonable interpretation of the text I published"—but that's a claim about what I mean, or a claim about the text; it's not a claim on them. I don't expect people to listen to me because I supposedly need them to.

III.

My favorite scene in Atlas Shrugged is the one where Cherryl Taggart (née Brooks) goes to see Dagny Taggart after discovering the truth about her marriage. Cherryl had married Dagny's brother James thinking that he was the intrepid industrialist responsible for the success of the Taggart Transcontinental railroad, only to later find out that James is a phony political actor who took credit for Dagny's accomplishments after the fact, despite having opposed her initiatives and made her work more difficult.

("I married Jim because I ... I thought that he was you," Cherryl tells Dagny. There is some very beautiful slash fanfiction that needs to be written picking up from that line, which is out of scope for this blog post.)

Cherryl intends only to briefly apologize to Dagny for earlier insulting remarks, not to make any further imposition—and is surprised when Dagny not only forgives her, but seems to take a genuine interest in her welfare. It's worth quoting at length:

"You've had a terrible time, haven't you?" [said Dagny.]

"Yes ... but that doesn't matter ... that's my own problem ... and my own fault."

"I don't think it was your own fault."

Cherryl did not answer, then said suddenly, desperately, "Look ... what I don't want is charity."

"Jim must have told you—and it's true—that I never engage in charity."

"Yes, he did ... But what I mean is—"

"I know what you mean."

"But there's no reason why you should have to feel concern for me ... I didn't come here to complain and ... and load another burden on your shoulders. ... That I happen to suffer, doesn't give me a claim on you."

"No, it doesn't. But that you value all the things I value, does."

"You mean ... if you want to talk to me, it's not alms? Not just because you feel sorry for me?"

"I feel terribly sorry for you, Cherryl, and I'd like to help you—not because you suffer, but because you haven't deserved to suffer."

"You mean, you wouldn't be kind to anything weak or whining or rotten about me? Only to whatever you see in me that's good?"

"Of course."

Cherryl did not move her head, but she looked as if it were lifted—as if some bracing current were relaxing her features into that rare look which combines pain and dignity.

"It's not alms, Cherryl. Don't be afraid to speak to me."

[...]

"You know, Miss Tag—Dagny," she said softly, in wonder, "you're not as I expected you to be at all. ... They, Jim and his friends, they said you were hard and cold and unfeeling."

"But it's true, Cherryl, I am, in the sense they mean—only have they ever told you in just what sense they mean it?"

"No. They never do. They only sneer at me when I ask them what they mean by anything ... about anything. What did they mean about you?"

"Whenever anyone accuses some person of being 'unfeeling', he means that that person is just. He means that that person has no causeless emotions and will not grant him a feeling which he does not deserve. He means that 'to feel' is to go against reason, against moral values, against reality. He means ... What's the matter?" she asked, seeing the abnormal intensity of the girl's face.

"It's ... it's something I've tried so hard to understand ... for such a long time. ..."

"Well, observe that you never hear that accusation in defense of innocence, but always in defense of guilt. You never hear it said by a good person about those who fail to do him justice. But you always hear it said by a rotter about those who treat him as a rotter, those who don't feel any sympathy for the evil he's committed or for the pain he suffers as a consequence. Well, it's true—that is what I do not feel. But those who feel it, feel nothing for any quality of human greatness, for any person or action that deserves admiration, approval, esteem. These are the things I feel. You'll find that it's one or the other. Those who grant sympathy to guilt, grant none to innocence. Ask yourself which, of the two, are the unfeeling persons. And then you'll see what motive is the opposite of charity."

"What?" she whispered.

“You’ll Never Persuade People Like That”

(originally published at Less Wrong)

Sometimes, when someone is arguing for some proposition, their interlocutor will reply that the speaker's choice of arguments or tone wouldn't be effective at persuading some third party.

This would seem to be an odd change of topic. If I was arguing for this-and-such proposition, and my interlocutor isn't, themselves, convinced by my arguments, it makes sense for them to reply about why they, personally, aren't convinced. Why is it relevant whether I would convince some third party that isn't here?

What's going on in this kind of situation? Why would someone think "You'll never persuade people like that" was a relevant reply?

"Because people aren't truthseeking and treat arguments as soldiers" doesn't seem like an adequate explanation by itself. It's true, but it's not specific enough: what particularly makes appeal-to-persuading-third-parties an effective "soldier"?


The bargaining model of war attempts to explain why wars are fought—and not fought; even the bitterest enemies often prefer to grudgingly make peace with each other rather than continue to fight.

That's because war is costly. If I estimate that by continuing to wage war, there's a 60% chance my armies will hold a desirable piece of territory, I can achieve my war objectives equally well in expectation—while saving a lot of money and human lives—by instead signing a peace treaty that divides the territory with the enemy 60/40.

If the enemy will agree to that, of course. The enemy has their own forecast probabilities and their own war objectives. There's usually a range of possible treaties that both combatants will prefer to fighting, but the parties need to negotiate to select a particular treaty, because there's typically no uniquely obvious "fair" treaty—similar to how a buyer and seller need to negotiate a price for a rare and expensive item for which there is no uniquely obvious "fair" price.


If war is bargaining, and arguments are soldiers, then debate is negotiation: the same game-theoretic structure shines through armies fighting over the borders on the world's political map, buyer and seller haggling over contract items, and debaters arguing over the beliefs on Society's shared map. Strong arguments, like a strong battalion, make it less tenable for the adversary to maintain their current position.

Unfortunately, the theory of interdependent decision is ... subtle. Although recent work points toward the outlines of a more elegant theory with fewer pathologies, the classical understanding of negotiation often recommends "rationally irrational" tactics in which an agent handicaps its own capabilities in order to extract concessions from a counterparty: for example, in the deadly game of chicken, if I visibly throw away my steering wheel, oncoming cars are forced to swerve for me in order to avoid a crash, but if the oncoming drivers have already blindfolded themselves, they wouldn't be able to see me throw away my steering wheel, and I am forced to swerve for them.

Thomas Schelling teaches us that one such tactic is to move the locus of the negotiation elsewhere, onto some third party who has less of an incentive to concede or is less able to be communicated with. For example, if business purchases over $500 have to be approved by my hard-to-reach boss, an impatient seller of an item that ordinarily goes for $600 might be persuaded to give me a discount.

And that's what explains the attractiveness of the appeal-to-persuading-third-parties. What "You'll never persuade people like that" really means is, "You are starting to persuade me against my will, and I'm laundering my cognitive dissonance by asserting that you actually need to persuade someone else who isn't here." When someone is desperate enough to try to get away with that, you know you've got them cornered. Go for the throat!

(Unless the belief you're arguing for is false. You checked that beforehand, right??)

“Rationalist Discourse” Is Like “Physicist Motors”

(originally published at Less Wrong)

Imagine being a student of physics, and coming across a blog post proposing a list of guidelines for "physicist motors"—motor designs informed by the knowledge of physicists, unlike ordinary motors.

Even if most of the things on the list seemed like sensible advice to keep in mind when designing a motor, the framing would seem very odd. The laws of physics describe how energy can be converted into work. To the extent that any motor accomplishes anything, it happens within the laws of physics. There are theoretical ideals describing how motors need to work in principle, like the Carnot engine, but you can't actually build an ideal Carnot engine; real-world electric motors or diesel motors or jet engines all have their own idiosyncratic lore depending on the application and the materials at hand; an engineer who worked on one, might not the be best person to work on another. You might appeal to principles of physics to explain why some particular motor is inefficient or poorly-designed, but you would not speak of physicist motors as if that were a distinct category of thing—and if someone did, you might quietly begin to doubt how much they really knew about physics.

As a student of rationality, I feel the same way about guidelines for "rationalist discourse." The laws of probability and decision theory describe how information can be converted into optimization power. To the extent that any discourse accomplishes anything, it happens within the laws of rationality.

Rob Bensinger proposes "Elements of Rationalist Discourse" as a companion to Duncan Sabien's earlier "Basics of Rationalist Discourse". Most of the things on both lists are, indeed, sensible advice that one might do well to keep in mind when arguing with people, but as Bensinger notes, "Probably this new version also won't match 'the basics' as other people perceive them."

But there's a reason for that: a list of guidelines has the wrong type signature for being "the basics". The actual basics are the principles of rationality one would appeal to explain which guidelines are a good idea: principles like how evidence is the systematic correlation between possible states of your observations and possible states of reality, how you need evidence to locate the correct hypothesis in the space of possibilities, how the quality of your conclusion can only be improved by arguments that have the power to change that conclusion.

Contemplating these basics, it should be clear that there's just not going to be anything like a unique style of "rationalist discourse", any more than there is a unique "physicist motor." There are theoretical ideals describing how discourse needs to work in principle, like Bayesian reasoners with common priors exchanging probability estimates, but you can't actually build an ideal Bayesian reasoner. Rather, different discourse algorithms (the collective analogue of "cognitive algorithm") leverage the laws of rationality to convert information into optimization in somewhat different ways, depending on the application and the population of interlocutors at hand, much as electric motors and jet engines both leverage the laws of physics to convert energy into work without being identical to each other, and with each requiring their own engineering sub-specialty to design.

Or to use another classic metaphor, there's also just not going to be a unique martial art. Boxing and karate and ju-jitsu all have their own idiosyncratic lore adapted to different combat circumstances, and a master of one would easily defeat a novice of the other. One might appeal to the laws of physics and the properties of the human body to explain why some particular martial arts school was not teaching their students to fight effectively. But if some particular karate master were to brand their own lessons as the "basics" or "elements" of "martialist fighting", you might quietly begin to doubt how much actual fighting they had done: either all fighting is "martialist" fighting, or "martialist" fighting isn't actually necessary for beating someone up.

One historically important form of discourse algorithm is debate, and its close variant the adversarial court system. It works by separating interlocutors into two groups: one that searches for arguments in favor of a belief, and another that searches for arguments against the belief. Then anyone listening to the debate can consider all the arguments to help them decide whether or not to adopt the belief. (In the court variant of debate, a designated "judge" or "jury" announces a "verdict" for or against the belief, which is added to the court's shared map, where it can be referred to in subsequent debates, or "cases.")

The enduring success and legacy of the debate algorithm can be attributed to how it circumvents a critical design flaw in individual human reasoning, the tendency to "rationalize"—to preferentially search for new arguments for an already-determined conclusion.

(At least, "design flaw" is one way of looking at it—a more complete discussion would consider how individual human reasoning capabilities co-evolved with the debate algorithm—and, as I'll briefly discuss later, this "bug" for the purposes of reasoning is actually a "feature" for the purposes of deception.)

As a consequence of rationalization, once a conclusion has been reached, even prematurely, further invocations of the biased argument-search process are likely to further entrench the conclusion, even when strong counterarguments exist (in regions of argument-space neglected by the biased search). The debate algorithm solves this sticky-conclusion bug by distributing a search for arguments and counterarguments among multiple humans, ironing out falsehoods by pitting two biased search processes against each other. (For readers more familiar with artificial than human intelligence, generative adversarial networks work on a similar principle.)

For all its successes, the debate algorithm also suffers from many glaring flaws. For one example, the benefits of improved conclusions mostly accrue to third parties who haven't already entrenched on a conclusion; debate participants themselves are rarely seen changing their minds. For another, just the choice of what position to debate has a distortionary effect even on the audience; if it takes more bits to locate a hypothesis for consideration than to convincingly confirm or refute it, then most of the relevant cognition has already happened by the time people are arguing for or against it. Debate is also inefficient: for example, if the "defense" in the court variant happens to find evidence or arguments that would benefit the "prosecution", the defense has no incentive to report it to the court, and there's no guarantee that the prosecution will independently find it themselves.

Really, the whole idea is so galaxy-brained that it's amazing it works at all. There's only one reality, so correct information-processing should result in everyone agreeing on the best, most-informed belief-state. This is formalized in Aumann's famous agreement theorem

That being the normative math, why does the human world's enduringly dominant discourse algorithm take for granted the ubiquity of, not just disagreements, but predictable disagreements? Isn't that crazy?

Yes. It is crazy. One might hope to do better by developing some sort of training or discipline that would allow discussions between practitioners of such "rational arts" to depart from the harnessed insanity of the debate algorithm with its stubbornly stable "sides", and instead mirror the side-less Bayesian ideal, the free flow of all available evidence channeling interlocutors to an unknown destination.

Back in late 'aughts, an attempt to articulate what such a discipline might look like was published on a blog called Overcoming Bias. (You probably haven't heard of it.) It's been well over a decade since then. How is that going?

Eliezer Yudkowsky laments:

In the end, a lot of what people got out of all that writing I did, was not the deep object-level principles I was trying to point to—they did not really get Bayesianism as thermodynamics, say, they did not become able to see Bayesian structures any time somebody sees a thing and changes their belief. What they got instead was something much more meta and general, a vague spirit of how to reason and argue, because that was what they'd spent a lot of time being exposed to over and over and over again in lots of blog posts.

"A vague spirit of how to reason and argue" seems like an apt description of what "Basics of Rationalist Discourse" and "Elements of Rationalist Discourse" are attempting to codify—but with no explicit instruction on which guidelines arise from deep object-level principles of normative reasoning, and which from mere taste, politeness, or adaptation to local circumstances, it's unclear whether students of 2020s-era "rationalism" are poised to significantly outperform the traditional debate algorithm—and it seems alarmingly possible to do worse, if the collaborative aspects of modern "rationalist" discourse allow participants to introduce errors that a designated adversary under the debate algorithm would have been incentivized to correct, and most "rationalist" practitioners don't have a deep theoretical understanding of why debate works as well as it does.

Looking at Bensinger's "Elements", there's a clear-enough connection between the first eight points (plus three sub-points) and the laws of normative reasoning. Truth-Seeking, Non-Deception, and Reality-Minding, trivial. Non-Violence, because violence doesn't distinguish between truth and falsehood. Localizability, in that I can affirm the validity of an argument that A would imply B, while simultaneously denying A. Alternative-Minding, because decisionmaking under uncertainty requires living in many possible worlds. And so on. (Lawful justifications for the elements of Reducibility and Purpose-Minding left as an exercise to the reader.)

But then we get this:

  1. Goodwill. Reward others' good epistemic conduct (e.g., updating) more than most people naturally do. Err on the side of carrots over sticks, forgiveness over punishment, and civility over incivility, unless someone has explicitly set aside a weirder or more rough-and-tumble space.

I can believe that these are good ideas for having a pleasant conversation. But separately from whether "Err on the side of forgiveness over punishment" is a good idea, it's hard to see how it belongs on the same list as things like "Try not to 'win' arguments using [...] tools that work similarly well whether you're right or wrong" and "[A]sk yourself what Bayesian evidence you have that you're not in those alternative worlds".

The difference is this. If your discourse algorithm lets people "win" arguments with tools that work equally well whether they're right or wrong, then your discourse gets the wrong answer (unless, by coincidence, the people who are best at winning are also the best at getting the right answer). If the interlocutors in your discourse don't ask themselves what Bayesian evidence they have that they're not in alternative worlds, then your discourse gets the wrong answer (if you happen to live in an alternative world).

If your discourse algorithm errs on the side of sticks over carrots (perhaps, emphasizing punishing others' bad epistemic conduct more than most people naturally do), then ... what? How, specifically, are rough-and-tumble spaces less "rational", more prone to getting the wrong answer, such that a list of "Elements of Rationalist Discourse" has the authority to designate them as non-default?

I'm not saying that goodwill is bad, particularly. I totally believe that goodwill is a necessary part of many discourse algorithms that produce maps that reflect the territory, much like how kicking is a necessary part of many martial arts (but not boxing). It just seems like a bizarre thing to put in a list of guidelines for "rationalist discourse".

It's as if guidelines for designing "physicist motors" had a point saying, "Use more pistons than most engineers naturally do." It's not that pistons are bad, particularly. Lots of engine designs use pistons! It's just, the pistons are there specifically to convert force from expanding gas into rotational motion. I'm pretty pessimistic about the value of attempts to teach junior engineers to mimic the surface features of successful engines without teaching them how engines work, even if the former seems easier.

The example given for "[r]eward[ing] others' good epistemic conduct" is "updating". If your list of "Elements of Rationalist Discourse" is just trying to apply a toolbox of directional nudges to improve the median political discussion on social media (where everyone is yelling and no one is thinking), then sure, directionally nudging people to directionally nudge people to look like they're updating probably is a directional improvement. It still seems awfully unambitious, compared to trying to teach the criteria by which we can tell it's an improvement. In some contexts (in-person interactions with someone I like or respect), I think I have the opposite problem, of being disposed to agree with the person I'm currently talking to, in a way that shortcuts the slow work of grappling with their arguments and doesn't stick after I'm not talking to them anymore; I look as if I'm "updating", but I haven't actually learned. Someone who thought "rationalist discourse" entailed "[r]eward[ing] others' good epistemic conduct (e.g., updating) more than most people naturally do" and sought to act on me accordingly would be making that problem worse.

A footnote on the "Goodwill" element elaborates:

Note that this doesn't require assuming everyone you talk to is honest or has good intentions.

It does have some overlap with the rule of thumb "as a very strong but defeasible default, carry on object-level discourse as if you were role-playing being on the same side as the people who disagree with you".

But this seems to contradict the element of Non-Deception. If you're not actually on the same side as the people who disagree with you, why would you (as a very strong but defeasible default) role-play otherwise?

Other intellectual communities have a name for the behavior of role-playing being on the same side as people you disagree with: they call it "concern trolling", and they think it's a bad thing. Why is that? Are they just less rational than "us", the "rationalists"?

Here's what I think is going on. There's another aspect to the historical dominance of the debate algorithm. The tendency to rationalize new arguments for a fixed conclusion is only a bug if one's goal is to improve the conclusion. If the fixed conclusion was adopted for other reasons—notably, because one would benefit from other people believing it—then generating new arguments might help persuade those others. If persuading others is the real goal, then rationalization is not irrational; it's just dishonest. (And if one's concept of "honesty" is limited to not consciously making false statements, it might not even be dishonest.) Society benefits from using the debate algorithm to improve shared maps, but most individual debaters are mostly focused on getting their preferred beliefs onto the shared map.

That's why people don't like concern trolls. If my faction is trying to get Society to adopt beliefs that benefit our faction onto the shared map, someone who comes to us role-playing being on our side, but who is actually trying to stop us from adding our beliefs to the shared map just because they think our beliefs don't reflect the territory, isn't a friend; they're a double agent, an enemy pretending to be a friend, which is worse than the honest enemy we expect to face before the judge in the debate hall.

This vision of factions warring to make Society's shared map benefit themselves is pretty bleak. It's tempting to think the whole mess could be fixed by starting a new faction—the "rationalists"—that is solely dedicated to making Society's shared map reflect the territory: a culture of clear thinking, clear communication, and collaborative truth-seeking.

I don't think it's that simple. You do have interests, and if you can fool yourself into thinking that you don't, your competitors are unlikely to fall for it. Even if your claim to only want Society's shared map to reflect the territory were true—which it isn't—anyone could just say that.

I don't immediately have solutions on hand. Just an intuition that, if there is any way of fixing this mess, it's going to involve clarifying conflicts rather than obfuscating them—looking for Pareto improvements, rather than pretending that everyone has the same utility function. That if something called "rationalism" is to have any value whatsoever, it's as the field of study that can do things like explain why it makes sense that people don't like concern trolling. Not as as its own faction with its own weird internal social norms that call for concern trolling as a very strong but defeasible default.

But don't take my word for it.

Don’t Double-Crux With Suicide Rock

(originally published at Less Wrong)

Honest rational agents should never agree to disagree.

This idea is formalized in Aumann's agreement theorem and its various extensions (we can't foresee to disagree, uncommon priors require origin disputes, complexity bounds, &c.), but even without the sophisticated mathematics, a basic intuition should be clear: there's only one reality. Beliefs are for mapping reality, so if we're asking the same question and we're doing everything right, we should get the same answer. Crucially, even if we haven't seen the same evidence, the very fact that you believe something is itself evidence that I should take into account—and you should think the same way about my beliefs.

In "The Coin Guessing Game", Hal Finney gives a toy model illustrating what the process of convergence looks like in the context of a simple game about inferring the result of a coinflip. A coin is flipped, and two players get a "hint" about the result (Heads or Tails) along with an associated hint "quality" uniformly distributed between 0 and 1. Hints of quality 1 always match the actual result; hints of quality 0 are useless and might as well be another coinflip. Several "rounds" commence where players simultaneously reveal their current guess of the coinflip, incorporating both their own hint and its quality, and what they can infer about the other player's hint quality from their behavior in previous rounds. Eventually, agreement is reached. The process is somewhat alien from a human perspective (when's the last time you and an interlocutor switched sides in a debate multiple times before eventually agreeing?!), but not completely so: if someone whose rationality you trusted seemed visibly unmoved by your strongest arguments, you would infer that they had strong evidence or counterarguments of their own, even if there was some reason they couldn't tell you what they knew.

Honest rational agents should never agree to disagree.

In "Disagree With Suicide Rock", Robin Hanson discusses a scenario where disagreement seems clearly justified: if you encounter a rock with words painted on it claiming that you, personally, should commit suicide according to your own values, you should feel comfortable disagreeing with the words on the rock without fear of being in violation of the Aumann theorem. The rock is probably just a rock. The words are information from whoever painted them, and maybe that person did somehow know something about whether future observers of the rock should commit suicide, but the rock itself doesn't implement the dynamic of responding to new evidence.

In particular, if you find yourself playing Finney's coin guessing game against a rock with the letter "H" painted on it, you should just go with your own hint: it would be incorrect to reason, "Wow, the rock is still saying Heads, even after observing my belief in several previous rounds; its hint quality must have been very high."

Honest rational agents should never agree to disagree.

Human so-called "rationalists" who are aware of this may implicitly or explicitly seek agreement with their peers. If someone whose rationality you trusted seemed visibly unmoved by your strongest arguments, you might think, "Hm, we still don't agree; I should update towards their position ..."

But another possibility is that your trust has been misplaced. Humans suffering from "algorithmic bad faith" are on a continuum with Suicide Rock. What matters is the counterfactual dependence of their beliefs on states of the world, not whether they know all the right keywords ("crux" and "charitable" seem to be popular these days), nor whether they can perform the behavior of "making arguments"—and definitely not their subjective conscious verbal narratives.

And if the so-called "rationalists" around you suffer from correlated algorithmic bad faith—if you find yourself living in a world of painted rocks—then it may come to pass that protecting the sanctity of your map requires you to master the technique of lonely dissent.

Idiot or Alien? Incompetence or Evil?

When you encounter someone who expresses a political or social opinion that you find absolutely abhorrent, it is instructive to consider the extent to which this person is making a mistake, and the extent to which they simply have different values from you. Is this opinion something that they would immediately relinquish, if only they knew they knew the true facts of which they are now ignorant?—or is it reflective of some quality essential to their agency, a basic motive far too sacred to be destroyed by the truth?

(Of course, it is also instructive to consider whether you're making a mistake. But that is not the subject of this post.)

Some would say that it is useless to consider such questions, that human cognition doesn't separate cleanly into beliefs and values, and that even if such a thing could be done, it is futile for any present-day human to consider the matter, given our ignorance of our own psychology. And yet, the question still seems to make sense to me. If I can't know, I can guess. And I don't guess the same thing every time.

It's hard to say which extreme is more terrifying. In one scenario, you want to cry out to them, "Oh, you fool! You beautiful, beautiful fool! I love you and I want to be your friend, trust that I will always want to be your friend, but don't you see that the path you're taking can only lead to disaster? If you give me some time I can explain my reasoning precisely, but all the evidence points to the same conclusion: you must turn back now, I beg you, for the sake of everything we hold dear!"

But you know that wouldn't work, so you say nothing.

In the other scenario, you instinctively know that appeals to emotion or common goals would be a waste of precious time, so you frantically search an argument, some sequence of facts and reasoning that will convince them, convince any halfway-rational creature, to stop doing this terrible thing—but it's clear that no such argument exists. Any fact or reason you could offer would only be interpreted as evidence about reality, incorporated into their world-model, and used to persue their monstrous goals that much more efficiently.

You say nothing, but as you look into your enemy's eyes as they hasten the destruction of your world, you can't shake the feeling of having been understood.