An Algorithmic Lucidity

a blog

Category: philosophy

Schelling Categories, and Simple Membership Tests

(originally published at Less Wrong)

Followup to: Where to Draw the Boundaries?

Or there might be social or psychological forces anchoring word usages on identifiable Schelling points that are easy for different people to agree upon, even at the cost of some statistical "fit" ...

The one comes to you and says, "That paragraph about Schelling points sounded interesting. What did you mean by that? Can you give an example?"

Sure. Previously on Less Wrong, in "The Univariate Fallacy", we studied points sampled from two multivariate probability distributions \(P_A\) and \(P_B\), and showed that it was possible to infer with very high probability which distribution a given point was sampled from, despite significant overlap in the marginal distributions for any one variable considered individually.

From the standpoint of "the way to carve reality at its joints, is to draw your boundaries around concentrations of unusually high probability density in Thingspace", the correct categorization of the points in that example is clear. We have two clearly distinguishable clusters. The conditional independence property is satisfied: given a point's cluster-membership, knowing one of the \(x_i\) doesn't tell you anything about \(x_j\) for j ≠ i. So we should draw a category boundary around each cluster. Obviously. We might ask hypophorically: what could possibly change this moral?

More constraints on the problem, that's what!

Suppose you needed to coordinate with someone else to make decisions about these points—that is, it's important not just that you and your partner make good decisions, but also that you make the same decision—but that each of you only got to observe one coordinate from each point. As we saw, the predictive work we get from category-membership in this scenario is spread across many variables: if you only get to observe a few dimensions, you have a lot of uncertainty about cluster-membership (which carries over into additional uncertainty about the other dimensions that you haven't observed, but which affect the ex post quality of your decision).

If you and your partner were both ideal Bayesian calculators who could communicate costlessly, you would share your observations, work out the correct probability, and use that to make optimal decisions. But suppose you couldn't do that—either because communication is expensive, or your partner was bad at math, or any other reason. Then it would be sad if you happened to see \(x_9\) = 2 and said "It's an A (probably)!", and your partner happened to see \(x_{27}\) = 3 and said "It's a B (I think)!", and the two of you made inconsistent decisions.

Okay, now suppose that there's actually a forty-first, binary, variable that I didn't tell you about earlier, distributed like so:

$$P_A(x_{41}) = \begin{cases} 3/4 & x_{41} = 0 \\ 1/4 & x_{41} = 1 \\ \end{cases}$$
$$P_B(x_{41}) = \begin{cases} 1/4 & x_{41} = 0 \\ 3/4 & x_{41} = 1 \\ \end{cases}$$

Observing \(x_{41}\) gives you \(\log_2 3\) ≈ 1.585 bits of evidence about cluster-membership, which is more than the

$$\frac{1/4 + 1/16}{2} \cdot |\log_2(4)| + \frac{7/16 + 1/4}{2} \cdot |\log_2(7/4)| + \frac{1/4 + 7/16}{2} \cdot |\log_2(4/7)| + \frac{1/16 + 1/4}{2} \cdot |\log_2(4)|$$

≈ 1.18 bits you can get from any one observation of one of the \(x_i\) for i ∈ {1...40}.

If you and your partner can both observe \(x_{41}\), you might end up wanting to base your shared categories and language on that—calling a point an "A" if it has \(x_{41}\) = 0, even though such points actually came from \(P_B\) a full quarter of the time—even if \(x_{41}\) itself has no effect on the quality of your decisions, and what you actually care about is wholely determined by the values of \(x_1\) through \(x_{40}\)! It's not the intension you would pick if you could make (and share) more observations—but ex hypothesi, you can't.

If you and your partner only get to observe one variable, \(x_{41}\) is your best choice—the single variable that gives you the most information about the "natural" cluster-membership. That also makes it a Schelling point—if you and your partner didn't get to commmunicate in advance about how you want to draw your shared category boundaries, you could pick \(x_{41}\) as your defining observation and be pretty confident your partner would make the same choice. We could imagine an even more pessimistic scenario in which the Schelling point category definition (a set of variables that "stuck out" from all the others) was less predictive than some other candidates—but if you couldn't coordinate to pick one of the more predictive category systems, you might be stuck with the Schelling point.

In conclusion, the right categories to use given constraints on communication and observation, might be different from the category boundaries you would draw from a "God's eye view", in part because consideration of which categories are easy for different agents to coordinate on is relevant, not just raw information-theoretic expressive power. Thus, "Schelling categories."

Thanks for reading!


The one says, "No, I meant, like, a real world example, not some dumb math thing for nerds. What is this post really about?"

It's about ... math? Or like, the relationship between math and human natural language? Like, I was wondering what "second-order" caveats or complications there might be to the basic "carve reality at the joints" moral of our standard Bayesian philosophy of language, and some of the people I've been collaborating with lately had been talking a lot about the importanace of intersubjective epistemology—that is, shared mapmaking, so—

"But where's the actionable takeaway? What's your real agenda here, huh?"

Oh. One of those readers, I see. Fine, I can probably think of some—how do you say?—"applications."

Ummmm ...

Let's see ...

Okay, here's something, maybe. What's the deal with the age of majority?

Society needs to decide who it wants to be allowed to vote, stand trial, sign contracts, serve in the military, &c. Whether it's a good idea for a particular person to have these privileges presumably depends on various relevant features of that person: things like cognitive ability, foresight, wisdom, relevant life experiences, &c. In particular, it would be pretty weird for someone's fitness to vote to directly depend on how many times the Earth has gone around the sun since they were born. What does that number have to do with anything?

It doesn't! But if Society isn't well-coordinated enough to agree on the exact prerequisites for voting and how to measure them, but can agree that most twenty-five-year-olds have them and most eleven-year-olds don't, then we end up choosing some arbitrary age cutoff as the criterion for our "legal adulthood" social construct. It works, but it's just a legal fiction—and not necessarily a particularly good fiction, as any bright teenagers reading this will doubtlessly attest.

If I told you that a particular fourteen-year-old was very "mature", that's a contentful statement: we have shared meaning attached to the word mature, such that my describing someone that way constrains your anticipations. But it's a really complicated meaning, a statistical signal in behavior that your brain can pick up on, but which isn't particularly verifiable to others who might have reasons to doubt my character assessment. In contrast, age is easy for everyone to agree on. We could imagine some hypothetical science-fictional Society that used brain scans and some sophisticated machine-learning classifer to determine which citizens get which privileges—but in our dumber, poorer world, calendars and subtraction will have to do.

In terms of Scott Garrabrant's taxonomy of applications of Goodhart's law, this is regressional Goodhart: Society wants to select for maturity, chooses age as a proxy, and in the process, ends up granting or withholding privileges that a more discriminating Society maybe wouldn't.

The age of majority is a case of replacing a complicated, illegible category ("maturity", the kind of abstract thing you might want to model as a cluster in a forty- or forty-one-dimensional space) with a simple membership test (an age cutoff that everyone knows how to compute). Different people might make make different subjective (but not arbitrary) judgements of the complicated, illegible category, so in order to get a more intersubjectively robust verdict on category-membership, we rely on an objective measurement that everyone can agree on.

If no convenient objective measurement is available, another strategy is possible: we can delegate to some canonical trusted authority, whose opinion of the complicated category will take precdence over everyone else's. An example of this is commodity grading standards. What is a "Grade AA" egg? Well, there's a complicated definition written down in a manual somewhere that you could try applying yourself—but for most people, Grade AA eggs are simply "those which have been certified as Grade AA by the USDA."1

It's even possible for the "simple objective measurement" and "delegate to an authority's subjective judgement" strategies to be combined. In "The Ideology Is Not the Movement", the immortal Scott Alexander writes about his model of the genesis of social groups—

Pre-existing differences are the raw materials out of which tribes are made. A good tribe combines people who have similar interests and styles of interaction even before the ethnogenesis event. Any description of these differences will necessarily involve stereotypes, but a lot of them should be hard to argue. [...] There are subtle habits of thought, not yet described by any word or sentence, which atheists are more likely to have than other people. [...]

The rallying flag is the explicit purpose of the tribe. It's usually a belief, event, or activity that get people with that specific pre-existing difference together and excited. Often it brings previously latent differences into sharp relief. People meet around the rallying flag, encounter each other, and say "You seem like a kindred soul!" or "I thought I was the only one!" Usually it suggests some course of action, which provides the tribe with a purpose.

Eliezer Yudkowsky's "A Fable of Science and Politics" depicts a fictional underground society split between two such tribes: an predominantly urban tribe that believes that the unseen sky is blue (and favors an income tax, strong marriage laws, and an Earth-centric cosmology), and predominanty rural one that believes that the sky is green (and favors merchant taxes, no-fault divorce, and a heliocentric cosmology). In this story, beliefs about the color of the sky are functioning as the "rallying flag" for tribe-formation in Alexander's model—and as a Schelling point for category definition.

We don't know how to talk about the preëxisting undefinable habits of thought that make social groups work—it's hard to explicitly articulate what exact statistical regularity our brains have detected in five-and-more-dimensional locale/sky-belief/tax-belief/divorce-belief/cosmology/&c.-space. (Although we could imagine some hypothetical science-fictional Society that did know how to articulate it, and consequently had richer forms of social and political organization than our own.) It's a lot simpler to talk about whether someone has pledged allegiance to the rallying flag: just ask someone, "What color do you believe the sky is?" (using sky-beliefs as as an "objective" simple membership test), or simply, "Are you a Blue or a Green?" (delegating the classification problem to the person themselves as the authority whose discernment is to be trusted)—and whatever they say, that's what they are.

Well, probably. We've seen that objective measurements like age are subject to regressional Goodhart, but the delegation-to-authority strategy is furthermore subject to adversarial Goodhart: once a category-membership test has been established, some agents might have an incentive to create examples that pass the test, but don't have the complicated, illegible properties than made the test a useful proxy in the first place.

We've seen this, for example, with title inflation: we expect the "job title" (the words that get printed on business cards or immigration sponsorship forms) to be the canonical description of what someone "does", even if the vagaries of the workday encompass many tasks,2 and an alien anthropologist tasked with observing the worksite and summarizing what each of the humans did might slice up her observations into categories with little resemblance to the company's formal org chart. But since we don't know how to do the obvious thing and average over all possible alien anthropologists weighted by simplicity, we can only rely on the org chart—which people have political incentives to manipulate, with the result that everyone in the finance industry is a "vice president" of some sort or another.

But "Vice President" has a literal meaning. Or it used to. Vice, "in place of; subordinate to." President, one who presides over some deliberative body. The adversarial-Goodhart pressures on language "exploit[ ] the trust we have in a functioning piece of language until it's lost all meaning".

So for readers who demand a takeaway beyond just an edge case in the math, perhaps take away this: coordination is costly. From the standpoint of language as an AI capability, the social constructions that feeble humans need in order to work together may be unavoidably dumbed-down for mass consumption, but that's no reason to not aspire to the true precision of the Bayes-structure to whatever extent possible.

(Thanks to Ben Hoffman for the etymology of "Vice President.")


  1. Or the analogous agency in your country. ↩

  2. When I worked in a supermarket, two days a week I did Tracy's bookkeeping/customer-service job while Tracy had her weekend, which entailed counting the money from last night's tills and swapping in new coinmags and completing the FSM report and answering the phone and selling money orders and covering the floral stand when the floral lady was on lunch, &c. I'm actually not sure what official name this role had in Safeway's official org chart. We just called it "the booth." ↩

Being Wrong Doesn't Mean You're Stupid and Bad (Probably)

(originally published at Less Wrong)

Sometimes, people are reluctant to admit that they were wrong about something, because they're afraid that "You are wrong about this" carries inextricable connotations of "You are stupid and bad." But this behavior is, itself, wrong, for at least two reasons.

First, because it's evidential decision theory. The so-called "rationalist" "community" has a lot of cached clichés about this! A blank map does not correspond to a blank territory. What's true is already so; owning up to it doesn't make it worse. Refusing to go to the doctor (thereby avoiding encountering evidence that you're sick) doesn't keep you healthy.

If being wrong means that you're stupid and bad, then preventing yourself from knowing that you were wrong doesn't stop you from being stupid and bad in reality. It just prevents you from knowing that you're stupid and bad—which is an important fact to know (if it's true), because if you don't know that you're stupid and bad, then it probably won't occur to you to even look for possible interventions to make yourself less stupid and less bad.

Second, while "You are wrong about this" is evidence for the "You are stupid and bad" hypothesis if stupid and bad people are more likely to be wrong, I claim that it's very weak evidence. (Although it's possible that I'm wrong about this—and if I'm wrong, it's furthermore possible that the reason I'm wrong is because I'm stupid and bad.)

Exactly how weak evidence is it? It's hard to guess directly, but fortunately, we can use probability theory to reduce the claim into more "atomic" conditional and prior probabilities that might be easier to estimate!

Let \(W\) represent the proposition "You are wrong about something", \(S\) represent the proposition "You are stupid", and \(B\) represent the proposition "You are bad."

By Bayes's theorem, the probability that you are stupid and bad given that you're wrong about something is given by—

$$P(S,B|W)=\frac{P(W|S,B)P(S,B)}{P(W|S,B)P(S,B)+P(W|S, \neg B)P(S, \neg B)+P(W| \neg S,B)P( \neg S,B)+P(W| \neg S, \neg B)P( \neg S, \neg B)}$$

For the purposes of this calculation, let's assume that badness and stupidity are statistically independent. I doubt this is true in the real world, but because I'm stupid and bad (at math), I want that simplifying assumption to make the algebra easier for me. That lets us unpack the conjunctions, giving us—

$$P(S,B|W)=\frac{P(W|S,B)P(S)P(B)}{P(W|S,B)P(S)P(B)+P(W|S, \neg B)P(S)P(\neg B)+P(W| \neg S,B)P( \neg S)P(B)+P(W| \neg S, \neg B)P( \neg S)P(\neg B)}$$

This expression has six degrees of freedom: \(P(S)\), \(P(B)\), \(P(W|S,B)\), \(P(W|S, \neg B)\), \(P(W|\neg S,B)\), \(P(W|\neg S, \neg B)\). Arguing about the values of these six individual parameters is probably more productive than arguing about the value of \(P(S,B|W)\) directly!

Suppose half the people are stupid (\(P(S) = 0.5\)), one-tenth of people are bad (\(P(B) = 0.1\)), and that most people are wrong, but that being stupid or bad each make you somewhat more likely to be wrong, to the tune of \(P(W|\neg S, \neg B) = 0.8\), \(P(W|S, \neg B) = P(W|\neg S,B) = 0.85\), and \(P(W|S,B) = 0.9\). So our posterior probabilty that someone is stupid and bad given that they were wrong once is

$$P(S,B|W) = \frac{(0.9)(0.5)(0.1)}{(0.9)(0.5)(0.1)+(0.85)(0.5)(0.9)+(0.85)(0.5)(0.1)+(0.8)(0.5)(0.9)}$$
$$\approx 0.0542$$

But the base rate of being stupid and bad is (0.1)(0.5) = 0.05. Learning that someone was wrong only raised our probability that they are stupid and bad by 0.0042. That's a small number that you shouldn't worry about!

“But It Doesn’t Matter”

(originally published at Less Wrong)

If you ever find yourself saying, "Even if Hypothesis H is true, it doesn't have any decision-relevant implications," you are rationalizing! The fact that H is interesting enough for you to be considering the question at all (it's not some arbitrary trivium like the 1923th binary digit of π, or the low temperature in São Paulo on September 17, 1978) means that it must have some relevance to the things you care about. It is vanishingly improbable that your optimal decisions are going to be the same in worlds where H is true and worlds where H is false. The fact that you're tempted to say they're the same is probably because some part of you is afraid of some of the imagined consequences of H being true. But H is already true or already false! If you happen to live in a world where H is true, and you make decisions as if you lived in a world where H is false, you are thereby missing out on all the extra utility you would get if you made the H-optimal decisions instead! If you can figure out exactly what you're afraid of, maybe that will help you work out what the H-optimal decisions are. Then you'll be a better position to successfully notice which world you actually live in.

Where to Draw the Boundaries?

(originally published at Less Wrong)

Followup to: Where to Draw the Boundary?

Figuring where to cut reality in order to carve along the joints—figuring which things are similar to each other, which things are clustered together: this is the problem worthy of a rationalist. It is what people should be trying to do, when they set out in search of the floating essence of a word.

Once upon a time it was thought that the word "fish" included dolphins ...

The one comes to you and says:

The list: {salmon, guppies, sharks, dolphins, trout} is just a list—you can't say that a list is wrong. You draw category boundaries in specific ways to capture tradeoffs you care about: sailors in the ancient world wanted a word to describe the swimming finned creatures that they saw in the sea, which included salmon, guppies, sharks—and dolphins. That grouping may not be the one favored by modern evolutionary biologists, but an alternative categorization system is not an error, and borders are not objectively true or false. You're not standing in defense of truth if you insist on a word, brought explicitly into question, being used with some particular meaning. So my definition of fish cannot possibly be 'wrong,' as you claim. I can define a word any way I want—in accordance with my values!

So, there is a legitimate complaint here. It's true that sailors in the ancient world had a legitimate reason to want a word in their language whose extension was {salmon, guppies, sharks, dolphins, ...}. (And modern scholars writing a translation for present-day English speakers might even translate that word as fish, because most members of that category are what we would call fish.) It indeed would not necessarily be helping the sailors to tell them that they need to exclude dolphins from the extension of that word, and instead include dolphins in the extension of their word for {monkeys, squirrels, horses ...}. Likewise, most modern biologists have little use for a word that groups dolphins and guppies together.

When rationalists say that definitions can be wrong, we don't mean that there's a unique category boundary that is the True floating essence of a word, and that all other possible boundaries are wrong. We mean that in order for a proposed category boundary to not be wrong, it needs to capture some statistical structure in reality, even if reality is surprisingly detailed and there can be more than one such structure.

The reason that the sailor's concept of water-dwelling animals isn't necessarily wrong (at least within a particular domain of application) is because dolphins and fish actually do have things in common due to convergent evolution, despite their differing ancestries. If we've been told that "dolphins" are water-dwellers, we can correctly predict that they're likely to have fins and a hydrodynamic shape, even if we've never seen a dolphin ourselves. On the other hand, if we predict that dolphins probably lay eggs because 97% of known fish species are oviparous, we'd get the wrong answer.

A standard technique for understanding why some objects belong in the same "category" is to (pretend that we can) visualize objects as existing in a very-high-dimensional configuration space, but this "Thingspace" isn't particularly well-defined: we want to map every property of an object to a dimension in our abstract space, but it's not clear how one would enumerate all possible "properties." But this isn't a major concern: we can form a space with whatever properties or variables we happen to be interested in. Different choices of properties correspond to different cross sections of the grander Thingspace. Excluding properties from a collection would result in a "thinner", lower-dimensional subspace of the space defined by the original collection of properties, which would in turn be a subspace of grander Thingspace, just as a line is a subspace of a plane, and a plane is a subspace of three-dimensional space.

Concerning dolphins: there would be a cluster of water-dwelling animals in the subspace of dimensions that water-dwelling animals are similar on, and a cluster of mammals in the subspace of dimensions that mammals are similar on, and dolphins would belong to both of them, just as the vector [1.1, 2.1, 9.1, 10.2] in the four-dimensional vector space ℝ⁴ is simultaneously close to [1, 2, 2, 1] in the subspace spanned by x₁ and x₂, and close to [8, 9, 9, 10] in the subspace spanned by x₃ and x₄.

Humans are already functioning intelligences (well, sort of), so the categories that humans propose of their own accord won't be maximally wrong: no one would try to propose a word for "configurations of matter that match any of these 29,122 five-megabyte descriptions but have no other particular properties in common." (Indeed, because we are not-superexponentially-vast minds that evolved to function in a simple, ordered universe, it actually takes some ingenuity to construct a category that wrong.)

This leaves aspiring instructors of rationality in something of a predicament: in order to teach people how categories can be more or (ahem) less wrong, you need some sort of illustrative example, but since the most natural illustrative examples won't be maximally wrong, some people might fail to appreciate the lesson, leaving one of your students to fill in the gap in your lecture series eleven years later.

The pedagogical function of telling people to "stop playing nitwit games and admit that dolphins don't belong on the fish list" is to point out that, without denying the obvious similarities that motivated the initial categorization {salmon, guppies, sharks, dolphins, trout, ...}, there is more structure in the world: to maximize the (logarithm of the) probability your world-model assigns to your observations of dolphins, you need to take into consideration the many aspects of reality in which the grouping {monkeys, squirrels, dolphins, horses ...} makes more sense. To the extent that relying on the initial category guess would result in a worse Bayes-score, we might say that that category is "wrong." It might have been "good enough" for the purposes of the sailors of yore, but as humanity has learned more, as our model of Thingspace has expanded with more dimensions and more details, we can see the ways in which the original map failed to carve reality at the joints.


The one replies:

But reality doesn't come with its joints pre-labeled. Questions about how to draw category boundaries are best understood as questions about values or priorities rather than about the actual content of the actual world. I can call dolphins "fish" and go on to make just as accurate predictions about dolphins as you can. Everything we identify as a joint is only a joint because we care about it.

No. Everything we identify as a joint is a joint not "because we care about it", but because it helps us think about the things we care about.

Which dimensions of Thingspace you bother paying attention to might depend on your values, and the clusters returned by your brain's similarity-detection algorithms might "split" or "collapse" according to which subspace you're looking at. But in order for your map to be useful in the service of your values, it needs to reflect the statistical structure of things in the territory—which depends on the territory, not your values.

There is an important difference between "not including mountains on a map because it's a political map that doesn't show any mountains" and "not including Mt. Everest on a geographic map, because my sister died trying to climb Everest and seeing it on the map would make me feel sad."

There is an important difference between "identifying this pill as not being 'poison' allows me to focus my uncertainty about what I'll observe after administering the pill to a human (even if most possible minds have never seen a 'human' and would never waste cycles imagining administering the pill to one)" and "identifying this pill as not being 'poison', because if I publicly called it 'poison', then the manufacturer of the pill might sue me."

There is an important difference between having a utility function defined over a statistical model's performance against specific real-world data (even if another mind with different values would be interested in different data), and having a utility function defined over features of the model itself.

Remember how appealing to the dictionary is irrational when the actual motivation for an argument is about whether to infer a property on the basis of category-membership? But at least the dictionary has the virtue of documenting typical usage of our shared communication signals: you can at least see how "You're defecting from common usage" might feel like a sensible thing to say, even if one's true rejection lies elsewhere. In contrast, this motion of appealing to personal values (!?!) is so deranged that Yudkowsky apparently didn't even realize in 2008 that he might need to warn us against it!

You can't change the categories your mind actually uses and still perform as well on prediction tasks—although you can change your verbally reported categories, much as how one can verbally report "believing" in an invisible, inaudible, flour-permeable dragon in one's garage without having any false anticipations-of-experience about the garage.

This may be easier to see with a simple numerical example.

Suppose we have some entities that exist in the three-dimensional vector space ℝ³. There's one cluster of entities centered at [1, 2, 3], and we call those entities Foos, and there's another cluster of entities centered at [2, 4, 6], which we call Quuxes.

The one comes and says, "Well, I'm going redefine the meaning of 'Foo' such that it also includes the things near [2, 4, 6] as well as the Foos-with-respect-to-the-old-definition, and you can't say my new definition is wrong, because if I observe [2, _, _] (where the underscores represent yet-unobserved variables), I'm going to categorize that entity as a Foo but still predict that the unobserved variables are 4 and 6, so there."

But if the one were actually using the new concept of Foo internally and not just saying the words "categorize it as a Foo", they wouldn't predict 4 and 6! They'd predict 3 and 4.5, because those are the average values of a generic Foo-with-respect-to-the-new-definition in the 2nd and 3rd coordinates (because (2+4)/2 = 6/2 = 3 and (3+6)/2 = 9/2 = 4.5). (The already-observed 2 in the first coordinate isn't average, but by conditional independence, that only affects our prediction of the other two variables by means of its effect on our "prediction" of category-membership.) The cluster-structure knowledge that "entities for which x₁≈2, also tend to have x₂≈4 and x₃≈6" needs to be represented somewhere in the one's mind in order to get the right answer. And given that that knowledge needs to be represented, it might also be useful to have a word for "the things near [2, 4, 6]" in order to efficiently share that knowledge with others.

Of course, there isn't going to be a unique way to encode the knowledge into natural language: there's no reason the word/symbol "Foo" needs to represent "the stuff near [1, 2, 3]" rather than "both the stuff near [1, 2, 3] and also the stuff near [2, 4, 6]". And you might very well indeed want a short word like "Foo" that encompasses both clusters, for example, if you want to contrast them to another cluster much farther away, or if you're mostly interested in x₁ and the difference between x₁≈1 and x₁≈2 doesn't seem large enough to notice.

But if speakers of particular language were already using "Foo" to specifically talk about the stuff near [1, 2, 3], then you can't swap in a new definition of "Foo" without changing the truth values of sentences involving the word "Foo." Or rather: sentences involving Foo-with-respect-to-the-old-definition are different propositions from sentences involving Foo-with-respect-to-the-new-definition, even if they get written down using the same symbols in the same order.

Naturally, all this becomes much more complicated as we move away from the simplest idealized examples.

For example, if the points are more evenly distributed in configuration space rather than belonging to cleanly-distinguishable clusters, then essentialist "X is a Y" cognitive algorithms perform less well, and we get Sorites paradox-like situations, where we know roughly what we mean by a word, but are confronted with real-world (not merely hypothetical) edge cases that we're not sure how to classify.

Or it might not be obvious which dimensions of Thingspace are most relevant.

Or there might be social or psychological forces anchoring word usages on identifiable Schelling points that are easy for different people to agree upon, even at the cost of some statistical "fit."

We could go on listing more such complications, where we seem to be faced with somewhat arbitrary choices about how to describe the world in language. But the fundamental thing is this: the map is not the territory. Arbitrariness in the map (what color should Texas be?) doesn't correspond to arbitrariness in the territory. Where the structure of human natural language doesn't fit the structure in reality—where we're not sure whether to say that a sufficiently small collection of sand "is a heap", because we don't know how to specify the positions of the individual grains of sand, or compute that the collection has a Standard Heap-ness Coefficient of 0.64—that's just a bug in our human power of vibratory telepathy. You can exploit the bug to confuse humans, but that doesn't change reality.

Sometimes we might wish that something to belonged to a category that it doesn't (with respect to the category boundaries that we would ordinarily use), so it's tempting to avert our attention from this painful reality with appeal-to-arbitrariness language-lawyering, selectively applying our philosophy-of-language skills to pretend that we can define a word any way we want with no consequences. ("I'm not late!—well, okay, we agree that I arrived half an hour after the scheduled start time, but whether I was late depends on how you choose to draw the category boundaries of 'late', which is subjective.")

For this reason it is said that knowing about philosophy of language can hurt people. Those who know that words don't have intrinsic definitions, but don't know (or have seemingly forgotten) about the three or six dozen optimality criteria governing the use of words, can easily fashion themselves a Fully General Counterargument against any claim of the form "X is a Y"—

Y doesn't unambiguously refer to the thing you're trying to point at. There's no Platonic essence of Y-ness: once we know any particular fact about X we want to know, there's no question left to ask. Clearly, you don't understand how words work, therefore I don't need to consider whether there are any non-ontologically-confused reasons for someone to say "X is a Y."

Isolated demands for rigor are great for winning arguments against humans who aren't as philosophically sophisticated as you, but the evolved systems of perception and language by which humans process and communicate information about reality, predate the Sequences. Every claim that X is a Y is an expression of cognitive work that cannot simply be dismissed just because most claimants doesn't know how they work. Platonic essences are just the limiting case as the overlap between clusters in Thingspace goes to zero.

You should never say, "The choice of word is arbitrary; therefore I can say whatever I want"—which amounts to, "The choice of category is arbitrary, therefore I can believe whatever I want." If the choice were really arbitrary, you would be satisfied with the choice being made arbitrarily: by flipping a coin, or calling a random number generator. (It doesn't matter which.) Whatever criterion your brain is using to decide which word or belief you want, is your non-arbitrary reason.

If what you want isn't currently true in reality, maybe there's some action you could take to make it become true. To search for that action, you're going to need accurate beliefs about what reality is currently like. To enlist the help of others in your planning, you're going to need precise terminology to communicate accurate beliefs about what reality is currently like. Even when—especially when—the current reality is inconvenient.

Even when it hurts.

(Oh, and if you're actually trying to optimize other people's models of the world, rather than the world itself—you could just lie, rather than playing clever category-gerrymandering mind games. It would be a lot simpler!)


Imagine that you've had a peculiar job in a peculiar factory for a long time. After many mind-numbing years of sorting bleggs and rubes all day and enduring being trolled by Susan the Senior Sorter and her evil sense of humor, you finally work up the courage to ask Bob the Big Boss for a promotion.

"Sure," Bob says. "Starting tomorrow, you're our new Vice President of Sorting!"

"Wow, this is amazing," you say. "I don't know what to ask first! What will my new responsibilities be?"

"Oh, your responsibilities will be the same: sort bleggs and rubes every Monday through Friday from 9 a.m. to 5 p.m."

You frown. "Okay. But Vice Presidents get paid a lot, right? What will my salary be?"

"Still $9.50 hourly wages, just like now."

You grimace. "O–kay. But Vice Presidents get more authority, right? Will I be someone's boss?"

"No, you'll still report to Susan, just like now."

You snort. "A Vice President, reporting to a mere Senior Sorter?"

"Oh, no," says Bob. "Susan is also getting promoted—to Senior Vice President of Sorting!"

You lose it. "Bob, this is bullshit. When you said I was getting promoted to Vice President, that created a bunch of probabilistic expectations in my mind: you made me anticipate getting new challenges, more money, and more authority, and then you reveal that you're just slapping an inflated title on the same old dead-end job. It's like handing me a blegg, and then saying that it's a rube that just happens to be blue, furry, and egg-shaped ... or telling me you have a dragon in your garage, except that it's an invisible, silent dragon that doesn't breathe. You may think you're being kind to me asking me to believe in an unfalsifiable promotion, but when you replace the symbol with the substance, it's actually just cruel. Stop fucking with my head! ... sir."

Bob looks offended. "This promotion isn't unfalsifiable," he says. "It says, 'Vice President of Sorting' right here on the employee roster. That's an sensory experience that you can make falsifiable predictions about. I'll even get you business cards that say, 'Vice President of Sorting.' That's another falsifiable prediction. Using language in a way you dislike is not lying. The propositions you claim false—about new job tasks, increased pay and authority—is not what the title is meant to convey, and this is known to everyone involved; it is not a secret."


Bob kind of has a point. It's tempting to argue that things like titles and names are part of the map, not the territory. Unless the name is written down. Or spoken aloud (instantiated in sound waves). Or thought about (instantiated in neurons). The map is part of the territory: insisting that the title isn't part of the "job" and therefore violates the maxim that meaningful beliefs must have testable consequences, doesn't quite work. Observing the title on the employee roster indeed tightly constrains your anticipated experience of the title on the business card. So, that's a non-gerrymandered, predictively useful category ... right? What is there for a rationalist to complain about?

To see the problem, we must turn to information theory.

Let's imagine that an abstract Job has four binary properties that can either be high or low—task complexity, pay, authority, and prestige of title—forming a four-dimensional Jobspace. Suppose that two-thirds of Jobs have {complexity: low, pay: low, authority: low, title: low} (which we'll write more briefly as [low, low, low, low]) and the remaining one-third have {complexity: high, pay: high, authority: high, title: high} (which we'll write as [high, high, high, high]).

Task variety and authority are hard to perceive outside of the company, and pay is only negotiated after an offer is made, so people deciding to seek a Job can only make decisions based the Job's title: but that's fine, because in the scenario described, you can infer any of the other properties from the title with certainty. Because the properties are either all low or all high, the joint entropy of title and any other property is going to have the same value as either of the individual property entropies, namely ⅔ log₂ 3/2 + ⅓ log₂ 3 ≈ 0.918 bits.

But since H(pay) = H(title) = H(pay, title), then the mutual information I(pay; title) has the same value, because I(pay; title) = H(pay) + H(title) − H(pay, title) by definition.

Then suppose a lot of companies get Bob's bright idea: half of the Jobs that used to occupy the point [low, low, low, low] in Jobspace, get their title coordinate changed to high. So now one-third of the Jobs are at [low, low, low, low], another third are at [low, low, low, high], and the remaining third are at [high, high, high, high]. What happens to the mutual information I(pay; title)?

I(pay; title) = H(pay) + H(title) − H(pay, title)
= (⅔ log 3/2 + ⅓ log 3) + (⅔ log 3/2 + ⅓ log 3) − 3(⅓ log 3)
= 4/3 log 3/2 + 2/3 log 3 − log 3 ≈ 0.2516 bits.

It went down! Bob and his analogues, having observed that employees and Job-seekers prefer Jobs with high-prestige titles, thought they were being benevolent by making more Jobs have the desired titles. And perhaps they have helped savvy employees who can arbitrage the gap between the new and old worlds by being able to put "Vice President" on their resumés when searching for a new Job.

But from the perspective of people who wanted to use titles as an easily-communicable correlate of the other features of a Job, all that's actually been accomplished is making language less useful.


In view of the preceding discussion, to "37 Ways That Words Can Be Wrong", we might wish to append, "38. Your definition draws a boundary around a cluster in an inappropriately 'thin' subspace of Thingspace that excludes relevant variables, resulting in fallacies of compression."

Miyamoto Musashi is quoted:

The primary thing when you take a sword in your hands is your intention to cut the enemy, whatever the means. Whenever you parry, hit, spring, strike or touch the enemy's cutting sword, you must cut the enemy in the same movement. It is essential to attain this. If you think only of hitting, springing, striking or touching the enemy, you will not be able actually to cut him.

Similarly, the primary thing when you take a word in your lips is your intention to reflect the territory, whatever the means. Whenever you categorize, label, name, define, or draw boundaries, you must cut through to the correct answer in the same movement. If you think only of categorizing, labeling, naming, defining, or drawing boundaries, you will not be able actually to reflect the territory.

Do not ask whether there's a rule of rationality saying that you shouldn't call dolphins fish. Ask whether dolphins are fish.

And if you speak overmuch of the Way you will not attain it.

(Thanks to Alicorn, Sarah Constantin, Ben Hoffman, Zvi Mowshowitz, Jessica Taylor, and Michael Vassar for feedback.)

Concerning Loyalty and Revenge

Retarget loyalty intuitions onto specific humans (never ideologies or collective identities). Retarget revenge intuitions onto patterns of incentives (never specific humans).

The Right to Life, Conjugated

She's a ward of the state; you have an inalienable right to live; I'm literally more useful alive rather than dead with respect to the values of powerful coalitions.

Concerning Motives for Cooperation

Always be peaceful and tell the truth to your friends because you love and trust them. Always be peaceful and tell the truth to cops, schoolteachers, psychiatrists, CPS agents, &c. because you're outgunned and bad at lying. Don't be confused about your reasons for doing things, even if you always end up doing the same thing.

Concerning Frame Control Via Salient Scenarios

"We need to institutionalize people in order to prevent them from hurting themselves" has the same memetic-superweapon structure as "We need to torture terrorists to get them to tell us where they've hidden the suitcase nuke." The scenario as stated obviously has consequentialist merit (death is worse than prison, megadeaths are worse than torture), so you'd have to be some kind of huge asshole—or a former suspected terrorist—to say, "I claim that this hypothetical scenario is not realized nearly as often as you seem to be implying and therefore falsifiably predict that many of your alleged real-world examples will fall apart on further examination."

Object vs. Meta Golden Rule

"I know it might seem like a lot to ask, but I wouldn't hesitate to do the same for you if our positions were reversed."

"I don't doubt that. But I can't help but notice that it would be easier for you to say it if the fact that they aren't reversed is—somehow—not a coincidence."

Means-Ends

Ayn Rand said that a Spanish proverb said that God said, "Take what you want, and pay for it."

But instructions from God would be redundant. Matter does not obey physical law out of fear of punishment or a sense of moral duty; what we call a "law" is a characterization of that which exists. So too with this.

An Intuition on the Bayes-Structural Justification for Free Speech Norms

We can metaphorically (but like, hopefully it's a good metaphor) think of speech as being the sum of a positive-sum information-conveying component and a zero-sum social-control/memetic-warfare component. Coalitions of agents that allow their members to convey information amongst themselves will tend to outcompete coalitions that don't, because it's better for the coalition to be able to use all of the information it has.

Therefore, if we want the human species to better approximate a coalition of agents who act in accordance with the game-theoretic Bayes-structure of the universe, we want social norms that reward or at least not-punish information-conveying speech (so that other members of the coalition can learn from it if it's useful to them, and otherwise ignore it).

It's tempting to think that we should want social norms that punish the social-control/memetic-warfare component of speech, thereby reducing internal conflict within the coalition and forcing people's speech to mostly consist of information. This might be a good idea if the rules for punishing the social-control/memetic-warfare component are very clear and specific (e.g., no personal insults during a discussion about something that's not the person you want to insult), but it's alarmingly easy to get this wrong: you think you can punish generalized hate speech without any negative consequences, but you probably won't notice when members of the coalition begin to slowly gerrymander the hate speech category boundary in the service of their own values. Whoops!

The Bayes-Structure in the Form of a Riddle

Left-wingers say torture is wrong because the victim will say whatever you want to hear.

Right-wingers say torture is right because the villain will tell the truth.

Q: What happens when you torture someone who only tells the truth?
A: They'll make noises in accordance with their personal trade-off between describing reality in clear language, and pain.

This is the whole of the Bayes-structure; the rest is commentary. Now go and study.

The Reason the World Sucks

"I think we should perform action A to optimize value V. The reason I think this is because of evidence X, Y, and Z, and prior information I."

"What?! Are you saying you think you're better than me?!"

"No, I don't think I'm better than you. But do I think I'm smarter than you in this particular domain? You're goddamned right I do!"

"You do think you're better than me! I guess I need to kill you now. Good thing I have this gun on me!"

"Wait, what?"

ka-BANG

Your Periodic Reminder I

Aumann's agreement theorem should not be naïvely misinterpreted to mean that humans should directly try to agree with each other. Your fellow rationalists are merely subsets of reality that may or may not exhibit interesting correlations with other subsets of reality; you don't need to "agree" with them any more than you need to "agree" with an encyclopædia, photograph, pinecone, or rock.

Late-Onset

the moment of liberating clarity when you resolve the tension between being a good person and the requirement to pretend to be stupid by deciding not to be a good person anymore ?

Bayesomasochism

Physical pain is the worst thing in the world, and the work of effective altruists will not be done until the last nociceptor falls silent and not a single moment of suffering remains to be computed across our entire future light cone.

But the emotional pain of discovering that your cherished belief is false, that everything you've ever cared about is not only utterly unattainable, but may in fact not even be coherent?—yeah, I'm pretty sadomasochistic about that. That's rationality; that's what it feels like to be alive.

The Roark–Quirrell Effect

Education increases altruism up to a point (as you increasingly understand that other people are real too and have moral value for the same reasons you do even if you don't experience it from the first person), until you accumulate so many seemingly unique insights that the entire rest of the world looks so abominably stupid that you no longer want to waste a single precious dollar or minute on the concerns of these creatures that can't even see the Really Obvious Thing.

(Or, maybe this is just a form of mental illness specific to high-psychoticism males that can be cured with the appropriate drugs. We'll find out!)

The Fundamental Theorem of Epistemology

$$P(H|E) = \frac{P(E|H)P(H)}{P(E)}$$

(more commonly known as Bayes's theorem, but I like my name better)

"What Can I Do For You?"

"I think we should set aside some time to discuss how I could provide more value to you."

"That's an awfully disingenuous way of proposing a negotiation; you're not that altruistic."

"The function of speech is to convey meaning to the listener. I speak of providing more value to you because that's all you should care about; if it happens that the means by which we arrange that I do so involves you providing more value to me, well, that's as irrelevant as it is obvious."

Identity

"You don't get to decide what I am! ... for the same reason that I don't get to decide what I am! 'What I am' is an empirical question to be settled by evidence and reasoning, the answer to which I can exert some limited control over in proportion to the strength of the self-modification techniques I have at my disposal!"