As a brash but anxious young programmer in the before-time, I affected a standoffish attitude on the topic of games. Why would I learn to play a game as a human, I said, when I could just write a program to do it? When an office chess game got started at my workplace, I wrote a chess engine (a standard minimax search with α–β pruning) rather than learn anything about strategy or tactics. When my insufficiently requited love started hosting a meetup to play Zendo—a deduction game in which one player makes up a rule to positively or negatively classify arrangements of plastic pyramids, while the others submit arrangements to be classified to help them guess the rule—I wrote a program to efficiently propose examples to deduce a true rule from dozens of candidates. (Given a distribution over hypotheses, it generated an arrangement whose positive or negative classification would eliminate half the remaining probability-mass.)

Now, in my old age (and the present end-time in which custom software is no longer such a flex), I no longer feel the need to pretend not to understand the appeal of games. I even learned chess properly back in 'twenty-four.

The one that's captured my interest lately is the arrow-stomping game—known by a couple other brand names in addition to the original, whose inessential differences of implementation accentuate rather than obscure its fundamental nature as the same game. (Except the South Korean one is trash because the timing windows and letter grades are too lenient.)

The arrow-stomping game is notable for being trivial from the perspective of my before-time impulse to write a program for it. When programming (say) chess, the implementation of the rules of the game is quite different from the implementation of the tree search agent to play the game. Moreover, a chess program is an exercise in logic: the program need only represent the abstract structure of the game; robotics to move physical pieces on a physical board would be considered superfluous if they were considered at all.

The arrow-stomping game is just the opposite. If you implemented the game as a sequence of timers (expect Left in 2.0 seconds, Up in 2.25 seconds, Right in 2.5 seconds ...), then the natural implementation of an agent to play the game would be—the same sequence of timers (input Left in 2.0 seconds, &c.). Arrow-stomping is only interesting as a robotics problem: the struggle of an embodied agent to convert visual cues to timing information and a motor plan to deliver a force to the appropriate panel or panels at the requested moment. It is thus particularly suited to the present end-time as an illustration of Moravec's paradox. The games that endure are sports, which appear cognitively trivial only because the form of cognition demanded is more ancient than language, with the effect that we can't talk about it and the pretraining set (for all the vastness of webtext) doesn't cover it.

A related way in which the pre-linguistic nature of arrow-stomping seems thematically suited to the present end-time has to do with how the process of learning it casts doubt on the dream of an intelligence shaped like Reason itself. In the before-time, people thought that the end-times would be brought about by "good old fashioned" engineering turned on itself: the code that writes the code that writes the code that ends the world.

To be clear, that's still plausible as a theoretical possibility for a later end-time, but it turned out to be unnecessary for the beginning of the end-times. For now, we don't know how to write the code that writes the code. We can only write code to train networks that write code. The kind of legible symbolic reasoning embodied in a chess engine is the output of a process that doesn't seem to natively have the shape of legible symbolic reasoning.

Likewise the solution to the robotics problem. Every arrow summons a judgment. In the original-brand implementation, they are Marvelous (±17 ms), Perfect (±34 ms), Great (±84 ms), Good (±124 ms), and Miss. (Apparently historical contingency forces us to suffer "Perfect" not being the best judgment. The American one at least has the decency to call the best judgment "Perfect!!" with two exclamation points, as distinguished from "Perfect".)

The judgments are a dense training signal. Beyond a few words of high-level advice (try to alternate your feet), one does not "study" arrow-stomping. In principle, we could imagine a mind inspecting a sequence of upcoming arrows and searching for a motor plan to fit them, like a chess engine calculating variations. As a human, there's no time for that. One can only let one's policy be shaped by the rewards of the judgments. (The American-brand implementation supplies gradient information by displaying "Early" or "Late" as part of the on-screen judgment; in the original, you have to toggle a config option.)

Arrows often come in patterns (like Left–Right–Left or Right–Up–Right), which can be "chunked" and loaded as one motor program, rather than read as individual notes. The chunking theory correctly predicts that losing the rhythm often results in correlated poor judgments on consecutive notes, in contrast to the independent "Gaussian" error on every step. A Great Fullcombo (completing a song with only Great or better judgments) isn't overwhelmingly harder than a Fullcombo (no Miss judgments): both imply having held the rhythm the whole way, with the Great Fullcombo "only" (only) requiring a tighter Gaussian variance.

The levels of the game provide a smooth progression. As the player masters simpler patterns, higher difficulty levels provide new ones. On level 6 of the original-brand implementation, notes spaced at closer intervals (blue arrows on odd eighths of a measure, contrasted to the red "quarter" notes) become common. Patterns like Left–Down–Right imply either a "crossover" putting one leg in front of the other (Left with left foot, Down with right foot, Right with left foot) or breaking the pattern of alternating feet ("double-stepping"). Novices tend to double-step, but as the arrows become denser at higher levels, crossovers become a practical necessity to keep up with the beat.

And yet I'm reconstructing all this knowledge in words after the fact. No one told me to chunk patterns. At no point did I "practice" crossovers (only realizing I had done them for sure after careful reading and video review). Competences emerge as solutions determined by the training signal. Good (late) tells you to step faster, Great (early), a little slower. A desperate reach to hit a panel in time gets reinforced when it works, until the combined force of thousands of gradient steps inscribe a dance routine which can only in retrospect be described in terms of "crossovers" or "double-stepping."

Ultimately, what makes the arrow-stomping game unusual isn't how impervious its robotics problem is to good old fashioned engineering, but how its rich training signal makes it permeable to anyone with enough time and quarters. I see online chess accounts with thousands of games over years and flat ratings. Apparently, mere experience isn't enough to solve the credit-assignment problem by which the difference between victory and loss might owe to positional weaknesses introduced in the opening. I also tried to learn gymnastics in 2024. It didn't take. Supposedly you can build up to a full handstand by starting against the wall, but even that wasn't enough of a gradient; I wasn't feeling out a clear path through parameter-space from kicking out off the wall for a few seconds to a clean, independent balance. When prodigies do make a better show of learning chess or gymnastics, it's with coaches and puzzles to smooth the way. The alternative to incrementally accreting complex skills from local rewards isn't designing them from clean principles—but not learning them at all.

Apprehension of this strange inversion of reasoning almost makes the present end-time seem like less of a grotesque surprise. The steps we were already taking implied the seeds of the revolution.