Skip to content
← Blog

You cannot crawl the physical world

Most arguments about a singularity are arguments about software improving software. A model writes a better training loop, the better loop produces a better model, the loop tightens, and somewhere in the tightening the curve leaves the page.

A survey of 1,250 arXiv papers from 2024 to 2026 separates two things that usually get discussed as one. Bounded self-refinement is convergent, evaluable, and already industrial practice. Open-ended recursive self-improvement is the thing the arguments are actually about, and the survey reports it as still bounded by three constraints: grounding requirements, collapse dynamics, and compute constraints.

Compute gets nearly all the attention. There are whole papers on whether compute bottlenecks alone can stop an intelligence explosion. Grounding gets far less, and it is the one worth sitting with, because it is the constraint that does not yield to spending.

Grounding runs through bodies

That last step is an inference rather than a finding. The survey names grounding without naming embodiment. But it is difficult to say what grounding is, for a system meant to act rather than only describe, without arriving at contact with the world.

And there the literature is blunt about the asymmetry. Work on the embodiment gap in robot foundation models puts it plainly: language and vision draw on internet-scale training data, whereas robotics cannot obtain observations paired with action commands at anything like that scale. Text was already written and already public. Images were already taken and already uploaded. Nobody has ever recorded a joint-torque trajectory as a side effect of doing something else.

The same work argues that useful robot data accumulates only after learning and engineering are combined well enough that robots operate reliably in the first place, which makes the bottleneck partly circular. And it names what does not transfer even when data is pooled: calibration, control alignment, contact correction, and safety mechanisms remain per-robot work.

So the physical corpus does not exist yet, cannot be scraped into existence, and has to be produced by machines operating in places, one episode at a time.

Why this changes the shape of the problem

A web crawl is a capital problem. One organisation, one budget, one datacentre, and the corpus is yours. This is the shape the last decade rewarded, and it is why frontier capability concentrated the way it did.

Physical interaction data is not that shape. It is a distribution problem. Ten thousand hours of contact-rich manipulation in ten thousand different rooms cannot be bought from a single vendor, because the thing being bought is presence in ten thousand rooms. Capital helps. It does not substitute.

This is the point where decentralisation stops being a preference and starts being the matching topology. Not because distributed systems are morally better, but because the input is already distributed and the collector has to be shaped like the thing collected.

Where this network comes in

The codex's answer to that is structural rather than aspirational, and it predates this argument.

The hourglass makes embodiment one of seven surfaces sharing a single learned latent, so an embodied episode contributed by one robot is not stranded in a robot-specific model. P2 puts the consequence in constitutional terms: representations are owned by the network, not by any model that uses them, and a checkpoint producing an incompatible latent is a fork rather than an update. The embodiment overview treats the surface as separate precisely because its constraints differ, with control loops at 20 to 100 Hz and continuous multimodal action distributions that a token softmax parameterises badly.

The collection mechanism is specified too. Phase 4 activates the manipulation and locomotion surfaces, and its scope includes a multi-embodiment registry, an embodied data contribution pipeline, and a per-episode reward. Which is to say: a way to pay for exactly the input that cannot be scraped, to whoever produced it, wherever they are.

That is the argument for why this design and this problem fit. It is not an argument that the problem is solved.

The objection worth answering

A technical governance analysis of distributed and decentralised training states the standard objection cleanly: centralised compute is attractive to governance precisely because it is detectable and concentrated, and moving to independent participants introduces verification difficulty that a single operator never faces.

The codex answers it in the design rather than in a rebuttal. P3 makes verification statistical and refuses to deploy any surface that cannot be checked by sampling. P7 requires every validator to run an identical binary on identical inputs and produce an attestable output, treating divergence as the signal and convergence as the substrate. The verification burden was priced in from the first page, and the validators chapter is where it is paid.

The frame the codex declines

Worth ending on the thing that is absent.

The word singularity does not appear anywhere in the codex. Neither does intelligence explosion, nor takeoff, nor recursive self-improvement. For a document this long about building a general intelligence, that is a choice.

What it says instead, at the end of the roadmap, is that after Phase 6 the network is in steady state, new surfaces and embodiments may be added at any time through the standard addition protocol, and the system has no terminal state by design.

No terminal state. Not an event to be first to, and not a finish line. Infrastructure that is supposed to outlast the people who started it, which is the same commitment P11 makes about operators and the roadmap makes about itself.

If a singularity happens, it will happen to a world that still has to move objects around in rooms. The question worth asking is not who gets there first. It is who owns the thing when it arrives.