← Back to blog index

What Platonic Representations Tell Us About the Race to AGI

There's an idea floating around AI research called the platonic representation hypothesis. The claim is simple but strange: as neural networks get bigger and better, whether they're trained on images, text, or something else entirely, they start converging on the same internal way of organizing the world. A vision model and a language model, given enough scale, begin measuring the "distance" between concepts in eerily similar ways.

Concretely: if you ask an image model "is a photo of a dog more similar to a photo of a wolf or a photo of a cat," and you ask a language model the same question using only text descriptions, they increasingly agree. The bigger and more capable the models get, the more their internal "sense of distance" between things lines up.

The name is a nod to Plato's cave. People chained in a cave only see shadows on a wall, cast by real objects behind them. They mistake the shadows for reality itself. Different senses (touch, sound, sight) are like different shadows of one underlying world. The theory suggests something similar is happening in AI. A model trained purely on text and a model trained purely on pixels are both staring at shadows of the same reality. Train them hard enough, and they start reconstructing the shape of the object casting those shadows.

This connects to something Greg Brockman said recently. In a podcast with Alex Kantrowitz, he argued that the debate over whether pure text models can reach AGI is settled, in his view they will get there. OpenAI has been narrowing its bets accordingly, folding video and world-model research (Sora) into a smaller robotics-focused effort rather than running it as a parallel consumer product, because splitting compute across multiple "branches of the tech tree" at once is too costly.

Put the two ideas together and you get a compelling picture. If different modalities are all shadows of one underlying statistical structure of reality, then it may not matter that much which door you walk through first, text, images, video, or robotics. Anthropic and OpenAI pushing hard on language, Google and World Labs pushing on multimodal and video, robotics labs building embodied models: these could all be different paths up the same mountain. The paths will differ in length. Robotics likely takes longer simply because real-world data is expensive to collect and the field only got serious momentum after the ChatGPT moment created enough capital and urgency to fund it properly. But if the platonic representation hypothesis holds, the destination each path is climbing toward may be the same one.

What Platonic Representations Tell Us About the Race to AGI – Mohammed Arham