A Blog by Jonathan Low

 

Aug 22, 2026

AI's next Big Leap Is Into Real World Robotics

The goal is not superintelligence or other cult-like musings, but training for practical applications that can actually produce and operate in real world settings in order to improve performance. 

Considered low-brow by some, many investors see this as the gateway to operating efficiencies and, ultimately, the real pay-off for AI. JL

Christopher Mims reports in the Wall Street Journal:
AI models are being trained on videogames and simulations, rather than literature, code and images. They have reached a point where they can navigate in three-dimensional space and even manipulate objects autonomously with just a set of instructions. World-model startups want to create a control system that’s as versatile when piloting robots as today’s LLMs are when crafting text, from sonnets to software. The bet is a model can pilot a robot in the real world as if it were a character in a videogame. One caveat: It has to send a live video feed to a user, and can be controlled with a videogame controller, or mouse and keyboard. This is typical criteria for mainstream robots, but it excludes most two-legged humanoid robots.

Today’s AI large language models might excel at pushing a pencil around, but proponents say doing real physical work requires large action models, aka world models.

In this approach intended for piloting robots, artificial-intelligence models are trained on videogames and simulations, rather than literature, code and images. They have recently reached a point where, in some circumstances, they can navigate in three-dimensional space and even manipulate objects autonomously with just a simple set of instructions.

It’s still early days for this tech, says Moritz Baier-Lentz, an investor in several companies in this area. “If we’re comparing this to large language models, this is like GPT-2,” he adds. (That’s the model OpenAI released way back in 2019.)

Despite the relatively primitive state of world models, tech luminaries are piling in. The bet is that a fundamentally new architecture has the potential to take AI to places that today’s LLM-based AIs can’t venture. And one day, pioneers hope to merge these two schools of artificial thought.

From words to worlds

For about half a billion years, animals have been evolving brains capable of modeling the world around them, and using that mental map to plan their next action. Language, on the other hand, dates back only about a hundred millennia, a shorthand humans created to process and relay concepts great and small.

“Text is just a lossy representation of the real world,” says Kent Rollins, chief product officer at General Intuition, and former director of the Fortnite ecosystem at Epic Games. “The world existed long before we had text, and using it to describe the world is going to essentially miss core aspects.” 

Roboticists have long known this. The control systems of today’s more sophisticated robots rely on physics-based simulations of the physical realm. These, too, are world models, though they are painstakingly coded and highly specialized. What works for one kind of robot doesn’t work on another.

Today’s world-model startups want to create a control system that’s as versatile when piloting robots as today’s LLMs are when crafting text, from sonnets to software.

New York-based world-model startup General Intuition is now finalizing a round of investment that would value it at more than $6 billion, making it the most valuable AI lab of this kind. The basis of its training is videogames: The company’s model captures every frame of a game, as well as the actions users take—every press on a button or nudge of a control stick. Unlike robots trained on video alone, this allows its AI to connect user actions to their consequences in virtual worlds, yielding a large action model. 

With a little fine-tuning, a model can pilot a robot in the real world as if it were a character in a videogame. One caveat: It has to be a quadruped, wheeled vehicle or flying drone that would normally send a live video feed to a user, and can be controlled with a videogame controller, or mouse and keyboard. This is fairly typical criteria for many mainstream robots, but it excludes most two-legged humanoid robots.

General Intuition’s model is trained on millions of hours of videogame play from real humans, gathered from its sister service Medal.tv, a platform for capturing and sharing gaming clips. 

In and around their offices in New York and Geneva, General Intuition shows off a robot dog. While a typical AI for driving such a robot might require enormous amounts of training, this one just requires a few minutes of fine-tuning. The system needs to know where it is, what kind of body it’s in and what its goal is—then away it goes, says Baier-Lentz, who is an investor in General Intuition.

Previous demos of world models, like the Genie models unveiled by Google DeepMind, focused on generating new worlds to train robots. This next generation perceives the world and decides what to do next. Jack Parker-Holder, who previously led those world-model efforts at Google, co-founded London-based Emulate. The fledgling lab is in talks to raise more than $500 million from investors, and its tech talent includes a half-dozen other former Googlers, according to documents reviewed by The Wall Street Journal.

Language + action = ?

Other world-model companies are reluctant to share details about the AIs they are building, but their acquisitions and the publications of their engineers give some hints. 

World Labs, the startup headed by Stanford professor and Google veteran—and “godmother of AI”—Fei-Fei Li, recently acquired robotics company Scenix. AMI Labs, headed by Meta’s former chief AI scientist Yann LeCun, appears to be working on something broader, exploring multiple architectures outside of traditional large language models. (Both companies declined to comment for this article.)

While General Intuition and other startups emphasize their capabilities in robotics, potential investors and business partners are asking another question: How can world models enhance the abilities of today’s language-based models?

Adam Jelley, Pim de Witte, Vincent Michelli, and James Swingos at General Intuition's offices in Geneva.
The team at General Intuition’s offices in Geneva, from left: Adam Jelley, CEO Pim de Witte, Vincent Micheli and James Swingos.

For LLMs to process images and audio, they have to be trained on that media directly, turning the inputs into tokens as they do with words. Trying to process a three-dimensional scene in this way has proven to be massively inefficient, yielding models that are too slow to direct a robot in most situations.

But world models come with their own set of issues: It was relatively easy to evolve ChatGPT from one generation to the next precisely because it started out only playing with text. As LLMs grow, and are crammed with more and more data, they get bigger and smarter, but they continue to make mistakes, says George Konidaris, a professor of robotics at Brown University, and a co-founder of Realtime Robotics, which builds systems for industrial robots.

With robots, there are far fewer scenarios where we would accept hallucinations and other oopsies. 

“The real world is very complex, and has a lot of special cases and hard edges, and you can’t approximately miss something,” says Konidaris. “If you hit something while moving your robot, everything changes.”

While some think increasing efficiencies in LLMs might make them enough to become the AI that powers robotics and other 3-D applications, a hybrid concept also exists: LLMs could call on world models, the way they call on other software to help them do their jobs.

In the meantime, both General Intuition and Emulate are conspicuously not based in the Bay Area, which has become enthralled by the potential of LLM-based AIs to lead to “superintelligence.”

When I ask if superintelligence is the goal of his company, General Intuition’s CEO, Pim de Witte, demurs, “I’m in New York because I want to stay away from all the cult-like behavior, and just focus on scientifically evaluating capabilities of models in robots,” he says. “We don’t need to make it more than what it is.”

0 comments:

Post a Comment