Saturday, August 15, 2026

More On Physical AI: "How world models became AI's next frontier"

From The Deep View, July 3:

Large language models can do a lot of things, as long as those things exist on a screen.

Having consumed practically all of the data on the internet, these models can tell you something about practically anything. They can analyze thousands of documents and pull out the most important parts. They can plan your trips, write poems, and mimic therapists in times of need. They can help researchers understand anything from protein structures to ancient history. And they’re the backbone of millions of agents that are the hottest ticket in tech right now.

But the world is bigger than a screen, and understanding it requires more than words.

That's why Nvidia’s Jensen Huang has said more than once that physical AI is due for its "ChatGPT moment."

It’s also the reason that world models, or AI models capable of understanding the physical environment, have gained significant momentum in 2026. Along with a flock of young startups entering the space, some of AI’s most prominent figures have homed in on the concept.

But as the momentum around world models grows, in tandem with physical AI and robotics, some leaders have begun to question their impact on LLMs, and ultimately, the ever-elusive path to artificial general intelligence (AGI).

"Seeing the world in a profound way, in a way that you participate in your movement, your interaction, in your communication, is critical for intelligence," said Dr. Fei-Fei Li, largely considered the godmother of AI, while also being the founder and CEO of World Labs, in a panel at the HumanX conference in April. "Not having that is intelligence in the dark."

Investors take notice

As world models catch the attention of some of AI’s biggest tastemakers, investors have become captivated. World model startups have been raking in billions in funding, some at incredibly early stages:

Beyond startups, several companies have pivoted into world models from one industry in particular: gaming.

Niantic, the maker of the beloved Pokémon Go app, sold its suite of mobile games to Scopely for $3.5 billion and spun out a lab last March called Niantic Spatial, focused on developing what it calls a "Large Geospatial Model," a world model that enhances spatial reasoning in LLMs.

And Roblox, the online gaming platform with more than 150 million daily active users, is developing its own version of a world model that it calls "real-time dreaming," that allows creators to generate and iterate on virtual environments through language prompts. In a panel at the HumanX conference in April, David Baszucki, Roblox founder and CEO, said that he envisions the company’s world models "not just as a play technology, but as a creation technology as well."

In January, Google DeepMind released Project Genie, an "experimental research prototype" that marks the latest iteration of its work in the world model space. The project is powered by its flagship Gemini model, its Nano Banana Pro image model, and Genie 3, its most powerful world model yet.

Nvidia, meanwhile, unveiled its Cosmos model at CES 2025, a world foundation model that’s aimed at accelerating the development and deployment of autonomous vehicles and robots. Since then, the company has expanded Cosmos further, including debuting the third generation of the Cosmos family and releasing world-generation models, controllable simulations for synthetic data generation, and multimodal reasoning models for physical AI.

Many bets, same problems....

....MUCH MORE 

August 14 - "World Models Are AI’s Next Frontier"