World Modeling, Explained: AI's Post-Language Bet in Numbers

Rachid Idali

by Rachid Idali

Two years ago, 880 people a month searched for "world modeling" in the US. As of July 2026 it is 6,600, up 247% in a year. That is the part every explainer gets right: a phrase from a 1943 psychology book and a 2018 machine learning paper has become one of the loudest ideas in AI.

Here is the part nobody has. The curve has two peaks of exactly the same height, in November 2025 and March 2026, and it has been drifting down since. Over the last three months the term is down 19% while it is still up 247% over the year. World modeling is no longer an emerging idea in search. It is an established one that moves when somebody ships something and sags when nobody does. That distinction matters if you are deciding whether to build on it.

Key takeaways:

  1. "world modeling" grew from 880 to 6,600 US monthly searches in two years, up 247% year over year, classified as EXPONENTIAL in the Rising Trends database (data as of July 2026).
  2. The peak is 9,900, hit twice: November 2025 and March 2026. The three-month growth column is negative, at -19%.
  3. A world model is an AI system that keeps an internal representation of an environment and predicts how it changes in response to action. World Labs splits the term into three jobs: renderers, simulators and planners.
  4. The vocabulary is rotating. Over the last year, world-model terms are up between 247% and 1,015% while llms is down 18%, agi down 33% and reasoning models down 63%.
  5. The catalysts are dated and public: Genie 3 in 2025, Fei-Fei Li's essay on November 10, 2025, Marble two days later, and World Labs' Atlas on September 1, 2026.
  6. The field cannot agree on a definition. A 58-page arXiv paper from July 2026 says outright that there is no consensus on what a world model fundamentally is.

Let's get into it.

The curve

Five years of monthly search volume, from our database. For most of that window "world modeling" was an academic phrase running under a thousand searches a month.

Line chart of monthly Google search volume for world modeling over 60 months, from 590 in Aug 2021 to 6,600 in Jul 2026, peak 9,900 in Nov 2025

Source: Rising Trends database, data as of Jul 2026

The breakout month is August 2025, when the term first crossed a quarter of its eventual peak. Then two spikes to 9,900, in November 2025 and March 2026, with a trough between them. July 2026 came in at 6,600. A term that doubles and holds is adoption. A term that spikes twice to the same ceiling is a news cycle with a floor underneath it, and the floor here is roughly seven times where the term sat in 2024.

Put it next to the words it is competing with and the scale problem is obvious.

Horizontal bar chart: Related searches around "world modeling". agi 90,500, llms 90,500, sora 2 49,500, world modeling 6,600, genie 3 2,900, spatial intelligence 1,900, world models ai 1,300, reasoning models 590

Source: Rising Trends database, data as of Jul 2026

agi and llms both draw 90,500 a month, roughly fourteen times world modeling. The new idea is not close to replacing the old vocabulary in public attention. What it is doing is taking the growth.

Horizontal bar chart: The AI vocabulary is rotating: searches over the last year. genie 3 +1,015%, world models ai +306%, world modeling +247%, llms -18%, agi -33%, spatial intelligence -57%, reasoning models -63%

Source: Rising Trends database, data as of Jul 2026

Every world-model term in our data is up over the year. Every term describing the previous generation is down. That is what a vocabulary change looks like before it finishes: the incumbent words are still bigger and already shrinking.

What a world model actually is

The short version: a world model is an AI system that holds a working representation of an environment and predicts how that environment changes in response to action. Where a language model learns the statistical structure of text, a world model learns the statistical structure of space and time.

The idea is older than the hardware. World Labs traces the phrase to Kenneth Craik, who proposed in 1943 that minds reason by running "small-scale models" of reality. The modern reference is a 2018 paper by David Ha and Jurgen Schmidhuber, published under the subtitle "Can agents learn inside of their own dreams?", which showed an agent training inside its own learned simulation of a game.

The useful modern definition is functional. In A Functional Taxonomy of World Models, World Labs splits the term by what the system outputs:

TypeOutputsWhat it is judged onExamples
RendererPixels for human eyesVisual fidelityText-to-video models, Genie 3
SimulatorState: geometry, physics, dynamicsStructural accuracyMarble, physics engines, digital twins
PlannerActionsWhether the robot succeedsVision-language-action models

The distinction is not pedantry. A renderer can produce a flawless drone shot of a city whose buildings collapse the moment you try to drive through them. World Labs' own line on this is blunt: renderers "cannot be trusted to design a building or train a robot."

Not everyone accepts that taxonomy. In July 2026 a team at Shanghai AI Laboratory posted a 58-page paper arguing the functional split classifies outputs rather than the underlying representation.

arXiv page for the paper A Definition and Roadmap for World Models, submitted 7 July 2026, showing the abstract stating there is no consensus on what a world model fundamentally is

Source: arxiv.org, captured 2026-09-11

The abstract of A Definition and Roadmap for World Models is the most honest sentence written about this category all year: researchers across AI subfields are building systems they call world models, "yet there is no consensus on what a world model fundamentally is, what it should predict, or how it should be built."

Why it broke out when it did

The dates line up with our curve almost exactly.

Google DeepMind put a working demo in front of the public with Genie 3, a model that generates environments you can walk around in real time at 24 frames per second and 720p, staying consistent for a few minutes. That is the month our data marks as the breakout.

Then November 2025, our peak month, delivered two events two days apart. On November 10, Fei-Fei Li published From Words to Worlds, arguing that spatial intelligence is AI's next frontier and world models are the road to it. On November 12, her company made Marble generally available: a model that turns text, an image, a video or a rough 3D layout into an explorable 3D world and exports it as Gaussian splats, meshes or video.

The people who had been working on this for decades noticed the sudden attention. Jurgen Schmidhuber, co-author of the 2018 paper, posted a history thread in February 2026 under the heading "World Model Boom."

His framing is the useful one: the concept "dates back millennia," and what changed in 2025 was not the idea but the compute and the demos. The post has 474 likes as of September 2026.

Who is building it, and what they shipped

World Labs is the pure play. On September 1, 2026 it introduced Atlas, an omni model pretrained from scratch on text, images, video and 3D. It generates up to one minute of video at 1440p with explicit camera control, reconstructs scenes from one to dozens of images, and supports real-to-sim workflows for robotics. It is early access by request and will power future versions of Marble.

Google DeepMind owns the renderer end with the Genie line, and the search data agrees: genie 3 is up 1,015% year over year, the fastest mover in the family, from a base of 2,900 a month as of July 2026.

NVIDIA is selling the substrate rather than the model. World Labs cites the company's own estimate that Omniverse targets more than a trillion dollars of addressable market across factories, warehouses, supply chains and digital twins.

The robotics buyers are the demand side. The agent platforms we track hit a wall that world models are meant to solve: an agent that operates software can learn from logs, but an agent that operates a gripper needs a world to practise in.

What changes if it works

Simulation stops being a capital expense. Today a faithful simulation of a warehouse or a hospital corridor is built by hand. If a model can generate one that holds up under physics, the cost of testing a robot, a building or an evacuation plan falls by orders of magnitude. Stanford HAI's brief opens with exactly that scenario, a wildfire incident commander who needs roads, grid and hospital capacity in one continuously updated picture.

Training data moves from the web to the fleet. The scarcest input is not video. It is action-labeled interaction data: robot trajectories and fleet logs that cannot be scraped. Whoever owns fleets owns the input, which is a very different competitive shape from the generative AI market where the training corpus was public.

The regulation gap is real and named. In July 2026 Stanford HAI published the first governance agenda for world models. Its sharpest finding: "No existing benchmark gives policymakers an adequate basis to evaluate a world model for safety-critical deployment." When a simulated environment stands in for the real one, the question regulators have to answer is whether it matches reality closely enough to train on, and nobody can measure that yet.

The skeptic's case

The most credible criticism comes from inside the field, not outside it.

World Labs, which sells world models, says this about the planner category: "Almost all have been confined to heavily constrained laboratory setups, with narrow object sets and short task horizons. None have been validated at the complexity, variability, or duration that real-world deployment demands."

Yann LeCun's long-running objection is more fundamental. He argues that pixel-level reconstruction is the wrong target entirely, and that models should predict in latent space instead. The Shanghai AI Laboratory paper engages with it directly and concedes the point partly: precise pixel reconstruction should not be the ultimate goal.

Then there are the unglamorous blockers the same paper lists: data asymmetry, compounding prediction errors over long rollouts, the sim-to-real gap, and the absence of benchmarks. Generated geometry can look correct while containing self-intersections or wrong scale that produce nonsensical physics. That is a specific, checkable failure mode, and it is why the search curve stalling at 9,900 twice is not surprising.

What to watch

Whether the floor rises. If world modeling clears 9,900 on a third spike and holds above it, the term has graduated from news cycle to category. If it keeps oscillating below, the idea is real but the audience is still researchers and reporters. The live world modeling trend page is where that shows up first.

A benchmark anybody trusts. Stanford HAI named the gap; the first credible measurement standard for simulation fidelity would be the moment procurement can start. Watch for it from a standards body rather than a vendor.

Robots, not videos. The renderer side is commercially mature and the planner side is not. The tell will be a planner shipped into an unstructured environment with published failure rates, not another demo reel. Until then, treat the category the way you would any frontier claim: impressive capability, unfinished infrastructure.

The honest read of this curve is that world modeling is a real idea in an early market, wearing the search profile of a story rather than a product. The words are rotating faster than the machines are.


Want to track a vocabulary shift while it is still small? Read our guide on how to identify market trends, follow the live world modeling dashboard page, or see what is breaking out right now on the Rising Trends dashboard.

Unlock More Trends & Insights

Never miss a trend again!

Thousands of Emerging Trends

Thousands of Breakout Apps

Mega Trends

Trend Analysis Tool

Get Access Now
Join +5,000 happy users
Limited time: 50% off annual with code

Written By

Rachid Idali

Founder of Rising Trends, helping entrepreneurs identify and capitalize on emerging market opportunities through expert trend analysis and insights.

World Modeling, Explained: AI's Post-Language Bet in Numbers