Introducing Live Models:A Frontier for World Models

Team of Visko

This video is generated by Visko Orbis.

← Research

For decades, we have built AI as passive Query Models: take an input, return a prediction, stop. That single paradigm gave rise to nearly everything we call modern AI, from perception 1 and media generation 2 to large language models 3 and coding agents 4. All of it rests on the same quiet assumption, that the world will hold still while the model computes.

The world does not hold still. Biological intelligence never have that luxury. The visual system evolved to turn a continuous stream of photons into meaning fast enough to survive, because an answer that arrives after the moment has passed is often no better than a wrong one. A static model, queried one turn at a time, sees this living world as disconnected snapshots, with no sense of time, no anticipation, no continuity a creature needs.

Today we are introducing Live Models, a new class of foundation models for continuous, real-time world intelligence. A Live Model is not queried once for a fixed answer. It runs. It holds a constantly refreshed understanding of its environment and of the hidden state beneath the surface. We mean “live” in three senses: 1) It generates in real time, under the deadline of a world in motion; 2) it interacts continuously, coupled to its environment and to the people in it; 3) it has an ongoing existence of its own (its “life”), a persistent process that keeps going whether or not anyone is watching.

We believe liveness is where general intelligence is headed, the enabler of the next generation of interactive systems across digital society, entertainment, robotics, and science. Our thesis, in one sentence: the next paradigm of AI is not a model you call but a process that runs, in step with the world itself.

Query Models

The modern era of deep learning began with perception. In 2012, a convolutional network trained on ImageNet cut recognition error dramatically 5, and recognition models grew steadily more generic until the emergence of Transformer 6, an architecture designed for language, was turned back onto images in Vision Transformer 7. Generation ran the same machinery in reverse. GANs 8 and diffusion models 2 learned to conjure images from noise, giving us DALL·E 2 9 and, in video, models such as Sora 10, Veo 11, and Seedance 12 that begin to capture physical dynamics. Language modeling traveled a parallel path, from BERT 13 predicting masked tokens to the GPT family 3 distilling everything into a single task, i.e., next-token prediction.

For all their differences, these systems share a single move: condition on an input, predict the rest, finish within one call. We call them Query Models. This works beautifully when the input can be treated as fixed, and a document does not change while the model reads it. The physical world does. It keeps moving while the model computes, observations arrive on their own schedule, and every action changes what will be observed next. Repeated queries rebuild the world from scratch on every call, yet they never sustain an understanding that evolves with the world itself.

World Models

A world model asks a deeper question: how does the world itself behave? The intuition is old. Kenneth Craik proposed in 1943 that a mind reasons by running a “small-scale model” of reality 14. Modern world models learn that model from data 15. As the functional taxonomy from Dr. Fei-Fei Li and the World Labs team makes clear 16,17, today’s world model systems can be divided by what they output. For example, renderers produce observations 18, simulators predict faithful states 19,20, and planners generate actions 21. The underlying approaches vary as well. Joint-Embedding Predictive Architectures (JEPA), for instance, learn a model of the world by predicting in an abstract representation space rather than in pixels 22. Beneath all these approaches lies a single aspiration: to learn the structure of space, physics, and dynamics itself. This is among the most important directions in physical AI today, and it is the foundation on which our work rests.

Our conviction is that world generation leads to a better understanding of the world. Recent work from Dr. Kaiming He and colleagues points the same way, showing that models trained simply to generate images and videos become superb world perceivers 23,24.

However, to be deployed into the world, e.g., to drive a car, guide a robot, hold a conversation as it happens, demands three things simultaneously: keeping pace with physical time, interacting in both directions at once, and persisting rather than being rebuilt on each request. None of these demands is new. Filtering and model-predictive control have met them for decades 25, wherever the dynamics were simple enough to write down by hand. What is new is meeting them in worlds that must be learned, at foundation-model scale.

Live Models: The Frontier of World Models

We introduce Live Models, a new frontier of world models that runs a world simulator in real time, maintaining an evolving belief about the world and carrying that belief forward in the world’s own time, for as long as it runs. This concept of “live” is built upon six functional pillars.

Figure 1

Figure 1. Six functional pillars of Live Models. (Created by Team of Visko)

1. Clocked: anchored to real time. Physical time does not pause for a computation to finish. A Live Model takes the real-world clock as its primary reference. Its understanding advances whether observations stream steadily, stutter, or vanish behind an occlusion, and it registers that time has passed even when nothing arrives. Delay is never neutral, because the world keeps moving through it, and a clocked model accounts for where things have gone since they were last seen.

2. Stateful: a persistent memory of the world. A Live Model carries one persistent internal state, a compressed and continually refreshed representation of its world that holds the permanence of objects, the momentum of motion, and the thread of what is happening. Because memory lives inside the model rather than in a growing transcript, the past is kept at roughly constant cost. That is what lets the model anticipate rather than merely react.

3. Interactive: a two-way loop with the world. A Live Model runs with its channels open. People, agents, and the environment can intervene in the process while it is underway, and the model folds them into what it is already doing instead of stopping and starting over. Its own actions matter in turn. Every move it makes changes the world, and the changed world is what it observes next, so influence keeps flowing in both directions. This is what makes a Live Model less an oracle you consult than a participant you act alongside.

4. Self-evolving: learning as a way of life. A Live Model continues learning after deployment. The model continually predicts what will happen next, compares its prediction with what actually happens, and adjusts itself based on the difference. On short timescales, this is test-time training 26, and it lets the model adapt to the specific scene, speaker, or machine in front of it. On long timescales this can be continual learning 27, and it means experience accumulates in the weights, so a model that has been running for a year can end up better than the one that shipped. The longer it lives, the more it knows.

5. Physical: grounded in the dynamics of the real world. A Live Model’s subject is the physical world itself: objects, forces, continuous cause and effect. From continuous streams of video, audio, and sensor data, it learns the regularities that govern how things move, how a cup tips and falls, and how traffic merges. Having absorbed those dynamics, it can run them forward, anticipating how a situation will unfold and how a different action would change the outcome. It can answer what happens next and what would happen instead. That is the difference between describing a world and possessing a working model of one.

6. Full-duplex: perceiving and generating at once. A Live Model does not take turns. It reads its input streams and produces its output streams at the same time, in the same network, so it keeps perceiving while it speaks or acts and can adjust its output as new observations arrive. This removes the read-then-respond cycle of turn-based systems, and with it the external machinery, such as voice activity detection, that determines when a turn has ended. Thinking Machines Lab’s recent interaction models are built on the same principle 28. With full-duplex in place, perception, memory, learning, and action run as one continuous computation.

We do not claim these pillars as ours alone. Pieces of liveness are already alive across the field, in real-time interactive world models 18, full-duplex speech systems 29, and streaming perception, and this essay stands on that shared ground. What we add is the synthesis: all six pillars held at once, in one foundation model, against the physical clock.

An Interpretation in Physics

Figure 2

Figure 2. A Live Model as a learned dynamical system. From a multimodal initial condition, the model integrates its world forward in step with real time, shaped by boundary conditions, bent by interactions, and corrected by each arriving observation. (Created by Team of Visko)

Stepping back, we offer a reading in the language of physics: a Live Model is a learned dynamical system (Figure 2). Physics describes a changing world not with a table of answers but with a law of motion, a rule for how one instant becomes the next, from Newton’s mechanics to Maxwell’s fields. A Live Model learns one, in the form of a neural stochastic differential equation (Neural SDE) 30: a learned operator carries the world’s latent state forward in continuous time, while a stochastic term absorbs what the law cannot fix in advance. The model’s state is therefore not a point estimate but a belief over possible worlds. This is what makes a Live Model a runnable world rather than a description of one. From an initial condition, a multimodal seed of text, image, sound, and physical state, it integrates the world forward, bent by interactions along the way.

One more thing worth noting is that, between observations, the model holds many possible worlds at once. An arriving observation acts like a measurement, collapsing them onto the worlds consistent with what was seen. It is fitting that “observation” is the word both Partially Observable Markov Decision Processes (POMDPs) 31 and quantum physics reach for, though the collapse here is Bayesian in spirit rather than identical to Bayesian conditioning in its mathematics. We make no claim about whether the universe is such a process. We only observe that a Live Model is one. To a process that lives inside time, holds a picture of its world, and is changed by everything it observes, that world is real. Building such processes is what we mean by live.

A Reference Architecture for Live Models

Figure 3

Figure 3. A reference architecture for Live Models. An interface of perception encoder, interaction channel, and rendering decoder faces the world, while the internal models hold perception, memory, and physics tokens in one latent space, paced by a world clock and rolled forward by latent autoregressive generation. (Created by Team of Visko)

How would one build a Live Model? Our reference architecture, sketched in Figure 3, joins an interface that faces the world with internal models that carry the world within, in a single loop that runs continuously along the clock. What passes between them is a common currency of tokens, compact carriers of meaning that let perception, memory, physics, and time share one latent space.

  • Perception Encoder. The model’s standing observer. It distills the raw, asynchronous streams arriving from the environment (video, audio, language, sensor signals) into perception tokens, converting an unbroken flow into a form the internal model can absorb without pause. We can train a VAE encoder to compress the world streams.

  • Interaction Interface. The channel through which an operator, another agent, or the environment reaches into a process already in motion, with a correction, a command, a nudge. Input arrives not as a reset but as a perturbation, and the model bends around it on the fly. We can train a multimodal condition encoder to realize this.

  • Rendering Decoder. Facing back toward the world, the decoder turns internal state into outputs: an observation rendered for a human, an action issued to the environment, a snapshot of current belief. Because any output is a partial view of a far richer state, it resolves detail only where needed, on demand and in time. We can train a VAE decoder or upscaler.

  • Perception, Memory, and Physics Tokens. Inside, the world is held as three kinds of tokens in one latent space. Perception tokens carry what is being observed now. Memory tokens carry what has been observed before, the compressed trace that spares the model from re-reading its history. Physics tokens carry what the world is doing beneath its appearance, the objects, forces, and relations that ground the model in real dynamics. We can use various representations to denote different levels of abstraction, such as DINO 32,33, SigLIP 34, and JEPA 22.

  • World Clock. The clock binds internal computation to external time, marking its passage whether or not new observations arrive, so that the model’s understanding never falls silently out of step. It is the heartbeat the rest of the architecture keeps time to.

  • Latent Autoregressive Generation. At the center, the model runs. Drawing on all three token streams and paced by the clock, it generates the next state of its world in latent space, folds in what actually arrives, and generates again, an unbroken cycle of prediction and correction in which the model forever anticipates the world, is answered by it, and is changed by the difference. We can train an autoregressive model.

Applications

Liveness will open a whole new era of applications: entire industries whose product is a real-time, running world, whether virtual or physical, will become possible for the first time. We sample the seven major application verticals below.

  • Entertainment will stop being content you request and become a place you inhabit. A game world will keep unfolding whether or not anyone is playing, and a film will become a live negotiation with its audience: the storyline can branch for a single viewer or follow the room’s vote, a fan can step into the cast, and a spin-off can be generated around whoever asks for one. Live streams themselves will be generated in the moment they are watched.

  • A live persona will never sign off. A virtual host can stream, sell, and banter with its audience in real time, and a lifelike idol can perform on stage and still meet fans afterward. A companion with continuous memory can help someone practice a foreign language, rehearse a speech, or take the fear out of a first conversation. And for a person missing someone, or missing a time in their life, steering a generated moment back to the old family dinner table is not about accuracy. It is about feeling close again.

  • Recommendation will stop ranking your past and start reading your present. The content, including ads, will be generated for you, not merely selected for you, and it will adapt within the session rather than overnight. A shopper can try clothing, a hairstyle, or makeup on their own live image, a magic mirror can dress a passerby in the store’s new collection, and a camera pointed at a bare room can show it renovated. Creative, recommendation, and conversion will collapse into one live loop.

  • For cars and robots, the world simulator will become the training ground, the test track, and the safety case. A live world model can synthesize the near-crash corner cases no fleet should have to collect, close the training loop at the latencies driving demands, and replay recorded logs with live interventions to ask what a robot or a car would have done. Automated red-teaming will hunt for failures before the world finds them.

  • Science will gain simulators that learn and run. A live model of the atmosphere or of a fusion plasma can be steered and questioned while it runs, continuously realigned against streaming measurement, and it can capture dynamics for which no clean equations exist. In medicine, the same loop will become a surgical simulator for training and an augmented view through the endoscope during diagnosis.

  • The best classroom will be a world a student can step inside. A lesson becomes runnable, so a student can slow a collision, rewind a reaction, or walk through ancient Rome as it is generated around them. A live tutor will watch the student work and reshape the world in response, fitting the pace, the difficulty, and the example to the learner in the moment.

  • The screen itself will become generative. An interface will no longer need to be designed in advance, since the system can render the exact view a task requires, from a disaster-response display assembled mid-crisis to a signing avatar translating any audio stream live for a deaf user. A planner or a teacher can render the scenario under discussion while it is being discussed. Pushed further, a generated layer will settle over raw reality, with overlays composed on the fly and scenes filtered by style, mood, or information density. We will interact less with pixels someone drew in advance and more with pixels a model is drawing now.

Mapping the Five-Layer Stack of the LivEconomy

Video streaming and downloads account for over 80% of all global internet traffic by Cisco’s estimate, and the share is still climbing 35. The internet, by volume, is mostly moving pictures of the world. We think current world models are living their BERT moment to prove that the idea works, and Live Models, we believe, will be their GPT moment. The scaling law, if applied to live models, will turn proof into an economy. When the unit of compute changes from the call (bounded, bursty, stateless) to the world-hour (sustained, stateful, always on), every layer beneath it must be rebuilt. Nvidia’s CEO, Jensen Huang, has described AI as a five-layer cake: energy, chips, infrastructure, models, and applications, with energy as the binding constraint at its base 36. We borrow his lens to ask what liveness changes at each layer, reading the stack from the top down. The result is what we call the LivEconomy.

  • Agents and Apps. The unit of value shifts from artifacts priced per token to the uptime of worlds: an hour of a character’s existence, a day of a factory twin, a mile of a vehicle’s held belief. Products should learn retention inside persistent worlds rather than optimization of single sessions, and pricing must move from usage to presence. The defining applications will not be tools people open, but worlds people join.
  • Models. At the center sit Live Models themselves, trained on continuous multimodal streams and scaled along a new axis, duration. The six pillars are the specification for this layer, and the open problems are the field’s hardest, e.g., objectives for continuous experience, benchmarks for persistence, safety for a system that never stops. Whoever solves them first will define the era’s standards, as the GPT series once did for language.
  • Infrastructure. A process that lives for months needs session-native infrastructure: world-state that persists and migrates, streaming input and output by default, and failover that never drops the thread of time. Billing, autoscaling, and fault tolerance were all designed for calls that end, and rethinking them is the challenge. The opening is a new platform layer, and platform layers are where the cloud’s largest franchises have always been built.
  • Chips. Live Models call for silicon of their own: bounded latency, persistent state held close to compute, and continuous low-batch inference. That profile is an invitation to new architectures, memory-centric designs and dedicated inference silicon sold on latency guarantees rather than peak throughput. Silicon for the LivEconomy will look less like a batch processor and more like a heartbeat that never stops.
  • Energy. Continuous world simulation draws power as baseload rather than bursts. Inference becomes a utility, its marginal cost quoted in joules per second of world. The challenge is capacity, since always-on inference stacks on top of training demand that already strains grids. The opportunity lies in the shape of the load: a flat, predictable draw is the ideal customer for always-on clean power such as nuclear and geothermal, making live inference a natural anchor tenant for new plants.

Each layer is a major industry in its own right, and no single company will build them all. The fastest way to make this economy real is for energy developers, chipmakers, cloud builders, model labs, and application teams to start building for liveness now, from the energy contract up. This is not a market to divide. It is a market to create.

Concluding Remarks

In this blog, we introduced Live Models, a new class of world models that run in real time. We defined liveness through six functional pillars, offered an interpretation in the language of physics as a learned dynamical system, and sketched a reference architecture built on shared tokens, a world clock, and latent autoregressive generation. We then surveyed the applications this unlocks, from entertainment and live personas to robotics, science, and education, and mapped the five-layer LivEconomy that will grow beneath them.

We will not pretend to know the path, and some of what this essay proposes will turn out to be wrong. We look forward to the corrections. But we are confident about where this leads. Live Models will set off the next wave of AI, this time in the physical world, and its impact will reach far beyond software: it will change how people work, play, learn, shop, and stay connected to one another, and it will reshape society as thoroughly as the internet did. We intend to help build that future, and we invite you to build it with us.

References

  1. [1]K. He, X. Zhang, S. Ren, and J. Sun. Deep Residual Learning for Image Recognition. CVPR, 2016.
  2. [2]J. Ho, A. Jain, and P. Abbeel. Denoising Diffusion Probabilistic Models. NeurIPS, 2020.
  3. [3]A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language Models are Unsupervised Multitask Learners. OpenAI Technical Report, 2019.
  4. [4]S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR, 2023.
  5. [5]A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet Classification with Deep Convolutional Neural Networks. NeurIPS, 2012.
  6. [6]A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention Is All You Need. NeurIPS, 2017.
  7. [7]A. Dosovitskiy et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ICLR, 2021.
  8. [8]I. Goodfellow et al. Generative Adversarial Networks. NeurIPS, 2014.
  9. [9]A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen. Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv:2204.06125, 2022.
  10. [10]OpenAI. Video Generation Models as World Simulators (Sora Technical Report). 2024.
  11. [11]Google DeepMind. Veo 3: A Video Generation Model with Native Audio. deepmind.google, 2025.
  12. [12]ByteDance Seed. Seedance 2.0: Advancing Video Generation for World Complexity. arXiv:2604.14148, 2026.
  13. [13]J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL, 2019.
  14. [14]K. J. W. Craik. The Nature of Explanation. Cambridge University Press, 1943.
  15. [15]D. Ha and J. Schmidhuber. World Models. arXiv:1803.10122, 2018.
  16. [16]F.-F. Li. From Words to Worlds: Spatial Intelligence is AI’s Next Frontier. drfeifei.substack.com, 2025.
  17. [17]F.-F. Li and the World Labs team. A Functional Taxonomy of World Models. drfeifei.substack.com, 2026.
  18. [18]J. Bruce, M. Dennis, A. Edwards, et al. Genie: Generative Interactive Environments. ICML, 2024.
  19. [19]World Labs. Marble: A Multimodal World Model for Generating Explorable 3D Environments. worldlabs.ai, 2025.
  20. [20]NVIDIA. Cosmos World Foundation Model Platform for Physical AI. arXiv:2501.03575, 2025.
  21. [21]D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering Diverse Domains through World Models. Nature, 2025.
  22. [22]Y. LeCun. A Path Towards Autonomous Machine Intelligence. OpenReview, 2022.
  23. [23]V. Gabeur, S. Long, S. Peng, et al. Image Generators are Generalist Vision Learners (Vision Banana). Google DeepMind, arXiv:2604.20329, 2026.
  24. [24]L. Wang, C. Zhang, R. Kabra, J. Uijlings, S. Waslander, A. Zisserman, J. Carreira, K. He, et al. Video Generation Models are General-Purpose Vision Learners. arXiv:2607.09024, 2026.
  25. [25]R. E. Kalman. A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering, 1960.
  26. [26]Y. Sun, X. Wang, Z. Liu, J. Miller, A. A. Efros, and M. Hardt. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts. ICML, 2020.
  27. [27]J. Kirkpatrick, R. Pascanu, N. Rabinowitz, et al. Overcoming Catastrophic Forgetting in Neural Networks. PNAS, 2017.
  28. [28]Thinking Machines Lab. Interaction Models: A Scalable Approach to Human-AI Collaboration. thinkingmachines.ai/blog, May 2026.
  29. [29]A. Défossez, L. Mazaré, M. Orsini, et al. Moshi: A Speech-Text Foundation Model for Real-Time Dialogue. Kyutai, arXiv:2410.00037, 2024.
  30. [30]R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. Duvenaud. Neural Ordinary Differential Equations. NeurIPS, 2018.
  31. [31]L. P. Kaelbling, M. L. Littman, and A. R. Cassandra. Planning and Acting in Partially Observable Stochastic Domains. Artificial Intelligence, 1998.
  32. [32]M. Oquab et al. DINOv2: Learning Robust Visual Features without Supervision. TMLR, 2024.
  33. [33]O. Siméoni et al. DINOv3. arXiv:2508.10104, 2025.
  34. [34]X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid Loss for Language Image Pre-Training. ICCV, 2023.
  35. [35]Cisco. Cisco Visual Networking Index: Forecast and Trends, 2017-2022. Cisco White Paper, 2018.
  36. [36]J. Huang. Remarks on the five-layer AI stack: energy, chips, infrastructure, models, and applications. World Economic Forum, Davos, and NVIDIA GTC keynote, 2026.

Citation

@online{visko2026livemodels,
  author = {Zhengzhong Tu and Jie Yang and Qing Yin},
  title = {Introducing Live Models: A Frontier for World Models},
  date = {2026-07},
  year = {2026},
  url = {https://www.visko.ai/blog/introducing-live-models},
}