WM@Booth brings together researchers from computer and social sciences to explore how world models can transform decision-making. The workshop will feature opinionated keynote talks on the future of AI, lightning talks, poster sessions, and culminate in hands-on sessions applying WMs to a variety of use cases.
When: August 31 – September 2, 2026
Where: The University of Chicago Gleacher Center, 450 Cityfront Plaza Dr, Chicago, IL 60611
Livestream: Free on YouTube — no registration required. Details.
Format:
Call for Papers: Submit via OpenReview by June 15, 2026 (AOE)
The organizers invite sponsors to connect with the community. In addition to supporting the workshop, you will be gratefully recognized in various media and materials, and have the opportunity to closely engage with workshop participants.
Information: Learn about various sponsorship levels, benefits, and opportunities. Sponsors are encouraged to contact the WM@Booth organizing committee.
Live on YouTube · August 31 – September 2
All three days stream live on YouTube — keynotes, lightning talks, and the panel included. Day 1 opens at 9:00 AM CT on Monday, August 31, and the recordings stay up afterward.
youtube.com/live/kP-CasXXiGE
Talks stream live on YouTube — free, no registration.
Mike Minnis, Deputy Dean for Faculty, Chicago Booth · Randall Balestriero, Brown University
Almost every leap in modern AI has come from doing more things at once: parallel hardware, then parallel training, and now parallel generation. Diffusion models, long the standard for images and video, are bringing that same parallelism to language. Instead of writing one token at a time, they draft an entire answer and refine it over a handful of passes. In this talk, I will trace how diffusion language models went from a contrarian bet to a mainstream paradigm, then turn to what it takes to extend them across modalities. Training a single diffusion model to both understand and generate over images and text surfaces problems that neither modality raises alone: high-dimensional visual detail is costly to represent, language-vision correspondences are sparse and easily missed during training, and understanding and generation present different modeling tradeoffs. I will describe LaViDa, a line of work that addresses these challenges and offers a new basis for large-scale multimodal world modeling. The larger claim is that the frontier worth optimizing for world modeling is universal intelligence per watt: universality demands joint reasoning over modalities, efficiency demands parallelism throughout, and diffusion is one of the few paradigms that offers both.
The timescale of a robot's decision determines how much test-time computation it can afford, and what has to be compiled into the policy beforehand. I will illustrate this with three works from my lab. At the millisecond level there is no time to deliberate: in-hand tool reorientation and precise assembly require closed-loop reaction at 60 Hz. Therefore, a world model's role is to generate the rollouts a reactive policy is trained on rather than to be queried online. I illustrate this with SimToolReal and Play2Perfect. At seconds-scale, explicit prediction becomes tractable and useful: MessyNav decides which obstacles in a cluttered scene can be moved and where they will end up. This raises the question of what representation this actually needs, since it mixes semantic affordance with dynamics prediction and it is not obvious that a world model in the usual sense is the right abstraction. When dynamics are unknown, the useful test-time computation is adaptation rather than search: Causal-PIK refines a physics-informed model in-context from failed attempts, solving tasks in few trials. I argue that the timescale of a task determines when prediction should be paid for, at training time or at test time, while the form that prediction takes is determined by the task itself.
Language models have become powerful engines for reasoning, but many embodied problems depend on knowledge that is difficult to express purely in language. In this talk, I will present an approach to embodied reasoning in which generative world models provide a space for imagining and evaluating possible futures. I will show how video models can reason through generation, using inference-time search and temporal backtracking to revise unsuccessful or physically implausible predictions. These models can also be combined with vision-language models, linking semantic reasoning and high-level planning with spatially and physically grounded predictions. When conditioned on actions, world models can simulate candidate controls, enabling model-predictive control and grounding high-level instructions into continuous robot actions. Across visual reasoning, long-horizon planning, and robotic manipulation, these results suggest that world models can serve as a general interface for imagining, evaluating, and refining actions before they are executed.
World Models: Enabling the Next AI Revolution
Recent advances in large language models (LLMs) have transformed human-AI interaction, however, building effective collaboration requires AI systems that truly understand the people they work with. In this talk, we first audit the U.S. workforce to assess the impact of automation and augmentation on the future of work, guiding the development of AI agents that reflect workers' perspectives. We then introduce General User Models (GUMs), which learn about users by observing any computer interaction and constructing propositions about user knowledge, preferences, and context. We further present NAP (Next Action Prediction), a framework for anticipating user intent by reasoning over rich multimodal sequences of human-computer interactions, where modeling long interaction histories enables significantly more accurate predictions. Overall, this talk highlights how to develop AI systems that are proactive and capable of fostering meaningful collaboration with human users.
Time-Series Modeling
Learning Representations of Assets and Investors with Applications to Financial Markets
The journey toward real-world autonomy in AI is fundamentally driven by the mastery of simulated environments and algorithmic self-improvement. This talk explores the evolution of agent training, tracing its roots in classical simulators where "learning to learn"—via automated curricula and AutoRL—first emerged as a critical lever. As the field transitioned to data-driven world models, these early meta-learning principles set the stage for today's frontier: generative foundation models that bootstrap their own capabilities. Next, we will examine how agents actively iterate and advance through self-improvement loops. We will dive into mechanisms like many-shot in-context learning, autonomous self-correction, and long-horizon reinforcement learning that allow models to escape the limits of static training data. By exploring applications across diverse domains—from web agents navigating dynamic digital interfaces, to foundation models achieving sub-angstrom precision in molecular design, to clinical agents mastering diagnostic reasoning via simulated patient encounters (ResidencyRL)—we will see how synthetic data and self-correction act as the engine of modern AI. Finally, by mapping these accelerating capabilities against levels of Artificial General Intelligence, we will analyze the critical interplay between performance, generality, and safe autonomy as we deploy agentic AI into high-consequence realities.
How to Train JEPA World Models Without Headache · Randall Balestriero, Brown University
Financial Data: Challenges, Evaluation, and Training · Bradford Levy, Chicago Booth
Amir Zadeh, Lambda
We invite submissions on world models from across disciplines — computer vision, reinforcement learning, robotics, finance, economics, and the physical sciences. Contributions may take any form: new architectures, empirical studies probing existing models in new domains, theoretical analyses, benchmarks, or novel applications. Work that bridges fields is especially encouraged.
| Format | NeurIPS 2026 style; submissions are up to 6 pages excluding references and appendix |
| Reviewing | Double-blind — submissions must be anonymized |
| Track | Single track; selected papers will receive best paper awards presented at the workshop |
| Archival | Non-archival — authors are free to submit elsewhere |
| Deadline | June 15, 2026 (AOE) |
| Author Notification | June 29, 2026 (AOE) |
Authors of accepted papers may request a letter of invitation for visa purposes.
Submit via OpenReview. Questions? Contact CAAI@chicagobooth.edu.