HyperAIHyperAI

Command Palette

Search for a command to run...

From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation

Quanyu Long Xiao Chen Jianda Chen Haozhen Zhang Qisheng Hu Jianzhu Bao Wenya Wang

Abstract

Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. We instantiate this paradigm with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs these traces into a reusable environment worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge. At runtime, the world model agent actively consults the worldbook together with persistent episodic state to infer each action’s observation and lasting state efects. Across nine environments, Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional promptbased language world models. In multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment, indicating that its simulated dynamics better preserve the consequences of earlier actions across successive turns. These results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.

One-sentence Summary

Researchers from Nanyang Technological University and The Hong Kong Polytechnic University propose Trace2Env, a learning-free agentic language world modeling framework that reconstructs historical interaction traces into a reusable environment worldbook, allowing a world model agent to consult environment schemas, grounded evidence, and induced behavioral knowledge alongside persistent episodic state, thereby improving next-observation fidelity and long-horizon interaction consistency across nine environments compared with conventional prompt-based language world models.

Key Contributions

  • The paper introduces agentic language world modeling, a paradigm where a world model agent serves as a stateful environment simulator for a task agent instead of requiring an executable replica of the original system.
  • The paper presents Trace2Env, a learning-free framework that reconstructs historical interaction traces into a reusable environment worldbook containing schemas, grounded evidence, and induced behavioral knowledge, then uses that worldbook with persistent episodic state to infer observations and lasting state effects.
  • Across nine environments, Trace2Env improves next-observation fidelity and long-horizon interaction consistency over prompt-based language world models, and task-agent actions generated against Trace2Env remain valid more often when replayed in the real environment.

Introduction

Interactive agents require environments that respond to actions and preserve their consequences, but building realistic replicas of terminals, software workspaces, or enterprise systems often requires code, data, or infrastructure that is unavailable. Historical interaction traces are often still accessible and contain behavioral evidence, yet prior trace-based approaches either produce skills for the acting agent or reconstruct executable workspaces, which may still depend on missing implementation details. Prompt-based language world models also flatten environment knowledge and long-horizon state into a single context, making them unreliable for maintaining state continuity over many turns. The authors address this by proposing Trace2Env, a learning-free framework that turns historical traces into a structured, non-executable environment worldbook and uses an agentic language world model to actively inspect that knowledge, maintain persistent episode state and memory, and simulate environment responses.

Method

The authors introduce Trace2Env, a framework designed for trace-based environment reconstruction that transfers behavioral knowledge across episodes without importing episode-specific facts. The system operates in two distinct phases: an offline phase that reconstructs an environment worldbook from recorded traces, and an online phase that combines this fixed knowledge with mutable episode state and interaction memory. A world model agent proposes each transition during the online rollout, while a shared harness controls which state changes are committed.

As shown in the figure below, the overall pipeline is divided into environment reconstruction, worldbook storage, and agentic simulation.

During the offline phase, the authors leverage recorded transitions to build a reusable environment worldbook. A recorded transition provides concrete evidence about a single interaction, but simulation requires knowledge that generalizes to unseen situations. The offline constructor processes the build traces to produce the worldbook KE=Constructϕ(Dbuild)K_{\mathcal{E}} = \mathrm{Construct}_{\phi}(\mathcal{D}_{\text{build}})KE​=Constructϕ​(Dbuild​). This worldbook comprises four complementary components. First, schemas describe the action interface and the state variables the simulator can maintain. Second, grounded evidence retains recorded transitions, selected demonstrations, and their original observations. The constructor aligns each action with its result to extract observed facts, candidate state effects, outcomes, and uncertainty, while withholding effects with ambiguous attribution. Third, induced abstractions capture recurring behaviors across traces, including conditional effects, state constraints, observation contracts, and descriptive conventions. The constructor proposes behavioral rules and reviews them against supporting and contrasting cases to refine these abstractions. Finally, provenance records link each artifact back to the specific traces and turns that support it, allowing the simulator to inspect the environment at varying levels of specificity.

To distinguish knowledge that transfers across episodes from facts established only within the current interaction, Trace2Env maintains a persistent episode workspace. Treating all signals as a single flat history would make their scope ambiguous. Instead, the workspace is defined as Wt=(KE,s^t,Mt)\mathcal{W}_{t} = (K_{\mathcal{E}}, \hat{s}_{t}, M_{t})Wt​=(KE​,s^t​,Mt​), where KEK_{\mathcal{E}}KE​ is the immutable environment worldbook shared across episodes, s^t\hat{s}_{t}s^t​ is the mutable represented state of the current episode, and MtM_{t}Mt​ is the episodic interaction memory. This separation ensures that cross-episode evidence informs how the environment behaves without establishing what exists in the current episode, treating missing state entries as unknown rather than absent.

In the online agentic simulation phase, the system produces the next transition in two stages. Given an action ata_tat​, the world model agent starts with a compact view of the current episode and selectively inspects additional information from the workspace. Because cross-episode evidence introduces the challenge that relevance does not imply applicability, the system separates retrieval from applicability. For each retrieved worldbook entry eee, an applicability gate assigns a classification:

gtapp(e)=Gapp(e∣at,Wt)∈{supporting, uncertain, format-only}g_{t}^{\mathrm{app}}(e) = G_{\mathrm{app}}(e \mid a_{t}, \mathcal{W}_{t}) \in \{\text{supporting, uncertain, format-only}\}gtapp​(e)=Gapp​(e∣at​,Wt​)∈{supporting, uncertain, format-only}

Supporting evidence informs concrete behavior, uncertain evidence is used cautiously, and format-only evidence contributes response structure without establishing episode-specific facts. The agent alternates between inspecting the workspace and reasoning about the transition until it is ready to externalize its decision as a transition proposal qt=(Δt,o~t+1)∼Pθ(⋅∣dE,at,Wt)q_{t} = (\Delta_{t}, \tilde{o}_{t+1}) \sim P_{\theta}(\cdot \mid d_{\mathcal{E}}, a_{t}, \mathcal{W}_{t})qt​=(Δt​,o~t+1​)∼Pθ​(⋅∣dE​,at​,Wt​). Here, Δt\Delta_{t}Δt​ is an ordered list of candidate state effects, and o~t+1\tilde{o}_{t+1}o~t+1​ is the proposed observation.

Rather than modifying the episode directly, the proposal is passed to the shared harness for validation. The harness applies a validation gate:

gtval(qt)=Gval(qt∣at,Wt)∈{accept, reject}g_{t}^{\mathrm{val}}(q_{t}) = G_{\mathrm{val}}(q_{t} \mid a_{t}, \mathcal{W}_{t}) \in \{\text{accept, reject}\}gtval​(qt​)=Gval​(qt​∣at​,Wt​)∈{accept, reject}

This gate checks the candidate effects on a state copy against the schemas, rule support, and applicable state constraints. Only accepted proposals are committed. The harness commits the staged effects and returns the resolved observation (s^t+1,o^t+1)=H(Wt,at,qt)(\hat{s}_{t+1}, \hat{o}_{t+1}) = \mathcal{H}(\mathcal{W}_{t}, a_{t}, q_{t})(s^t+1​,o^t+1​)=H(Wt​,at​,qt​). Once committed, the updated state and interaction memory persist into the next turn, providing continuity across the online rollout.

Experiment

The experiments evaluate Trace2Env across nine environments in two settings: next-observation prediction on AgentWorldBench and EnvScaler, and multi-turn interaction on ALFWorld and SciWorld. Comparisons with direct prompting, trace retrieval, worldbook prompting, and harness-only baselines show that the full system improves prediction quality and consistency, with reconstructed worldbook knowledge and the runtime contributing complementary gains. Long-horizon interaction results indicate that high simulated success alone does not guarantee a faithful world model, while Trace2Env achieves stronger transfer back to real environments by preserving grounded state constraints. Ablations further show that worldbook evidence and abstraction are complementary and that fidelity improves with more construction traces, and qualitative results highlight how persistent state tracking prevents incorrect transitions from derailing interaction.

For the gpt-5.6-sol backbone, Trace2Env achieves the highest average next-observation prediction score across the seven evaluated environments. Prompting with retrieved traces or worldbook context also improves over direct prompting, while the harness-only variant shows smaller and less consistent gains. The strongest scores occur in application-style environments such as Food, Shopping, and Benefits. Trace2Env leads overall and records the best score in most environments, with especially large gains over direct prompting in Terminal and Benefits. Worldbook prompting and trace RAG prompting both outperform direct prompting, with worldbook prompting slightly ahead on average and in most environments. Harness-only results are mixed, improving over direct prompting in some environments but trailing in others.

Direct Prompting achieves high task success in simulated world models, but its action sequences often fail when replayed in the real environment. Trace2Env has lower simulated success but much higher world-model-to-real consistency. The results indicate that simulated success alone is not a reliable measure of world model fidelity. In ALFWorld, Direct Prompting drops from 97% simulated success to 3% when its induced actions are replayed in the real environment. Trace2Env maintains substantially higher W2R consistency in both ALFWorld and SciWorld even though its simulated success is lower than Direct Prompting.

On Terminal, evidence provides the stronger standalone contribution to worldbook performance, while abstraction is most useful when paired with evidence rather than used alone. The full combination of abstraction and evidence achieves the best result. Worldbook performance also improves as more construction traces are used, though construction cost rises with trace volume, so a moderate trace count is chosen as a practical trade-off. Evidence improves the schema-only variant substantially, whereas abstraction alone does not help and slightly lowers performance. Adding evidence to an abstracted worldbook produces a larger gain than adding abstraction on top of evidence, and the full worldbook performs best. Performance becomes more stable and continues rising with additional construction traces, but one-time construction cost increases with trace volume.

The experiments evaluate trace-based environment induction against direct prompting and retrieval or worldbook baselines for next-observation prediction across seven environments, finding that Trace2Env performs best overall and especially in application-style settings such as Food, Shopping, and Benefits. A second study on simulated versus real replay shows that direct prompting can achieve high simulated success but poor world-model-to-real consistency, while Trace2Env trades some simulated success for substantially more faithful action transfer in ALFWorld and SciWorld. Ablations on Terminal indicate that evidence contributes more than abstraction to worldbook performance, the full combination of abstraction and evidence works best, and using more construction traces improves stability at added construction cost.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp