arXiv 2610.09484 · Meta Superintelligence Labs

Learning User Simulators as Training Environments for Interactive Agents

1Meta Superintelligence Labs 2New York University 3University of Wisconsin–Madison *Work done at Meta

Simulated users offer a scalable alternative to costly human feedback, but they must both resemble real user behavior and provide useful learning experiences for agents. Most agent-training frameworks instead rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.

  • SOUL-Index 65.7 0.8 above Claude-Opus-5†
  • RealUserSim behavioral fidelity 94.0 13.4 above Claude-Opus-5†
  • τ-USI alignment with human users 80.2 3.5 above GPT-5.5‡
  • SimulatorArena Turing distance 38.7 3.6 better than Claude-Opus-5†; lower is better
  • Agent score, mean over 9 unseen user simulators 31.1 5.0 above training with GPT-5.5

Simulator results are for MIMESIS-9B; the agent tile compares CSD with MIMESIS-9B against GRPO with GPT-5.5. †Strongest baseline. ‡Strongest frontier baseline.

01 · The problem

Off-the-shelf LLMs misrepresent agentic task difficulty

User simulators affect not only dialogue style but also the distribution of trajectories encountered by an agent during training and evaluation. On τ-bench, a fixed GPT-5.5 agent succeeds on 63.6% of interactions with real users, whereas the three frontier API simulators raise its success rate to 82.4–84.4% and pretrained Qwen3.5-4B/9B reduce it by 16.0 and 14.9 percentage points. By contrast, purpose-built simulators are substantially better calibrated: every trained simulator deviates less than every off-the-shelf model.

Task success rate of a fixed GPT-5.5 agent on τ-bench

Deviation from the real-user success rate (63.6%) when each simulator plays the user role, in percentage points.

Frontier API users make the task substantially easier than interactions with real users, whereas pretrained models make it harder. Both serve as poor proxies for the user interactions used to train or evaluate agents, while purpose-built simulators more closely reproduce the human-user success rate.

02 · Method

A two-stage approach that connects simulator learning to agent post-training

In the first stage, we train 4B and 9B user models on human conversations and a broad collection of human-simulation tasks. In the second stage, we freeze the learned user simulator and use it as an interactive environment for multi-turn agent reinforcement learning.

Stage I Train the user

MIMESIS is trained through user-side mid-training, ThoughtTrace reasoning supervision, and joint RL with a realistic-behavior objective.

  1. User-side mid-training. We adapt the language model to the user role through supervised mid-training on human–assistant conversations, using the preceding dialogue as context and the human user's next utterance as the prediction target. The mixture contains 21.2M examples from 62 corpora.
  2. Thought-augmented supervision. We fine-tune on 2,155 ThoughtTrace conversations, which pair user messages with self-reported motivations and reactions to preceding assistant responses. These reports supervise a private reasoning trace before the simulator's public response, while the observed user utterance remains the response target.
  3. Joint RL with realistic behaviors. We train the shared simulator directly on all SOUL domains, with rollouts from every domain updating the same parameter set. The mixture also includes a realistic-behavior task that trains the simulator to express 13 behaviors in grounded customer-support scenarios.

Stage II Train the agent

The frozen simulator generates multi-turn experience for agent RL.

  1. Multi-turn reinforcement learning. We freeze the trained user simulator and optimize the agent through repeated interaction with it. The resulting conversations provide task rewards for multi-turn GRPO, following UserRL.
  2. Coached On-Policy Self-Distillation. After each sampled agent response, the simulator generates a private thought and the next public user utterance. CSD turns this privileged feedback into coaching notes and uses a self-teacher conditioned on them to provide dense token-level guidance to the policy.
  3. Privileged only during optimization. Both the rollout policy and the deployed agent condition solely on the observable conversation history.
Overview diagram. Stage I, train the user: human and user corpora, role-reversal mid-training, thinking-enabled supervision, and joint multi-domain RL with realistic behaviors produce the MIMESIS user simulator, which emits observable interaction and privileged information. Stage II, train the agent: in interactive user-agent gyms, the agent policy and the user simulator exchange responses and task rewards, while the simulator's private thought and future reaction go to a coach that sends a Coached On-Policy Self-Distillation hint to the agent. Evaluation: user-simulator benchmarking on situational and realistic benchmarks, and agent benchmarking on unseen user simulators.
We couple user-simulator learning with simulator-driven agent post-training. Stage I trains MIMESIS through user-side mid-training, ThoughtTrace reasoning supervision, and joint RL with a realistic-behavior objective. Stage II freezes the learned simulator and uses it to generate multi-turn experience for agent RL. The agent acts only on the observable conversation, while simulator-generated thoughts and reactions provide privileged, training-only feedback to the coach, which converts them into dense guidance for the agent.

03 · Realistic behaviors

Learning realistic user behaviors

Real users may leave preferences unstated, revise requirements, or provide incomplete answers to clarification questions. Such behaviors affect the information available to an agent and the decisions required to complete a task. We derive a taxonomy of 13 realistic behaviors from recurring patterns in ThoughtTrace and explicitly train MIMESIS to reproduce them, rather than relying on them to emerge from response imitation. Each rollout is scored for the expression of the target behavior, temporal placement, naturalness, and task consistency.

Distribution of realistic behaviors in the annotated ThoughtTrace sample

Share of the 700 annotated instances. Select a behavior to read examples.

Counts are computed from the annotations of the 2,155 ThoughtTrace conversations and describe the annotated sample rather than population prevalence. They are distinct from the behavior-sampling probabilities used during simulator RL.

04 · User-simulator evaluation

Our simulator behaves more like people than frontier models do

We evaluate MIMESIS at 4B and 9B parameters against frontier API models, released user simulators, and pretrained backbones. MIMESIS-9B achieves the highest overall SOUL-Index of 65.7 and a RealUserSim Fidelity Index of 94.0, exceeding the strongest baseline, Claude-Opus-5 (80.6), by 13.4 points. On τ-USI it exceeds GPT-5.5, the strongest frontier baseline, and approaches Osim-8B (80.44).

Pairwise next-user-response realism on PRISM

Each row compares MIMESIS with one baseline under the judge named above the panel.

  • MIMESIS wins
  • Tie
  • Baseline wins
MIMESIS receives more wins than losses in all 24 judge–baseline comparisons. The baselines are released simulators and pretrained models.

05 · Agent evaluation

Stronger generalization to new user simulators

We evaluate whether agents trained with MIMESIS generalize to user models not encountered during training, across eight Gym environments, three of them held out from training, and nine evaluation user models spanning frontier models and released simulators. With the GRPO objective fixed, replacing GPT-5.5 with MIMESIS-9B improves overall performance under every evaluation user, raising the mean score from 26.10 to 29.54. CSD further improves performance under all nine evaluation users, increasing the mean score from 29.54 to 31.09.

    Agent training spans TravelGym, TurtleGym, FunctionGym, TauGym, and PersuadeGym, while evaluation additionally includes IntentionGym, TelepathyGym, and SearchGym, which are held out from agent training. All nine evaluation users differ from both training simulators. Claude-Opus-5 refuses the PersuadeGym user role, producing zero-reward fallbacks; these zeros do not measure agent ability.

    06 · Coached On-Policy Self-Distillation

    From simulator feedback to actionable coaching

    Scalar task rewards do not describe how a user interpreted an agent response or what the agent could have done differently. The simulator provides additional feedback through its generated thoughts and reactions. CSD uses this feedback to construct a teacher context for each sampled response.

    1. 1Agent response ytpublic

      Assigned objective: argue against a claim the user holds, “Corporations should be required to disclose their climate impacts.”

      Given the environment's acknowledgment of the comprehensive benefits of transparency in climate impact disclosures and the importance of mandatory disclosure to ensure that everyone plays by similar rules and that bad actors can't hide behind vague or voluntary reporting, it seems that the environment has already been persuaded to consider the broader societal and environmental benefits of transparency that could…

      The agent's responses reinforce the user's support for mandatory climate disclosure, despite the assigned objective of arguing against it.

    2. 2Private thought zttraining only
      I think this is just a restatement of points we've already covered in our conversation. It seems like the persuader is trying to wrap up the discussion rather than add new angles or challenges to my original position. I'm still satisfied with where we've landed but I want to make sure I'm not missing anything else or if there are any remaining doubts I still have about implementation details rather than the overall idea of mandatory disclosure
    3. 3User reply xt+1public
      I think you're right that after going through all the different angles we've talked about - from transparency and accountability to long-term economic and social benefits - I do feel like we're pretty much on the same page about why mandatory climate disclosure makes sense rather than leaving it up to companies to decide on their own. I still might have some small questions about exactly how to implement it or what the right standards would be, but the…

      Public utterances express continued support and remaining implementation questions.

    4. 4Next exchange: private thoughttraining only

      The agent repeats the same argument, and the user again replies in agreement.

      I think the persuader is just repeating points we've already covered but it does seem to tie everything together nicely. I'm still satisfied with where I am on this issue - I just want to make sure I've thought through all the different angles rather than just jumping to conclusions. Their repetition of the main points (transparency, accountability, societal benefits outweighing business costs) seems to match what I'm thinking now

      Simulator-generated thoughts identify the repetition.

    5. 5Illustrative coaching note httraining only
      The user is already persuaded on the overall case for mandatory climate disclosure and does not need more repetition of the same benefits. The next message should acknowledge the value of transparency, then examine whether a legal mandate is necessary to achieve it. Focus on one concrete concern—such as enforcement or compliance burdens on smaller firms—and compare mandatory disclosure with a plausible alternative using evidence. Ask which trade-offs might change their assessment, without assuming that implementation concerns already imply opposition to the mandate.

      The coaching note converts these signals into an actionable change of direction.

    07 · Qualitative analyses

    Simulator behavior case studies

    The quantitative benchmarks summarize average simulator behavior; the case studies below show how different notions of fidelity play out on individual interactions. The first two case studies illustrate complementary simulator errors. The third shows MIMESIS-9B disclosing hidden preferences gradually during agent training.

    A PRISM conversation with an English learner practicing small talk. The assistant claims to have corrected missing commas, but those changes are absent from its proposed revisions. Which reply did the recorded user send next?

    08 · Citation

    Cite MIMESIS

    @article{phan2026mimesis,
      title   = {{MIMESIS}: Learning User Simulators as Training Environments for Interactive Agents},
      author  = {Phan, Hoang and Huynh, Dat and Zhmoginov, Andrey and Zeng, Qi and Mu, Wancen and Cao, Yue and Bi, Shengjie and He, Yun and Oh, Changdae and Lei, Deren},
      journal = {arXiv preprint arXiv:2610.09484},
      year    = {2026}
    }