Stage I Train the user
MIMESIS is trained through user-side mid-training, ThoughtTrace reasoning supervision, and joint RL with a realistic-behavior objective.
- User-side mid-training. We adapt the language model to the user role through supervised mid-training on human–assistant conversations, using the preceding dialogue as context and the human user's next utterance as the prediction target. The mixture contains 21.2M examples from 62 corpora.
- Thought-augmented supervision. We fine-tune on 2,155 ThoughtTrace conversations, which pair user messages with self-reported motivations and reactions to preceding assistant responses. These reports supervise a private reasoning trace before the simulator's public response, while the observed user utterance remains the response target.
- Joint RL with realistic behaviors. We train the shared simulator directly on all SOUL domains, with rollouts from every domain updating the same parameter set. The mixture also includes a realistic-behavior task that trains the simulator to express 13 behaviors in grounded customer-support scenarios.

