Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Conceição Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
- Published
- Sep 24, 2026 — 17:07 UTC
{'Problem': 'The paper addresses a gap in the capability of customer experience (CX) agents, particularly in regulated industries, where enhancing agent performance is critical. The authors highlight the challenges in improving these agents and propose a novel simulation-based method to evaluate and refine their effectiveness. This work is presented as a preprint and has not undergone peer review.', 'Method': "The authors introduce a hypothesis-driven simulation workflow designed to screen CX agents effectively. They utilize the Snowglobe simulator to conduct their experiments. The data for the simulation is sourced from Nubank's Card Delivery agent and Card Management agent, with over 16,000 simulated conversations generated for evaluation. The performance of the agents is assessed using binary evaluator scores, comparing the simulated versions against their production counterparts.", 'Results': 'The simulation approach led to a notable improvement in customer experience metrics. Specifically, the tNPS (transaction Net Promoter Score) increased by 36.69 points compared to the previous version of the agent. Additionally, the SSR (Service Success Rate) saw an improvement of 8.82 percentage points over the prior level, indicating enhanced agent performance in customer interactions.', 'Limitations': 'The authors do not report any limitations in their study, and no obvious limitations are identified in the available text.', 'Why it matters': 'This work has significant implications for the development of AI agents in customer service, particularly in regulated sectors. By demonstrating the effectiveness of simulation for screening and improving agent performance, it opens avenues for further research into simulation-driven methodologies for agent training and evaluation, potentially leading to better customer experiences and compliance with industry regulations.'}
By Callan Zhang · Sep 24, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
