What happened
- On September 24, a Nubank team published on arXiv the details of how it tests its customer service agents before putting them into production.
- The method replaces manual testing with simulation: synthetic personas react to the agent’s responses and tools return simulated results, so the entire flow runs without touching the bank’s real systems.
- It was applied to the Card Delivery agent and its successor, Card Management, which the paper identifies as the bank’s highest-volume chat support agent in Brazil.
- Reported results: across four deployed versions, simulated and production scores showed high correlation. Simulation-guided iteration raised the transactional satisfaction indicator by 36.69 points in a live A/B test.
- In a second round, open-weights configurations were evaluated across more than 16,000 simulated conversations. The chosen model raised the self-service rate by 8.82 percentage points, the highest level recorded at the bank, with no statistically significant change in satisfaction.
Why it matters
- It’s production evidence from a bank in the region, measured in Portuguese and on real traffic, not a lab demo. For any team running automated customer service in Chile, the contribution isn’t the number but the method: you can measure a version before exposing it.
- The correlation between simulated and real is the point that makes everything else usable. Without that correlation, simulating is theater. The paper reports it over four deployed versions, which is a small sample and should be treated as such.
- There’s a consequence the paper doesn’t develop: if the synthetic personas are generated by the same kind of model that answers, the system is evaluated against its own idea of how people behave. The complaints nobody modeled keep reaching the human channel.
The number
The self-service rate rose 8.82 percentage points after an open-weights configuration was chosen through simulation.
Context
Banks in the region have spent months moving analysis tasks to language models; the note on ChatGPT taking over junior analysts’ work documents the other end of the same process. Nubank applies the same logic to the direct contact channel.
What’s next
- No timelines announced for releasing the simulator or the set of synthetic conversations.
- The paper is in version 1 and doesn’t state peer review.
Bottom line
Measuring agents on long tasks remains the open problem Taste-Bench laid out, with its best model at 59.7%. Nubank didn’t solve it: it sidestepped it by measuring the commercial outcome instead of judgment.
Sources
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


