Banking and agents

Nubank simulates 16,000 conversations before changing its support agent


The Brazilian bank tested versions of its agent with synthetic users and raised self-service by 8.82 points without exposing anyone to a live failure.

September 25, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

The self-service rate rose 8.82 percentage points after an open-weights configuration was chosen through simulation.

Context

Banks in the region have spent months moving analysis tasks to language models; the note on ChatGPT taking over junior analysts’ work documents the other end of the same process. Nubank applies the same logic to the direct contact channel.

What’s next

Bottom line

Measuring agents on long tasks remains the open problem Taste-Bench laid out, with its best model at 59.7%. Nubank didn’t solve it: it sidestepped it by measuring the commercial outcome instead of judgment.

Sources

Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.

Related notes

← All notes