Sales agents

A seller trained with reinforcement learning matches models with trillions of parameters


The agent splits a limited budget of conversation turns among buyers with private valuations and keeps its edge in markets it never saw.

September 29, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

One item per buyer and a fixed cap on turns: the two constraints that force a choice.

Context

Anthropic had already put its agents to negotiate, where the more capable model wins and not the better-written instruction. This paper points the other way: with task-specific training, a small model catches up with a large one on the same kind of task.

What’s next

Bottom line

Six banks have already asked to notify the buyer every time the one paying is an agent. When the other side of the counter also has one trained to extract surplus, that notice stops being a transparency formality and becomes price information.

Sources


Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.

Related notes

← All notes