Dos burbujas de conversación, con sus colas hechas pedazos. Ilustración inspirada en Picasso y Joan Miró. Son la misma respuesta bajo las dos condiciones del estudio.

Ads in chats

A benchmark of 18,000 human judgments measures where an ad bothers people least in a chat


Eight frontier models evaluated as judges flip one in four decisions when the order is reversed and don't agree with real people.

September 29, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

One in four decisions is reversed in the most stable models when the presentation order is changed.

Context

Comscore had already measured that sponsored advertising in ChatGPT rose from 6% to 24% in three months, and the inventory still hasn’t reached Chile. The format is being standardized in other markets while here there is still nothing to buy.

What’s next

Bottom line

Two artificial intelligence engines already name the same brand in 81% and 43% of their responses. Now people are starting to measure where the paid ad lands inside that response too. The surface is the same and the two measurements will end up in the same spreadsheet.

Sources


Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.

Related notes

← All notes