What happened
- A team published LLMAdBench on September 29, a human-preference benchmark for studying advertising inserted into language model responses.
- The design isolates a single decision: given a conversation, a response and a relevant ad, where the ad goes. The pairs compared differ only in that position; the query, the base response and the ad piece are held fixed.
- There are more than 18,000 human judgments under two transparency conditions: the ad labeled as sponsored, or blended into the response without saying so. Each pair is evaluated on six criteria, from the standpoint of both the advertiser and the reader.
- When eight frontier models are used as judges, the result is that they do not replace people: the most stable ones flip about a quarter of their decisions when the presentation order is changed, there is little agreement between models and their placement preferences differ systematically from the human ones.
Why it matters
- The finding that most changes a business decision is not the optimal placement but that disclosing the sponsorship systematically changes where people tolerate the ad. Transparency and placement stop being two separate discussions.
- Any team that wants to test ad formats inside an assistant using a model as judge is measuring noise. The paper quantifies it: a quarter of the decisions flip on their own.
- For the Chilean market, where advertising in assistants is not yet sold, this arrives before the inventory does. It is the window to set your own transparency standard instead of inheriting the vendor’s.
The number
One in four decisions is reversed in the most stable models when the presentation order is changed.
Context
Comscore had already measured that sponsored advertising in ChatGPT rose from 6% to 24% in three months, and the inventory still hasn’t reached Chile. The format is being standardized in other markets while here there is still nothing to buy.
What’s next
- A Qwen3-8B model fine-tuned on the human preferences outperforms the eight frontier judges on the held-out prediction task. No release date announced for that fine-tune.
- No announced timeline to extend the benchmark to other languages or other ad formats.
Bottom line
Two artificial intelligence engines already name the same brand in 81% and 43% of their responses. Now people are starting to measure where the paid ad lands inside that response too. The surface is the same and the two measurements will end up in the same spreadsheet.
Sources
- LLMAdBench: A Human Preference Benchmark for Advertising in LLM Responses, arXiv, announced September 29, 2026.
Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.




