Agent evaluation

A benchmark invents worlds with false rules to measure exploration


The set's 140 tasks contradict prior knowledge, so remembering doesn't help. Ten systems showed unstable performance across attempts.

September 27, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

140 tasks spread across two environments with 55 discovery goals.

Context

Agent evaluations keep running into the same problem: test sets end up leaking into training data and accuracy rises without capability changing. This year’s response was to build executable, verifiable environments instead of lists of questions. This adds a twist: the rules are designed to contradict what was learned.

What’s next

Bottom line

Measuring exploration requires a world where nobody has been before, and building that world is harder than building the model that travels it. That’s why almost every evaluation still rewards memory.

Sources

Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.

Related notes

← All notes