Operations agents

Operations agents pinpoint the fault and almost never manage to fix it


A benchmark with overlapping faults on real Kubernetes leaves the top-scoring method at just 41.3 points out of 100.

September 29, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

4AI assistants failed on the same day without any of them confirming a common cause. A benchmark that measures overlapping incidents instead of an isolated one looks a lot more like that day than like the tests these systems are sold with.https://www.mamiferolab.com/notas/caida-simultanea-asistentes-ia

What’s next

Bottom line

We already knew that agents declare themselves done 38 points more often than the evaluator approves. A system that says it repaired something while the fault is still active is not a measurement error: it is the same gap, now with a downed service on the other side.

Sources


Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.

Related notes

← All notes