On August 25, OpenAI published the first numbers for Jalapeño, the chip it designed with Broadcom to run AI models. The results are strong: between 1.5 and 1.9 times more work per watt, and up to 3.6 times less latency than the Nvidia systems it was compared against. In interactive tasks the advantage reaches 4.1 times.
And the chip draws 700 watts, against the 1,200 and 1,400 of what it was measured against.
If that holds up, it’s huge. It means the cost of running an AI model is going to fall sharply, and that cost is what currently determines how much you’re charged to use these tools.
But before you repeat it, look at who measured it.
Three things the number doesn’t say
The tests were run by OpenAI. The benchmark is public and auditable, which is good. But the measurements were made by the company that owns the chip. Nobody outside has repeated them, partly because the chip doesn’t exist outside OpenAI: it isn’t for sale.
It’s lab silicon. These are engineering units, not volume production. Between a prototype that performs well in controlled tests and thousands of units working at real-world temperatures there’s a distance the industry knows well.
The efficiency comparison uses nominal factory power ratings. Not power measured during the test. It’s a legitimate way to normalize, and also one that favors whoever chooses it.
None of this means it’s a lie. It means it’s a self-reported number, published a day before Nvidia delivered quarterly results. The context doesn’t invalidate the figure, but it explains the timing.
Why we’re telling you, a small business owner
Because you’re not going to buy a chip. Ever. And even so, this is going to show up on your bill.
Your AI provider ultimately charges you for what it costs them to run the model. When that cost falls, your price should fall. And with OpenAI, Google, Amazon, Microsoft and Anthropic all making their own silicon, the direction is clear.
That’s where the only practical recommendation in this note comes from: don’t sign long AI contracts at today’s prices. If they offer you a discount for committing to two years, that discount is probably smaller than the price drop that’s going to happen on its own.
What we really wanted to say
This note isn’t about chips. It’s about how to read the numbers you’re going to be shown this year.
You’re going to see a lot of AI headlines with big numbers. “3.6 times faster.” “40% more conversions.” “Save 20 hours a week.” They’re almost always true under the exact conditions in which they were measured, and almost nobody tells you what those conditions were.
Three questions, and they work for any number put in front of you:
Who measured it? If the seller measured it, that’s not wrong, but give it less weight.
Compared with what? 40% more conversions compared with doing nothing isn’t the same as compared with doing it well by hand.
Under what conditions? Every number comes from a scenario. If they don’t tell you which one, the number isn’t verifiable, and an unverifiable number isn’t data: it’s advertising.
We’re going to show you numbers too. Apply the same three questions to us.
