AI Tools4 Sep 2026

Why the best benchmark often recommends the wrong model

Kemal Duran, alone in the office, above him the line YOUR OWN TEST COUNTS MORE
Created with AI, image made with Higgsfield (affiliate link).

This week a new model arrived with early access, ahead of the official rollout. On the aggregated leaderboard it ranks 5th, level with its predecessor. In the demo I watched, it didn't feel like 5th place. It felt like a leap.

That's the trap half the market is walking into right now. The number says one thing, the work says another, and most people believe the number because it's a number.

The table is an average, your job isn't

Take a current model from the top group. On a broad coding benchmark it scores 74.1 percent. A rival model scores 75.4 percent. On the aggregate leaderboard our model lands in 5th place, exactly level with its own predecessor.

Now the same model on a design and graphics test. 1st place. Ahead of the rival, ahead of two fast variants from other vendors. The AI judge that scores the results gives the second-place model 5.7 of the possible points on a single task, and anyone who looks at the visual result next to it sees right away: it's noticeably weaker.

Two scales, two completely different answers to the question "is this good?".

An aggregate ranking averages that away. It takes code, world knowledge, design, math, throws it all into one pot and stirs a rank out of it. The rank is mathematically correct and almost worthless for your decision, because you never do the average of all tasks. You do yours.

If you build graphics and interfaces all day, a model in 1st place on the design test is the right one, even if it drifts around in 5th place in the overall aggregate. If you run pure coding pipelines, the math comes out differently. The average knows nothing about your job.

Why "the best model" is the wrong question

The obvious answer sounds reasonable: take the model at the top. Saves you the thinking.

It falls short for two reasons.

First, the good benchmarks are saturating. Two of the well-known test suites, one for world knowledge at PhD level, one for advanced mathematics, are now considered saturated. All top models sit so close together that the number no longer separates them. A score of 99 versus 98 tells you exactly nothing about everyday usefulness. On another test, ARC AGI 3, one model jumps from 7.8 to 99.9 percent, while the average human tester sits at 48 percent. Superhuman, clean. And still the number doesn't help you one bit with the question of whether the thing can operate your Blender.

Second, no benchmark measures the ability that makes the real difference right now: a model controlling third-party software on your computer on its own. That doesn't show up on any leaderboard, because it's hard to pour into a percentage. But that's exactly where the practical leap is.

What's new when you use it

The test that shows progress cleanly is always the same: have the model rebuild a specific small game as a clone. A few model generations ago that took an hour and a half to two hours. Most recently: eight minutes. Part of that is probably lower server load from limited access and not pure model strength. Even if you factor that out, a clear difference remains.

The real moment came elsewhere. The model took control of Blender, a 3D software the person at the computer couldn't operate. It built a humanoid wolf figure, in eight minutes. Then it rigged the figure with 50 bones and put an eight-second walk animation on top, in six minutes. After that it handed the whole thing over to the Unreal Engine and turned it into a walkable forest world with the wolf as a playable character, in 35 minutes.

No human operated any of these programs along the way. I watched, with the clock running.

That's the point no table captures. The value isn't that the model got two percent better on a test. The value is that an ability jumped from "doesn't work" to "works". And you only find these jumps when you put the model on your real task, not on the task a benchmark team came up with.

How to test a new model before you approve it

This needs a fixed routine. Not reading benchmarks, but half an hour of honest work with the model on things you do every day.

Model entrance exam (30 minutes, before every approval)

1. Write down your three most frequent real tasks.
 Not "can it code", but
 "will it build me the Resend mail pipeline that I
 built by hand".

2. Define a standard task that you repeat for
 EVERY new model. Always the same one.
 It is your ruler across generations.

3. Stop the clock. Raw wall-clock time from prompt to
 finished, usable result. Not the time
 the model spends "thinking".

4. Log the costs. Note the token count of the task
 and convert it into money. A single test document
 recently cost 1.94 dollars at 63,858 tokens.
 That adds up faster than you think.

5. Test the ability that NO benchmark measures:
 Can it operate an unfamiliar tool on its own,
 one you have not mastered yourself? That is the test
 that decides between "incremental" and "leap".

6. Only then look at a leaderboard. As
 a cross-check, not as the basis.

The step that matters most is the second one. A task you repeat for every model is worth more than any outside table, because it measures your progress in your own language. The game clone from above is one such ruler. Build your own.

Where this gets overrated

Now the counterargument, and it belongs here, otherwise this is just enthusiasm.

Take the speed advantage from above with a grain of salt. Eight minutes instead of two hours sounds like a revolution, but part of it is plainly an empty server during limited early access. Once everyone is on it, the clock looks different. If you base your calculations on preview speed, you're getting rich on paper.

Second, the cost. The new model costs more per task than its predecessor. The jump in ability is real, and so is the jump in price. For a wolf figure in a forest world that doesn't matter. For a process that runs a hundred times a day, it decides everything.

Third, and this is the part nobody talks about in the launch videos: when a model controls software on your computer on its own, you hand over control. It clicks, it types, it saves. In the wolf demo that's great. In an environment with real data and real access, it's a question you want settled before approval and not after. Computer use belongs in a sealed-off environment first, with no access to anything that can hurt. Only once it's clear what the thing does on its own does it get anywhere near real folders.

And finally: the benchmark that shows a jump from 7.8 to 99.9 percent is probably just done. A test that a model "pretty much solves" measures nothing after that. A score close to 100 is not a reason to celebrate but a sign that this test won't tell you anything from now on.

FAQ

Are AI benchmarks useless?

No, but they measure the average across many task types, and your working day isn't an average. Aggregated leaderboards like Artificial Analysis scores weight code, world knowledge and design together, so a model can take 1st place on design-heavy tasks and still land only in 5th place in the overall aggregate. Use benchmarks as a cross-check, not as a reason to buy.

What is computer use in AI models?

Computer use is a model's ability to operate third-party software on a computer on its own, meaning it clicks, types and controls programs like a 3D software or a game engine itself. The user doesn't need to know how to use these programs. This ability barely shows up in classic benchmarks, but in practice it often makes the biggest difference.

How often should I repeat my test setup?

For every new model you're seriously considering, and always with the same standard task. That's the only way progress across model generations becomes measurable in your own language instead of in someone else's numbers. A fixed test case is your ruler, a changing task is no comparison.

Why is a benchmark score close to 100 percent a warning sign?

A score close to 100 generally means the test is saturated: all top models sit close together and the number no longer separates them. Two well-known suites for world knowledge and advanced mathematics are already considered saturated. A jump like the one from 7.8 to 99.9 percent on a single test means, as a rule, that this test won't tell you anything about differences between models from now on.

Does a stronger model automatically cost more?

Mostly, yes. The current top model costs more per task than its predecessor. Whether that pays off depends on how often you run it: for a one-off complex task the price is a side issue, for a process that runs a hundred times a day it decides whether the numbers work.

The one sentence

Never buy a model by its place in the table, but by what it delivers on your own standard task in thirty minutes.

Text and image were created with AI. Image made with Higgsfield (affiliate link).

← All insights