If you get an AI output you didn't expect, it can be tempting to explain it away in hindsight. The result was bad? Well, maybe better instructions would help next time. And if the result was good? Well, maybe it got lucky.
Results, and how to achieve them, often seem obvious once you've seen the outcome. But that's not a scientific approach to understanding something. If you can't design the experiment – or configure the AI agent – ahead of time so that it reliably produces the desired result, then you haven't really demonstrated that it works. You're just creating stories about outcomes after the fact.
This is why it's so useful to write down predictions about performance before you see the result. Recently, we made a quiz that allowed participants to do exactly this: predict how well Claude, Gemini and ChatGPT would do when given 10 attempts at a range of tasks.
What happened when people put their judgement to the test?
In some cases, human optimism was justified. When it came to Gemini extracting text from an image, the most common human prediction was 10/10 correct task completions, and the model met expectations:

For context, this was the image in question:

In other cases, people were correctly pessimistic. When it came to ChatGPT drawing an accurate hopscotch game, the most popular human prediction was 3/10, and the model… er.. also met these low expectations:
But for other tasks, like Claude multiplying two four-digit numbers, people were over-optimistic:
(The incorrect answer 16,665,069 was given four times.)
People were also too optimistic about ChatGPT counting which word had more letters.
(The incorrect answer 'they are the same length' was given three times.)
In the case of Claude tallying up happy responses in a simple dataset, human predictions were scattered, and so was the model's performance:
(There were 400 strongly happy statements in total, and 400 sad. Incorrect answers included 406, 408, 410 and 438.)
When it came to adding rows and columns of a small CSV, Gemini was generally good, although not perfect:
The above echoes a lot of what we hear from teams we work with. They know AI can do some incredible things. They just don't know which specific things on a given day, and they really don't want to be the person who trusted it the time it was wrong. That uncertainty – the not knowing – is what stops so many teams from using these tools effectively for data crunching at big scales, where manual checking just isn't an option.
Rather than a lack of capability, it's a lack of confidence in when that capability will turn up to work. This is why stability and being uncertainty-aware are must-haves in the analysis we do at Wholesum.