Plate · Star trails, long exposure — repeated measurement made visible
Note
What should "confidence: 0.86" actually mean?
We attach confidence values to claims, so we owe readers a definition: a stated confidence is a promise about frequency — of the claims we mark 0.86, about 86 in 100 should turn out true. Here is the standard we hold those numbers to, and where it comes from.
Several of our published claims carry a number — confidence: 0.86, confidence: 0.41. Numbers like these are easy to print and easy to abuse, so this note states plainly what ours are required to mean, and what would prove them wrong.
The definition: confidence is a promise about frequency
A stated confidence is calibrated when it matches reality in the long run: of all the claims marked 0.86, about 86 in 100 turn out true; of those marked 0.41, about 41 in 100. That is the entire definition. A confidence number is not a mood, an emphasis, or a rhetorical softener — it is a checkable prediction about a batch of claims, and it can be audited by anyone who keeps score.
This standard was not invented by the AI industry. It comes from the one profession that has published probabilistic judgments, in public, under scorekeeping, for half a century: weather forecasting. When researchers examined U.S. precipitation forecasts, they found forecasters’ stated probabilities tracked observed frequencies remarkably well — when they said 70% chance of rain, it rained close to 70% of the time (Murphy & Winkler, 1977). Forecasters did not become calibrated by being humble or brilliant. They became calibrated because someone kept score.
The same discipline built the strongest results in judgment research: forecasters who track their own accuracy and update on misses reliably outperform credentialed experts who never check themselves (Tetlock & Gardner, Superforecasting, 2015).
What the machines can and cannot promise
Can AI models produce calibrated confidence? Partly — and the partial answer is exactly why a human standard matters. Research has shown that large models can estimate the probability that their own answers are correct with better-than-chance calibration (Kadavath et al., 2022). But calibration is fragile: OpenAI’s own technical report noted that the training used to make GPT-4 more helpful measurably worsened its calibration (OpenAI, 2023). A model can be tuned toward confidence the way a spokesperson can — and for the same reasons.
So we treat machine-stated confidence as an input, never a verdict. The number a claim ships with is one a person reviewed and was willing to stake a track record on.
The bands we use, in plain words
| Stated confidence | What we are telling you |
|---|---|
| 0.90 and above | We would be surprised to be wrong. Multiple independent sources agree. |
| 0.70 – 0.89 | Solid, with named gaps. Act on it, but read the assumptions. |
| 0.40 – 0.69 | Genuinely uncertain. The claim is useful mainly for what it tells you to check. |
| Below 0.40 | Stated to be challenged. We publish it because hiding a weak belief is worse than showing one. |
Two commitments follow. First: these numbers are auditable — every claim with a confidence value is public, so anyone can keep our score, and our corrections policy applies when a batch drifts from its promise. Second: we will not print false precision. A claim we cannot honestly band gets words (“we don’t know”), not a decimal.
A long-exposure photograph of the night sky shows what repeated measurement looks like: single points become arcs, and the pattern only appears across many exposures. Calibration is the same. No single claim can prove a confidence number honest — only the record can. That is why we publish ours.
Sources
- Murphy, A. H., & Winkler, R. L. (1977). Reliability of Subjective Probability Forecasts of Precipitation and Temperature. Journal of the Royal Statistical Society, Series C. doi.org/10.2307/2346866
- Kadavath, S., et al. (2022). Language Models (Mostly) Know What They Know. arxiv.org/abs/2207.05221
- OpenAI (2023). GPT-4 Technical Report — calibration before and after post-training. arxiv.org/abs/2303.08774
- Tetlock, P., & Gardner, D. (2015). Superforecasting: The Art and Science of Prediction.