You will not have a reference book open when the claim lands
The next bold AI claim you meet will arrive in a sales call, a launch email, or a founder's post at eleven at night, and you will have to decide on the spot how much of it to believe. I spent 2026 scoring 132 public claims from fifty named AI leaders, and the only way to do that at scale was to use the same test every time. This post is that test, in full. It takes about five minutes per claim and needs nothing but the claim itself and a search engine.
The four categories
Every claim lands in exactly one of these.
Confirmed means the claim is independently checkable and it happened as stated. Independently is the operative word. A company's own report of its own success is not confirmation. If the only evidence is the claimant's own unaudited word, note that rather than upgrading it.
Refuted means the claim did not hold, either because the same person's later actions contradicted it or because another dated fact plainly conflicts with it. Refuted does not require malice. Plenty of refuted claims were sincere when made.
Pending means it cannot be resolved yet, because the resolution date has not arrived or the evidence to check it does not exist. It is the largest category in my ledger, and that is not a weakness.
Unfalsifiable means the claim offers no fixed way to resolve it, ever. An aspiration, a values statement, a comparison with no baseline. "We believe in a positive AI future" is unfalsifiable. "Our model will pass this benchmark by this date" is not. Unfalsifiable claims are the ones most often mistaken for promises.
The Claim-Scoring Canvas
Five steps, in order. Do not skip step two. It is the one everyone skips and the one that decides everything after it.
Step 1. Write the claim in one sentence, in the claimant's own words. Do not paraphrase yet. Paraphrasing is where vagueness creeps in. Write down exactly what was said.
Step 2. Define the contested term. If the claim uses a word like "AGI," "open," "faster," or "safe," write down what that word would have to mean for the claim to be checkable at all. If the claimant never defined it, say so. An undefined term is usually the entire reason a claim survives scrutiny.
Step 3. Find the closest independent check, and name what kind of evidence it is. Ask what would have to be true in the world, separate from the claimant's own statement, for the claim to hold. Then ask whether that evidence exists yet and who produced it.
Step 4. Assign the category, and write the one sentence that would change your mind. If you cannot name the new fact that would move the claim to a different category, it is very likely unfalsifiable, whatever you were about to call it.
Step 5. Ask who benefits if you believe it, and say so out loud. This does not change the category. It changes how much weight you put on anything marked pending or unfalsifiable, because those are the two categories most often repeated as if they were confirmed.
Worked example: "the AGI era"
Here is the canvas run on the loudest claim of the year.
Step 1, the claim. On September 3, 2026, Sam Altman framed the launch of GPT-6 Astra as the arrival of "the AGI era."
Step 2, the term. What does AGI mean here? Altman's own January 2025 reflections said OpenAI was "now confident we know how to build AGI as we have traditionally understood it." No operational definition accompanies that statement anywhere in the record I reviewed. Without one, "the AGI era" cannot be checked against anything. It can only be believed or not believed.
Step 3, the independent check. In the same week, the independent benchmark tracker Artificial Analysis scored Astra behind a competing model, Claude Fable 5.1, on its own evaluation. Three days after the launch, OpenAI's own chief scientist, Jakub Pachocki, published an essay urging caution about exactly the kind of triumphant framing his CEO had used. Both pieces of evidence are independent of Altman's framing, and both cut against it.
Step 4, the category. Refuted or disputed, not merely pending, because the claim was undercut within the week by an outside evaluator and by evidence from inside the company itself. What would change my mind: a defined, agreed operational test for AGI that Astra then passed, repeated by an independent evaluator. No such test exists in the record, which is itself informative.
Step 5, who benefits. Altman and OpenAI benefit directly, in fundraising and in positioning against Anthropic and Google DeepMind, from the public believing AGI arrived under their roof first. That does not make the claim false. It means it was made by the person with the largest financial incentive on earth to make it, and that belongs in the assessment even after the category is assigned.
Your turn: a second claim to score
Now try one yourself. Here is a claim from a source I treat far more sympathetically, to show the method does not sort by how much you like the speaker.
The claim: on July 30, 2026, at its Epoch developer conference, Sarvam AI priced its Sarvam-105B model at $0.80 per million blended tokens, framed as roughly 5.5 times cheaper than GPT-5.4 Mini and 11 times cheaper than Gemini 3.5 Flash.
Run the five steps before reading on. What exactly was claimed? Is "cheaper" defined, and against what? What independent evidence exists, and who produced it? Which category, and what single fact would move it? Who benefits if you believe it?
Here is where I landed, so you can check your work. "Cheaper" is unusually well defined here: price per million blended tokens against two named competitor products. The independent check is available immediately, because GPT-5.4 Mini and Gemini 3.5 Flash pricing is public and set by companies with no reason to help Sarvam look cheap. Category: confirmed. What would date it: either competitor cutting price below Sarvam's stated multiple, which would not retroactively falsify the July 30 claim but would make it stale. Who benefits: Sarvam, and the wider argument its founders make that India should treat foundational AI as sovereign infrastructure. That agenda does not weaken the claim, because the check does not depend on trusting the source.
Both claims come from founders with obvious incentives. One fails the check. One passes it. The incentive was not the deciding factor in either, which is the method working as intended.
Four mistakes to watch for
Treating a large, confident number as proof of itself. Specificity is not verification. A precise percentage feels credible purely because it is precise. It only helps once you have completed step three and found something outside the claimant to check it against.
Letting repetition substitute for a second source. TIME's TIME100 AI profile reported Sarvam's Series B as $75 million from Nvidia. HCLTech's own release put it at $234 million with no Nvidia involvement. A wrong number in a credible publication gets carried forward by everyone who does not think to check. Five outlets copying one report is still one source.
Scoring the person instead of the sentence. Every claimant made strong and weak claims in the same year. Score each sentence on its own evidence, not on the speaker's record on a different sentence.
Closing the file too early on a pending claim. Pending is not a resting state. It is a claim with a date attached. Write the date down and revisit it.
The whole thing, stripped to five lines
- What exactly was claimed, in the claimant's own words?
- What contested term needs a definition, and did the claimant supply one?
- What independent evidence exists, separate from the claimant, and who produced it?
- Confirmed, refuted, pending, or unfalsifiable, and what single new fact would move it?
- Who benefits if you believe this, and does that change how loudly you should repeat it?
Find one claim from this week's AI news and run it through those five today. If you cannot name the fact that would move it, you just caught an unfalsifiable claim before it had the chance to sound like a promise.
If you want the method applied 132 times
This canvas is Chapter 8 of The Claims Ledger, the book I wrote from a year of doing exactly this to fifty of the most powerful people in AI. The book has the full ledger of 132 scored claims in the back, worked chapters on the open-weights reversal and the launch-day walkback, and twelve pending claims to re-check next year. It is $19, PDF and EPUB, at /book. The method above is complete on its own and you do not need the book to use it.