The claim that started this
On September 3, 2026, Sam Altman announced GPT-6 Astra and called it the arrival of what he termed "the AGI era." Twenty-four hours later, he was on X apologizing for a "messy rollout" that had locked paying subscribers out of the very model he had just called a new capability level for humanity. Same person, same week, same product, two very different stories.
That gap between the announcement and what happened next is where I started paying closer attention. Across the public records I reviewed, I noticed the same pattern again and again across the AI industry: a company or its leader makes a bold, specific, checkable claim. Weeks or months later, without much fanfare, the claim quietly resolves, one way or the other, and almost nobody goes back and checks.
So I built a ledger. For fifty named people, drawn from frontier lab leaders to policy voices to platform builders, I traced their major public statements from January through September 10, 2026, and scored each checkable one against what else had happened by that date. I turned it into a book, The Claims Ledger, and it is free through September 30, 2026 (UTC), then $19. This post is the short version: the method, three of the sharpest disagreements the ledger surfaced, and how to run the same test yourself on the next AI claim in your inbox.
The four-category test
Most coverage of AI predictions sorts claims into two piles: true or false. That instinct is wrong, and it is worth explaining why before I get to any specific person.
Take Sam Altman's January 2025 statement that OpenAI was "now confident we know how to build AGI as we have traditionally understood it." Is that true or false? Neither, because he never said what "AGI as we have traditionally understood it" means in checkable terms. No benchmark, no date, no task. A claim without a defined resolution condition is not a lie and not a truth either. It is a third thing.
Here is the test I ended up using on every claim in the ledger, and the one you can run on the next bold AI claim you see.
1. What exactly is being claimed, in one sentence, with no adjectives?
Strip "revolutionary" and "game-changing" and anything else that
describes a feeling rather than a fact.
2. What would have to be true for this to be checkable?
A date, a number, a named comparison, a defined outcome. If none
exists, the claim is unfalsifiable: treat it as a belief, not a
prediction.
3. Who is the source, and are they also the beneficiary?
A company reporting its own benchmark is not lying by default, but
it is not a neutral witness either.
4. Has this same person's later action already contradicted this claim?
This is the fastest way to catch a reversal before it becomes a
pattern.
5. If it is genuinely pending, what specific event resolves it, and when?
Write that event down. Then actually check back.
Run any claim through those five questions and it lands in one of four buckets: confirmed, refuted, pending, or unfalsifiable. "Pending" turned out to be the largest category by a wide margin in my ledger, which is itself a finding, not a gap in the research. Many of 2026's biggest claims, when AGI arrives, whether the current compute spending pays off, will not resolve for years.
Three disagreements the ledger surfaced
Jensen Huang versus Dario Amodei on jobs
On June 1, 2026, at a keynote in Taipei, Jensen Huang was asked about AI and employment. He did not hedge: "People talk about AI reducing jobs, complete nonsense." Five months earlier, in a January essay called "The Adolescence of Technology," Dario Amodei had put a number on the other side of that sentence: AI may displace up to 50 percent of entry-level white-collar jobs within one to five years.
Both men run companies that profit from enterprises adopting AI right now. Huang's company sells the chips; a denial that AI costs jobs is, among other things, a sales pitch to a procurement committee worried about headline risk. Amodei's company has built its identity and its fundraising pitch, including a $30 billion round at a $380 billion valuation, around being the safety-forward lab. Neither claim is resolvable from public data as of this writing. But the size of the gap, an unqualified "complete nonsense" against a specific "50 percent," is itself worth noticing, and neither man is the one who will actually lose a job if either is right.
The scaling debate nobody quite agrees on
Ilya Sutskever declared on the Dwarkesh Podcast in late 2025 that the field had moved from "the age of scaling" back to "the age of research." Then, in July 2026, his company Safe Superintelligence took a multi-billion-dollar compute deal explicitly to scale up "research worthy of scaling up," a partial reversal of its own two-year position. Meanwhile Sholto Douglas argued on a podcast titled "Sonnet 4.5 and the AI Plateau Myth" that progress has not slowed at all. Yann LeCun took a third position entirely: that scaling large language models cannot reach human-level intelligence regardless of how much research goes into it, because the architecture itself is wrong, a claim he backed by raising over a billion dollars for a competing approach.
None of the three address each other by name. The disagreement is real, drawn from independently stated positions, not a manufactured debate.
The regulation fight, explained for operators
David Sacks has spent much of 2026 calling proposals for AI oversight a "Trojan horse of a FINRA for AI" and "a DMV for AI." His central evidence: a Chinese open-weight model, Kimi K3, topped a coding benchmark in July 2026, which he argues proves that US regulation would hand the lead to competitors facing no equivalent restriction. Five days after he made that argument, OpenAI and Anthropic, bitter competitors for the same customers and compute, jointly warned about the risks of open-weight models, using the same underlying evidence for the opposite conclusion. That does not settle whether any specific regulatory proposal is well designed. It does suggest that a good benchmark score is not, by itself, the argument against oversight it is sometimes presented as.
What this means if you are not tracking any of this professionally
If you run a small business or manage a team, you do not get a vote on which of these disagreements resolves which way, but you will operate under whatever the industry and eventually the law settle on. The useful move is not picking a side today. It is building the habit of running the five-question test on the next confident claim that reaches you, whether it is a vendor's launch email, a keynote clip, or a headline. Ask what exactly is being claimed, whether it is checkable, who benefits from you believing it, and whether the same source's own actions already contradict it.
That habit is the actual product The Claims Ledger is built to teach, across sixteen chapters covering the AGI-era claim and its 24-hour walkback, the open-weights retreat, who pays when jobs are on the line, and a full reference table of every scored claim in the back matter. It is free through September 30, 2026 (UTC), no card required, then $19, and it comes with the method above turned into a repeatable worked framework in Chapter 8.
If prompts, not predictions, are more your problem this week, the Claude Cheat Sheet is a separate $19 reference: 120 prompt patterns with worked examples and honest notes on when each one does nothing. And if you are new to Claude entirely, the free 75-page guide is the place to start, no cost, no time limit.