The tally, first
Over 2026 I read the public record of fifty named people in AI, drawn from a compilation anchored on the TIME100 AI 2026 list, and pulled out every specific, checkable claim they made in public: predictions, benchmark numbers, promises, reversals. That came to 132 claims. I scored each one against what else the record showed as of September 10, 2026, the day I closed the ledger. Four categories: confirmed, refuted, pending, unfalsifiable.
Here is how the 132 break down.
Pending: 61. Confirmed: 42. Unfalsifiable: 16. Refuted: 12. One entry sits outside the four, a vendor marketing line marked unverified and likely overstated.
Two things about those numbers before the examples. First, pending is the largest category by a wide margin, bigger than confirmed and refuted put together. That is the finding, not a gap, and I will come back to why. Second, confirmed is softer than it looks. Of the 42 confirmed entries, only 18 carry the plain label. The other 24 carry a qualifier: company-reported, as reported, self-admission, partially confirmed, needs primary verification. The same is true down the list. Of the 16 unfalsifiable entries, 11 are plain. Of the 12 refuted, 6 are plain and the rest are marked partial, contested, superseded or disputed. I kept those qualifiers on purpose. A company blog post about its own benchmark is a claim that was made, not a fact that was checked, and the ledger says so every time.
Confirmed: 42, and only 18 of them the strong way
The strong kind of confirmation rests on evidence nobody with a stake in the answer produced. Three examples.
Andrej Karpathy, January 29, 2026, in a public GitHub discussion: he trained a model that beat GPT-2's CORE score for about $73 in 3.04 hours on eight H100s, and posted the numbers for anyone to reproduce. Alex Rives, May 27, 2026, at the Chan Zuckerberg Biohub: computationally designed protein binders against five named disease targets, with wet-lab results published the same day, hit rates of 36 to 88 percent for minibinders and 15 to 29 percent for antibody-derived designs, and PD-L1 binders that restored T-cell signaling in the lab. That is the most falsifiable claim in the whole corpus. And the SpaceX IPO in June 2026, priced at $135.00 a share across 555,555,555 shares, which matches the reported outcome exactly.
Now the qualified kind, which is more common. Demis Hassabis's team reported on May 7, 2026 that AlphaEvolve cut DNA-variant detection errors by 30 percent and raised power-flow feasibility from 14 to 88 percent. Sundar Pichai said at Google I/O on May 19, 2026 that Google processes over 3.2 quadrillion tokens a month, up 7x year over year. Both are marked confirmed, company-reported. I recorded them because the company said them and I have no evidence they are false. Nobody outside those companies has audited them, and that is a different thing from true.
Refuted: 12
Refuted means the claim did not hold, either because the same person's later action contradicted it or because a dated fact in the record conflicts with it directly. It does not require anyone to have lied.
Mark Zuckerberg wrote on July 23, 2024 that "open source is necessary for a positive AI future" and committed Meta to open frontier weights. On April 8, 2026, Meta shipped Muse Spark, its first frontier model on a new stack, closed. Elon Musk told a Baron Capital investor conference in November 2025 that Grok 5 would ship in the first quarter of 2026 with roughly six trillion parameters. As of September 10, 2026 it had not shipped. And TIME's TIME100 AI profile on August 27, 2026 reported Sarvam's Series B as $75 million from Nvidia, when HCLTech's own release of June 15, 2026 put the round at $234 million, led by HCLTech, with no Nvidia involvement. That error appears twice in the ledger because two Sarvam executives, Vivek Raghavan and Pratyush Kumar, each carry it in their public record.
Pending: 61
Pending means the claim cannot be resolved yet, either because its own date has not arrived or because the evidence to check it does not exist. It is not a weak verdict. It is the honest one for most of the biggest claims of the year.
Dario Amodei's January 2026 essay put a number on job losses: up to 50 percent of entry-level white-collar jobs within one to five years. There is no labor-market data in the record to check it yet. Demis Hassabis said in May 2026 that AGI arrives around 2030, plus or minus a year, an explicit disagreement with Amodei's much shorter window. Jensen Huang told the GTC San Jose keynote on March 16, 2026 that the AI buildout means "at least $1 trillion in revenue from 2025 through 2027." None of those three can resolve before the ledger closed, and I am not going to pretend they did.
Why pending dominates
There are two mechanical reasons, and they are both in the ledger's own rules. A claim is pending when its resolution date is after September 10, 2026, or when the corpus lacks the data to check it. The biggest claims in AI are also the longest-dated ones. AGI arrival, the payoff on trillions of dollars of compute spending, the humanoid robot timeline, Brett Adcock's line that in ten years "pretty close to every home will have a humanoid robot."
There is a third reason that is less mechanical. Several of the year's loudest questions are head-to-head disagreements between people who both have money riding on their own answer. Huang called the idea that AI reduces net employment "complete nonsense" on June 1, 2026. Amodei's essay says half of entry-level white-collar jobs. Ilya Sutskever, Sholto Douglas and Yann LeCun reached three different conclusions about whether scaling still works. When two credible people contradict each other and the record cannot adjudicate, I mark it pending or contested rather than manufacture a winner. That inflates the pending count, and it should.
Unfalsifiable: 16
This is the category people confuse with pending, and the difference matters. A pending claim will resolve someday. An unfalsifiable claim is built so it cannot.
Sam Altman wrote in January 2025 that OpenAI was "now confident we know how to build AGI as we have traditionally understood it." No definition of AGI, no benchmark, no date. Jeff Bezos told CNBC on May 20, 2026 that even if the AI investment boom is a bubble, "the bubble is driving investment" and that is fine. There is no future event that makes that false. Zuckerberg's August 10, 2026 essay argues that the core AI risk is concentration of control, not capability, which is a values position, not a prediction. None of these are worthless. They tell you what a person believes. They just cannot be scored as promises, and the trick of dressing a belief up as a prediction is the most common move in AI communication this year.
What an operator should do with a pending claim
Most people either round a pending claim up to confirmed because the speaker sounded confident, or down to refuted because they distrust the speaker. Both are how you end up building a plan on a claim with nothing behind it yet. Here is what I do instead.
Do not round it. Confidence is not evidence. Amodei's 50 percent figure and Huang's "complete nonsense" are both confident, and neither is proven.
Write down the specific event that resolves it and the date. Amodei's one-to-five-year window started in January 2026. Hassabis's date is 2030, give or take a year. Huang's trillion resolves against actuals through 2027. If you cannot name the resolving event, you are probably holding an unfalsifiable claim, not a pending one.
Note who benefits if you believe it. Amodei's company raised $30 billion at a $380 billion valuation in February 2026 partly on its credibility about advanced AI arriving soon. Huang's revenue depends on adoption continuing without friction. That does not make either of them wrong. It belongs in the file next to the claim.
Then actually check back. This is the step nobody does. I put twelve of the ledger's pending claims on a calendar for September 2027, and I re-score three of them at the six-month mark. A pending claim with a date attached is a task, not a conclusion.
Where the full ledger lives
All 132 claims, with the date, the source, the status and the note on each, sit in the back of The Claims Ledger, the book I wrote from this research. It covers fifty people across the window that closed on September 10, 2026, walks through the scoring method chapter by chapter, and ends with the twelve pending claims worth re-checking next year. It is $19, delivered as PDF and EPUB, at /book.