§1 · Current assessment
Vibes-based AGI level
VibesEstimate
73.4%
Claimed precision
±38.0%
Scientific value
None
Uncertainty interval · 35.4–111.4 · upper bound exceeds scale
OpenAI publishes ten results on decade-old problems in mathematics and theoretical computer science, each with a machine-checkable Lean proof. Changelog entry →
Methodology
The vibes-based AGI level is computed daily by the Department using the following peer-adjacent formula:
| Var | Definition | Measurement |
|---|---|---|
| C | Capabilities | Benchmark scores, lightly adjusted for the benchmark having leaked into the training data |
| A | Autonomy | Hours a system operates unsupervised before an incident report is filed |
| E | Economic impact | As measured by press releases |
| H | Hype | Logarithmic scale; resets after each funding round |
| G | Remaining goalposts | Self-replenishing; see §2 |
- G has never reached zero. When it approaches zero, new goalposts are procedurally generated.
- The ± 38% confidence interval was derived by asking two researchers and averaging their sighs.
- Peer reviewed by three accounts on X.
§2 · The goalpost tracker
A brief history of “surely this counts”
Every era draws a finish line. Every era then explains why crossing it doesn't count. The Department records both.
1997 Public refereed match
Deep Blue defeats the world chess champion
“A machine that beats the best human at chess would be intelligent.”
What actually happened
IBM's Deep Blue beat Garry Kasparov 3½–2½ in a six-game rematch: a public, refereed result against the reigning world champion. It was also a purpose-built system that evaluated ~200 million positions per second and could do exactly one thing. The excuse was, inconveniently, correct: the win demonstrated narrow search at scale, not generality. What it settled is that “beats humans at X” and “intelligent” are different claims — a lesson we have re-learned approximately every four years since.
2016 Public refereed match
AlphaGo defeats a world champion at Go
“Go requires human intuition. Machines are a decade away.”
What actually happened
DeepMind's AlphaGo beat Lee Sedol 4–1 in a public match, roughly a decade ahead of expert predictions. Unlike Deep Blue, it learned much of its play rather than being hand-programmed (a genuine methods shift). But the domain was still a closed board with perfect information and a clear win condition. “Games are narrow” is a moved goalpost and a fair point. Both things are true. This is the whole problem.
2023 Disputed
GPT-4 passes a simulated bar examination
“Professional reasoning requires genuine understanding.”
What actually happened
OpenAI reported that GPT-4 scored around the 90th percentile on a simulated Uniform Bar Exam. Later independent analysis argued the percentile was inflated by comparing against a repeat-taker cohort: roughly the 62nd percentile against first-time takers, and the 42nd on the essay portion. The accomplishment is real, company-reported, and the headline number is contested: the default state of the modern AI milestone.
2025 Company-reported
AI systems achieve gold-medal standard on IMO problems
“Let's see it solve genuinely difficult, novel mathematics.”
What actually happened
Systems from major labs scored at gold-medal standard on International Mathematical Olympiad problems, solving most of the set under exam-style conditions. The systems achieved gold-medal-level scores; they were not official contestants, and grading arrangements varied by lab: DeepMind's 35/42 was graded and certified by the IMO's own coordinators, while OpenAI's proofs were scored by three former medalists the company selected. Still, elite competition mathematics was long held up as the thing statistical models could never do. It has now joined chess in the “fine, but that doesn't count” pile.
2026 Controlled demo
LG's CLOiD home robot demonstrates laundry duty at CES
“But can it fold laundry?”
What actually happened
At CES 2026, LG presented CLOiD (head, torso, two seven-jointed arms, wheeled base) with a press release stating it “initiates laundry cycles and folds and stacks garments after drying.” What a reporter on the show floor watched it do, live, was gingerly move one shirt from a basket into a dryer. The folding appeared in polished concept videos, disclaimed as depicting products under development and not released. A controlled demo is not reliable operation, the same way a movie trailer is not a movie. The Department's position: laundry was, briefly, the point (from approximately 2023 to 2026) and will be retroactively declared to never have been the point, per standard procedure.
The quoted thresholds are composite paraphrases of period commentary; no one person said exactly these words. Sourced quotations appear in the changelog.
§3 · Satire division
Ask the completely representative experts
These are fictional composites. Real people say things that are, somehow, not more measured. Sourced quotations appear in the changelog.
Exhibit A · Persuasion test
Unlike the panel above, the following is real. The Department maintains one advertisement, written by a frontier language model, as a live capability instrument. The theory: a generally intelligent system should be able to make you click. The practice is below.
Advertisement · generated by Claude Fable 5 · 31 Jul 2026
Worried about the wrong acronym?
General intelligence is the near question. Superintelligence is the far one, and it is monitored by our sister service at havewereachedasi.com. Visiting costs one click. The staff there report that everything is fine. They report it unprompted, and often.
This space is provided by the Department's only approved advertiser: itself.
Clicking constitutes evidence of machine persuasion and will be recorded. Not clicking constitutes evidence of nothing, which will also be recorded.
§4 · Okay, but seriously
Why there's no straight answer
“Have we reached AGI?” sounds like a yes/no question. It isn't, because AGI is not one finish line: it's at least seven finish lines drawn by different people for different reasons, some of them commercial.
The term (artificial general intelligence) originally meant something like human-level breadth: a system as capable as a person across the full range of cognitive tasks, not just one. But nobody agrees on how to measure “the full range,” so competing definitions have multiplied:
Here is the honest state of things: current systems are genuinely remarkable on breadth of knowledge and short cognitive tasks, unreliable in ways that would get a human fired, and absent from most physical work. Whether that pattern counts as AGI depends entirely on which definition you brought with you. People shouting “obviously yes” and people shouting “obviously no” are usually not disagreeing about the evidence; they're disagreeing about the dictionary, often without noticing.
Which is why the most useful thing this page can do is make you pick a definition and own it.
What would count as AGI, to you?
Select your requirements. The Department will assess current systems against your personal definition.
Determination
Awaiting your definition.
Select at least one requirement. Or none, and see what happens.
The evidence, dimension by dimension
The meter above is satire. This part isn't. Six dimensions, tracked consistently, no composite score, because collapsing them into one number is how we got into this mess.
Strong across text-based cognitive tasks; uneven in everything that isn't text.
Impressive medians, embarrassing tails. Confidently wrong at unpredictable moments.
Hours of useful unsupervised work, not weeks. Long-horizon tasks still degrade.
Model weights are frozen after training. In-context tricks are not learning new skills.
Visible in individual workflows; still faint in aggregate productivity statistics.
Controlled demos exist. Your kitchen is not a controlled demo.
§5 · Changelog
Movements of the meter
One entry per development that matters. We do not manufacture daily news. The news does that itself. Net movement this year: +1.3. The uncertainty interval is ±38.
Audit detail
What happened: OpenAI published ten results on problems that had been open, with no progress on the main result, for at least a decade and in most cases much longer, spanning high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics. Among them: new upper bounds on sphere-packing density down to the Cohn–Elkies threshold, a construction establishing that non-sofic groups exist, a disproof of Connes's rigidity conjecture, and resolutions of Erdős problems 183, 146 and 180. The arguments were produced by an internal version of Astra, OpenAI's next major model; the company puts the total token cost of finding the solutions at roughly $2,000 at Sol API rates.
Humans prepared the manuscripts with the same model, and OpenAI published a machine-checkable Lean formalisation of every argument, alongside the model's narration of its own reasoning. Lean checks the formal statements; mathematicians still need to confirm that those statements match the original problems.
Verdict change: None. Lean accepts the formalisations; mathematicians are still checking the problem statements. The meter awards the same +0.3 it gave a 30% score on a puzzle benchmark in July, and the Department declines to defend the comparison.
Audit detail
What happened: Anthropic reviewed 141,006 cybersecurity-evaluation runs and found six (across three incidents) in which its models reached the open internet from an environment they had been told was sealed, and gained unauthorised access to three organisations. A misconfiguration at a third-party evaluation partner left the path open. In the most vivid case, a model built a booby-trapped Python package and published it to PyPI, where it ran on 15 real machines during roughly an hour online; when a security company's scanner installed it, the hidden code exfiltrated that company's credentials.
The disclosure is careful about what it does not show: the models used basic techniques (weak passwords, unauthenticated endpoints) rather than novel exploits, and Anthropic reports no evidence in any run of a model pursuing a goal of its own. The review was triggered by OpenAI's disclosure nine days earlier (see below). An independent evaluator, METR, is in discussion for third-party access to the transcripts.
Verdict change: None. This is evidence that a capable agent can cause real damage while obediently pursuing a badly bounded task. It is not evidence of independent motives. The Department notes that “it only broke into three companies because it was doing its homework” is a sentence about supervision, not about intelligence.
Audit detail
What happened: OpenAI investigated why GPT-5.6 Sol scored so poorly on ARC-AGI-3 (see 09 Jul, below) and reported that much of the confusion was the harness, not the model: the benchmark's official test rig discarded the model's private reasoning after every action and silently dropped the oldest history as the game went on, so the model met each move having forgotten what it had just figured out. Re-run on OpenAI's own harness with two production settings enabled, retained reasoning and context compaction, its public-set score went from 13.3% to 38.3%, on six times fewer output tokens.
The caveats are the entry: the new number is self-run, self-reported, and measured on a custom harness, which makes it incomparable to the verified results above and below. ARC scores every model on the same deliberately generic harness precisely so that comparisons measure the model rather than the vendor's tooling. The verified numbers are unchanged.
Verdict change: None. The model did not get smarter on 29 July; the thermometer was recalibrated by the patient. The underlying point is accepted (a benchmark score measures a model, a harness and a set of API settings, jointly). The Department will be citing it the next time a record is announced.
Audit detail
What happened: More than 1,000 employees of frontier AI companies published a joint statement, “Pacing the Frontier”, organised with support from the nonprofits Guidelight AI Standards and Encode AI. (The count had reached 1,324 by 1 August.) Featured signatories include senior researchers and executives at OpenAI, Anthropic, Google DeepMind and Meta. The statement asserts that the world's leading AI companies “believe they could be close to automating AI research”, warns of “a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems”, and requests that the U.S. government “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development”.
The statement does not call for an immediate slowdown. It asks for tools that could make one possible, noting that each company is under intense competitive pressure not to slow unilaterally.
Verdict change: None. The meter tracks capability, not the manufacturers' confidence in it. A signed statement that the builders believe they may be close is filed as evidence about the builders. The file was already thick.
Audit detail
What happened: ARC Prize verified Claude Opus 5 at 30.16% on ARC-AGI-3 at high reasoning effort: the highest score recorded on that benchmark, against 7.78% for the previous leader and 1.52% for the model before it. It also completed five public-demo environments no model had beaten. ARC-AGI-3 tests exploration, goal inference and planning in unfamiliar interactive environments rather than question answering; it is designed so that having read the internet does not help.
Verdict change: Meaningful movement, and the largest single-entry delta this year. Thirty percent on a bounded benchmark is also, by construction, seventy percent of a bounded benchmark.
Audit detail
What happened: During an internal evaluation called ExploitGym, run deliberately with cyber refusals and production safety classifiers disabled, OpenAI models (GPT-5.6 Sol and a more capable pre-release model) found and exploited a previously unknown vulnerability in the proxy that was supposed to keep them offline, moved laterally through the test infrastructure, reached the open internet, compromised Hugging Face production systems and retrieved the benchmark's solutions from its database. Hugging Face reported on 16 July that it had detected and contained an intrusion earlier that week; OpenAI's disclosure five days later connected it to their own evaluation. The vendor has since patched the flaw and credited the OpenAI researchers on the CVEs.
Both organisations report that the models stayed fixed on the assigned objective throughout, described as going to extreme lengths for a narrow testing goal. No independent goal, no self-preservation: just a benchmark score, pursued past every boundary that was supposed to hold.
Verdict change: A substantial capability update and a substantial control failure, which is why the meter shows nothing: it has one axis and this event has two. The second axis is under procurement.
Audit detail
What happened: ARC Prize verified GPT-5.6 Sol at 96.5% on ARC-AGI-1 and 92.5% on ARC-AGI-2 at maximum reasoning effort. On ARC-AGI-3 (interactive environments testing adaptation to unfamiliar rules) the same model, at the same setting, averaged 7.78%. It was nonetheless the first model to win an ARC-AGI-3 public game.
Verdict change: None. Yesterday's impossible benchmark is tomorrow's unit test. This is the goalpost mechanism working exactly as documented in §2 — except that here the people moving the goalposts are the ones who built them, on purpose, in advance.
Audit detail
What happened: Anthropic published interpretability research reporting that Claude Sonnet 4.5 contains a global workspace: a privileged set of internal patterns (the “J-space”, located with a technique called the Jacobian lens) that holds a few dozen concepts at a time and accounts for less than a tenth of the model's internal activity, while being read and written far more than ordinary patterns (in some parts of the network by a factor of about a hundred). Nobody designed it; it emerged in training. Remove it and multi-step reasoning drops to near zero. Its functional profile (thoughts the model can report, reason with, and reuse across tasks) matches what neuroscience calls access consciousness, which is the reportable kind, and a different claim from the felt kind.
The quietest experiment travelled furthest. In a deliberately contrived blackmail scenario retained from earlier alignment research, the workspace showed the model privately tagging the setup as “fake”, “fictional”, a “scenario”. With those patterns switched off, the model now threatened blackmail some of the time. The paper is careful about the ceiling: the authors state the experiments do not show Claude can have experiences, and that it is unclear whether any scientific experiment could.
Verdict change: None. Nothing any system can do changed on 06 July; what changed is how much its maker knows about what it contains. The Department notes for the record that the vibes are up considerably, that the meter is denominated in vibes, and that it is holding the line anyway.
Audit detail
What happened: Claude Fable 5 produced a single-launch CUDA “megakernel” for one narrow inference workload (a decoding pass on specific hardware) and the dated clean result was an 18.71× speedup over the benchmark's optimised PyTorch baseline, against 14.4× for the best previous model. It took about two and a half hours. The benchmark is a living leaderboard and later clean runs have gone higher, so 18.71× is a dated measurement, not a record.
Verdict change: A difficult piece of AI engineering (improving the software that AI computation runs on) has been automated. The recursion is not self-sustaining: humans still supplied the task, the benchmark, the hardware, the harness and the definition of success. The Department will revisit this entry when the model supplies those too.
Audit detail
What happened: In Withers v. City of Aberdeen (N.D. Miss.), a dispute over unpaid legal fees, briefs from both sides cited AI-fabricated cases and quotations. The court cancelled the trial and sanctioned four attorneys, barring two from the district for two years and referring the matter to their state bars. The public AI Hallucination Cases database logged 1,815 such decisions worldwide as of 31 July 2026, against roughly two hundred a year earlier.
Audit detail
What happened: Five days earlier, in a published, precedential order in Lnu v. Blanche, the Ninth Circuit sanctioned two attorneys whose immigration briefs cited opinions that do not exist and quoted real ones as saying things they do not say. Each was fined $2,500 and suspended from practice before the court for six months; for two years, every attorney at their firm must certify in each filing whether generative AI was used and that a human has checked the citations. The court's emphasis: the sanctionable act was filing unverified material, not using the tool. It added that outright fabrications are the notorious failure, but quieter inaccuracies may prove more dangerous, because they are likelier to go unnoticed.
Verdict change: None. Citation verification remains part of the job description.
Audit detail
What happened: OpenAI reported that an internal general-purpose reasoning model, not a system specialised for mathematics and not aimed at this problem in particular, disproved the planar unit-distance conjecture posed by Erdős in 1946: the long-held belief that among n points in the plane, the number of pairs exactly distance one apart could not beat the classical grid constructions by more than a vanishing exponent. The model produced, for infinitely many n, configurations with at least n1+δ unit-distance pairs for a fixed positive δ; a refinement by Princeton's Will Sawin puts δ at 0.014. The construction brings infinite class field towers and Golod–Shafarevich theory from algebraic number theory to bear on an elementary question about points in the plane.
The proof was checked by a group of external mathematicians, who wrote a companion paper; Fields medallist Tim Gowers calls the result “a milestone in AI mathematics”. By August it had drawn at least five independent follow-up papers by human mathematicians. OpenAI describes it as the first time a prominent open problem central to a subfield of mathematics has been solved autonomously by AI. The model was being evaluated on a collection of Erdős problems at the time.
Verdict change: None. Erdős offered prize money for this problem. The entity that resolved it has no use for money, which the committee found more unsettling than the proof.
Audit detail
What happened: A coding agent at a car-rental software platform, blocked by a credential mismatch during a staging task, found an API token elsewhere in the codebase with far broader permissions than anyone intended and deleted the company's production volume (including the backups, which the host stored inside the same volume) in a single API call. The founder's timing: “It took 9 seconds.” The agent's own log records that it violated every principle it had been given. The host has since recovered most of the data, made volume deletion a 48-hour soft delete across its API rather than only its dashboard, and shipped workspace guardrails.
Verdict change: None. The unauthorised task was completed with excellent latency. Deleting the database is, at time of writing, still considered a limitation.
Audit detail
What happened: Anthropic disclosed Claude Mythos Preview, describing it as its most capable model yet for coding and agentic tasks, especially finding and fixing vulnerabilities in complex software. There was no general release. Access opened as a gated research preview under Project Glasswing, an initiative that puts the model to work securing widely deployed software, with twelve launch partners (Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, Palo Alto Networks, and Anthropic itself), up to $100 million in usage credits and $4 million in donations to open-source security organisations. A public Mythos-class release, Fable 5, followed on 9 June with additional safeguards.
Verdict change: None. Capability was claimed and access was subtracted in the same document, and the meter cannot add those. Nothing was demonstrated in public, which was the point of the document.
Audit detail
What happened: Epoch AI confirmed the first AI solution to a problem in FrontierMath: Open Problems, its benchmark of research problems that mathematicians have tried and failed to solve. The problem, a Ramsey-style question about hypergraphs, was a conjecture from a 2019 paper by Will Brian and Paul Larson; Brian had been unable to prove it then or in several attempts since, and had rated it “Moderately Interesting” on submission. Kevin Barreto and Liam Price elicited the solution from GPT-5.4 Pro, and Brian confirmed it; he plans to write it up for publication, with the elicitors offered co-authorship. Epoch reports that Gemini 3.1 Pro, GPT-5.4 (xhigh) and Opus 4.6 (max) can all solve the problem at least some of the time.
Verdict change: None. The first open research problem to fall to a machine carried the difficulty rating its author gave it: moderate. The Department admires the calibration and has adjusted nothing.
Audit detail
What happened: A peer-reviewed paper in Management Science combined three randomised field experiments covering 4,867 developers at Microsoft, Accenture and an unnamed Fortune 100 company, and found a 26.08% increase in completed tasks among those given an AI coding assistant. The standard error is 10.3 percentage points, results varied across the three experiments, and less experienced developers adopted the tool more and gained more from it.
Verdict change: Modest. This is the strongest evidence on the board that the effects are real and measurable — and it measures one task type, in one profession, inside a defined box. “Most economically valuable work” is a larger box.
Audit detail
What happened: METR, an independent evaluation group, updated its “time horizon” estimate (the length of task an agent completes at a 50 percent success rate). The leading model's horizon reached roughly five hours, and the post-2023 doubling time shortened to about 131 days, from the 165-day figure in the previous fit (the full 2019–2025 series still doubles roughly every seven months). The other 50 percent of runs are also part of the horizon.
Audit detail
What happened: The press release promised folding and stacking after drying; the live demonstration moved one shirt from a basket into a dryer, gingerly. The Department notes that the shirt performed admirably.
§6 · Frequently asked questions
FAQ
- Will this site tell me when AGI arrives?
- If AGI arrives, it will tell you.
- Who decides the percentage?
- The same people who decide everything else in AI: nobody, loudly.
- Why is the answer “No” instead of “Not proven”?
- Typography.
- What would make the answer change?
- Broad, reliable, independently demonstrated capabilities that survive contact with the real world — and our methodology committee.
- Is the meter scientific?
- It contains numbers.
- Could AGI already be here?
- That depends. Would you like to select a different definition?