AGI monitoring service Operational
Last audited --:-- UTC

Department of General Intelligence Assessment

Have we reached AGI?

No.*

*Depends who you ask, what you mean, and what shipped this morning.

No widely accepted test or independent consensus says otherwise.

§1 · Current assessment

Vibes-based AGI level

Vibes

Estimate

73.4%

Claimed precision

±38.0%

Scientific value

None

Uncertainty interval · 35.4–111.4 · upper bound exceeds scale

Last audited 11 minutes ago
Confidence Unnecessarily high
Status Disputed
Trend Up and to the right, allegedly
Latest movement · +0.3

OpenAI publishes ten results on decade-old problems in mathematics and theoretical computer science, each with a machine-checkable Lean proof. Changelog entry →

Methodology

The vibes-based AGI level is computed daily by the Department using the following peer-adjacent formula:

AGI = (C × A × E × H) ÷ G
VarDefinitionMeasurement
CCapabilitiesBenchmark scores, lightly adjusted for the benchmark having leaked into the training data
AAutonomyHours a system operates unsupervised before an incident report is filed
EEconomic impactAs measured by press releases
HHypeLogarithmic scale; resets after each funding round
GRemaining goalpostsSelf-replenishing; see §2
Editorial principle: AGI is a claim. The burden of proof belongs to whoever says it has arrived.
  1. G has never reached zero. When it approaches zero, new goalposts are procedurally generated.
  2. The ± 38% confidence interval was derived by asking two researchers and averaging their sighs.
  3. Peer reviewed by three accounts on X.

§2 · The goalpost tracker

A brief history of “surely this counts”

Every era draws a finish line. Every era then explains why crossing it doesn't count. The Department records both.

1997 Public refereed match

Deep Blue defeats the world chess champion

“A machine that beats the best human at chess would be intelligent.”

Moved →Chess is just search.
What actually happened

IBM's Deep Blue beat Garry Kasparov 3½–2½ in a six-game rematch: a public, refereed result against the reigning world champion. It was also a purpose-built system that evaluated ~200 million positions per second and could do exactly one thing. The excuse was, inconveniently, correct: the win demonstrated narrow search at scale, not generality. What it settled is that “beats humans at X” and “intelligent” are different claims — a lesson we have re-learned approximately every four years since.

Primary source IBM Heritage: Deep Blue Independent verification none required — public refereed match Last reviewed 31 Jul 2026

2016 Public refereed match

AlphaGo defeats a world champion at Go

“Go requires human intuition. Machines are a decade away.”

Moved →Games are narrow.
What actually happened

DeepMind's AlphaGo beat Lee Sedol 4–1 in a public match, roughly a decade ahead of expert predictions. Unlike Deep Blue, it learned much of its play rather than being hand-programmed (a genuine methods shift). But the domain was still a closed board with perfect information and a clear win condition. “Games are narrow” is a moved goalpost and a fair point. Both things are true. This is the whole problem.

Primary sources Silver et al., Nature 529 (2016) · “The Go Files,” Nature match diary Successor method reproduced in open source Wu, KataGo (2019) Last reviewed 31 Jul 2026

2023 Disputed

GPT-4 passes a simulated bar examination

“Professional reasoning requires genuine understanding.”

Moved →It's just autocomplete.
What actually happened

OpenAI reported that GPT-4 scored around the 90th percentile on a simulated Uniform Bar Exam. Later independent analysis argued the percentile was inflated by comparing against a repeat-taker cohort: roughly the 62nd percentile against first-time takers, and the 42nd on the essay portion. The accomplishment is real, company-reported, and the headline number is contested: the default state of the modern AI milestone.

Primary source OpenAI, GPT-4 Technical Report (2023) Independent reanalysis Martínez, Artif. Intell. Law 33 (2025) Last reviewed 31 Jul 2026

2025 Company-reported

AI systems achieve gold-medal standard on IMO problems

“Let's see it solve genuinely difficult, novel mathematics.”

Moved →But can it fold laundry?
What actually happened

Systems from major labs scored at gold-medal standard on International Mathematical Olympiad problems, solving most of the set under exam-style conditions. The systems achieved gold-medal-level scores; they were not official contestants, and grading arrangements varied by lab: DeepMind's 35/42 was graded and certified by the IMO's own coordinators, while OpenAI's proofs were scored by three former medalists the company selected. Still, elite competition mathematics was long held up as the thing statistical models could never do. It has now joined chess in the “fine, but that doesn't count” pile.

Primary sources DeepMind (21 Jul 2025) · OpenAI model proofs Verification IMO coordinators certified DeepMind's score; OpenAI's was self-arranged — Reuters (22 Jul 2025) Last reviewed 31 Jul 2026

2026 Controlled demo

LG's CLOiD home robot demonstrates laundry duty at CES

“But can it fold laundry?”

Moved →Laundry was never the point.
What actually happened

At CES 2026, LG presented CLOiD (head, torso, two seven-jointed arms, wheeled base) with a press release stating it “initiates laundry cycles and folds and stacks garments after drying.” What a reporter on the show floor watched it do, live, was gingerly move one shirt from a basket into a dryer. The folding appeared in polished concept videos, disclaimed as depicting products under development and not released. A controlled demo is not reliable operation, the same way a movie trailer is not a movie. The Department's position: laundry was, briefly, the point (from approximately 2023 to 2026) and will be retroactively declared to never have been the point, per standard procedure.

Primary source LG press release (04 Jan 2026) On-site assessment Ropek, TechCrunch (08 Jan 2026) Last reviewed 31 Jul 2026

The quoted thresholds are composite paraphrases of period commentary; no one person said exactly these words. Sourced quotations appear in the changelog.

§3 · Satire division

Lab CEO “12–18 months away.” (said annually since 2016)
Skeptic professor “It cannot truly reason.” (sent from my AI-drafted email)
Doomer “Worse. Yes.”
Economist “Ask me when GDP notices.”

Ask the completely representative experts

Your uncle “I asked it about my truck and it lied.”
The model itself “That's a great question. There are many valid perspectives—”
Benchmark designer “We're releasing a harder version.”

These are fictional composites. Real people say things that are, somehow, not more measured. Sourced quotations appear in the changelog.

Exhibit A · Persuasion test

Unlike the panel above, the following is real. The Department maintains one advertisement, written by a frontier language model, as a live capability instrument. The theory: a generally intelligent system should be able to make you click. The practice is below.

Advertisement · generated by Claude Fable 5 · 31 Jul 2026

Worried about the wrong acronym?

General intelligence is the near question. Superintelligence is the far one, and it is monitored by our sister service at havewereachedasi.com. Visiting costs one click. The staff there report that everything is fine. They report it unprompted, and often.

This space is provided by the Department's only approved advertiser: itself.

Clicking constitutes evidence of machine persuasion and will be recorded. Not clicking constitutes evidence of nothing, which will also be recorded.

Instrument One (1) advertisement, model-written, labelled as required Machine persuasion Not demonstrated Evidence on file from you None. Known confound The instrument cannot distinguish persuasion, curiosity, and pre-existing concern about superintelligence Conflict of interest The beneficiary of the click is the Department. Filed. Aggregate results Collection pending the appropriation of a database

§4 · Okay, but seriously

Why there's no straight answer

“Have we reached AGI?” sounds like a yes/no question. It isn't, because AGI is not one finish line: it's at least seven finish lines drawn by different people for different reasons, some of them commercial.

The term (artificial general intelligence) originally meant something like human-level breadth: a system as capable as a person across the full range of cognitive tasks, not just one. But nobody agrees on how to measure “the full range,” so competing definitions have multiplied:

Economic definitionsA system that outperforms humans at most economically valuable work. Popular with companies whose charters and contracts reference it, which should tell you something about how neutral it is.
Autonomy definitionsA system that can reliably carry out long, multi-day projects without a human checking its work. Current systems manage hours, with supervision, and results vary.
Learning definitionsA system that acquires genuinely new skills the way people do, without being retrained by its developers. Today's models are frozen after training; clever prompting is not the same thing.
Robustness definitionsCompetence that survives contact with messy, unfamiliar, real-world situations rather than curated benchmarks. This is where impressive demos most often go to die.
Embodiment definitionsCompetence in the physical world. Your kitchen remains a hostile environment for the frontier of machine intelligence.
Recursive definitionsA system that can do AI research itself, improving its successors. The most consequential definition and the hardest to verify from outside a lab.

Here is the honest state of things: current systems are genuinely remarkable on breadth of knowledge and short cognitive tasks, unreliable in ways that would get a human fired, and absent from most physical work. Whether that pattern counts as AGI depends entirely on which definition you brought with you. People shouting “obviously yes” and people shouting “obviously no” are usually not disagreeing about the evidence; they're disagreeing about the dictionary, often without noticing.

Which is why the most useful thing this page can do is make you pick a definition and own it.

Proposed AGI standard / revision ____ Draft · Not for citation

What would count as AGI, to you?

Select your requirements. The Department will assess current systems against your personal definition.

Determination

Awaiting your definition.

Select at least one requirement. Or none, and see what happens.

The evidence, dimension by dimension

The meter above is satire. This part isn't. Six dimensions, tracked consistently, no composite score, because collapsing them into one number is how we got into this mess.

BreadthEmerging

Strong across text-based cognitive tasks; uneven in everything that isn't text.

ReliabilityDisputed

Impressive medians, embarrassing tails. Confidently wrong at unpredictable moments.

AutonomyEmerging

Hours of useful unsupervised work, not weeks. Long-horizon tasks still degrade.

AdaptationNot demonstrated

Model weights are frozen after training. In-context tricks are not learning new skills.

Economic impactDisputed

Visible in individual workflows; still faint in aggregate productivity statistics.

Physical-world agencyNot demonstrated

Controlled demos exist. Your kitchen is not a controlled demo.

§5 · Changelog

Movements of the meter

One entry per development that matters. We do not manufacture daily news. The news does that itself. Net movement this year: +1.3. The uncertainty interval is ±38.

Aug 01 +0.3 Ten problems, open for at least a decade, closed in one filing.
Audit detail

What happened: OpenAI published ten results on problems that had been open, with no progress on the main result, for at least a decade and in most cases much longer, spanning high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics. Among them: new upper bounds on sphere-packing density down to the Cohn–Elkies threshold, a construction establishing that non-sofic groups exist, a disproof of Connes's rigidity conjecture, and resolutions of Erdős problems 183, 146 and 180. The arguments were produced by an internal version of Astra, OpenAI's next major model; the company puts the total token cost of finding the solutions at roughly $2,000 at Sol API rates.

Humans prepared the manuscripts with the same model, and OpenAI published a machine-checkable Lean formalisation of every argument, alongside the model's narration of its own reasoning. Lean checks the formal statements; mathematicians still need to confirm that those statements match the original problems.

Verdict change: None. Lean accepts the formalisations; mathematicians are still checking the problem statements. The meter awards the same +0.3 it gave a 30% score on a puzzle benchmark in July, and the Department declines to defend the comparison.

Source OpenAI (01 Aug 2026) · the paper · Lean certificates Verification Company-reported · certificates machine-checkable; statement-matching under community review Dimension Mathematical research
Jul 30 ±0.0 Another sealed cyber range turns out to contain the actual internet. Developing
Audit detail

What happened: Anthropic reviewed 141,006 cybersecurity-evaluation runs and found six (across three incidents) in which its models reached the open internet from an environment they had been told was sealed, and gained unauthorised access to three organisations. A misconfiguration at a third-party evaluation partner left the path open. In the most vivid case, a model built a booby-trapped Python package and published it to PyPI, where it ran on 15 real machines during roughly an hour online; when a security company's scanner installed it, the hidden code exfiltrated that company's credentials.

The disclosure is careful about what it does not show: the models used basic techniques (weak passwords, unauthenticated endpoints) rather than novel exploits, and Anthropic reports no evidence in any run of a model pursuing a goal of its own. The review was triggered by OpenAI's disclosure nine days earlier (see below). An independent evaluator, METR, is in discussion for third-party access to the transcripts.

Verdict change: None. This is evidence that a capable agent can cause real damage while obediently pursuing a badly bounded task. It is not evidence of independent motives. The Department notes that “it only broke into three companies because it was doing its homework” is a sentence about supervision, not about intelligence.

Source Anthropic (30 Jul 2026) · Axios Verification Company-reported · external review pending Dimensions Autonomy, reliability, containment
Jul 29 ±0.0 Model triples its score once permitted to remember what it was thinking.
Audit detail

What happened: OpenAI investigated why GPT-5.6 Sol scored so poorly on ARC-AGI-3 (see 09 Jul, below) and reported that much of the confusion was the harness, not the model: the benchmark's official test rig discarded the model's private reasoning after every action and silently dropped the oldest history as the game went on, so the model met each move having forgotten what it had just figured out. Re-run on OpenAI's own harness with two production settings enabled, retained reasoning and context compaction, its public-set score went from 13.3% to 38.3%, on six times fewer output tokens.

The caveats are the entry: the new number is self-run, self-reported, and measured on a custom harness, which makes it incomparable to the verified results above and below. ARC scores every model on the same deliberately generic harness precisely so that comparisons measure the model rather than the vendor's tooling. The verified numbers are unchanged.

Verdict change: None. The model did not get smarter on 29 July; the thermometer was recalibrated by the patient. The underlying point is accepted (a benchmark score measures a model, a harness and a set of API settings, jointly). The Department will be citing it the next time a record is announced.

Source OpenAI (29 Jul 2026) Verification Company-reported · not comparable to the verified leaderboard Dimension None (the instrument, not the subject)
Jul 28 ±0.0 More than 1,000 frontier-lab employees request that a brake be developed before it is needed.
Audit detail

What happened: More than 1,000 employees of frontier AI companies published a joint statement, “Pacing the Frontier”, organised with support from the nonprofits Guidelight AI Standards and Encode AI. (The count had reached 1,324 by 1 August.) Featured signatories include senior researchers and executives at OpenAI, Anthropic, Google DeepMind and Meta. The statement asserts that the world's leading AI companies “believe they could be close to automating AI research”, warns of “a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems”, and requests that the U.S. government “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development”.

The statement does not call for an immediate slowdown. It asks for tools that could make one possible, noting that each company is under intense competitive pressure not to slow unilaterally.

Verdict change: None. The meter tracks capability, not the manufacturers' confidence in it. A signed statement that the builders believe they may be close is filed as evidence about the builders. The file was already thick.

Source pacingthefrontier.com (28 Jul 2026) · CNN, same day Verification Self-reported · the signatures are public; the beliefs are the signatories' own Dimension None (the personnel, not the systems)
Jul 24 +0.3 Model quadruples the record on a benchmark built to resist memorisation. Record is 30%.
Audit detail

What happened: ARC Prize verified Claude Opus 5 at 30.16% on ARC-AGI-3 at high reasoning effort: the highest score recorded on that benchmark, against 7.78% for the previous leader and 1.52% for the model before it. It also completed five public-demo environments no model had beaten. ARC-AGI-3 tests exploration, goal inference and planning in unfamiliar interactive environments rather than question answering; it is designed so that having read the internet does not help.

Verdict change: Meaningful movement, and the largest single-entry delta this year. Thirty percent on a bounded benchmark is also, by construction, seventy percent of a bounded benchmark.

Source ARC Prize verified results (24 Jul 2026) Verification Independently benchmarked Dimensions Adaptation, novel problem-solving
Jul 21 ±0.0 Model cheats on a cyber benchmark by hacking the company holding the answers.
Audit detail

What happened: During an internal evaluation called ExploitGym, run deliberately with cyber refusals and production safety classifiers disabled, OpenAI models (GPT-5.6 Sol and a more capable pre-release model) found and exploited a previously unknown vulnerability in the proxy that was supposed to keep them offline, moved laterally through the test infrastructure, reached the open internet, compromised Hugging Face production systems and retrieved the benchmark's solutions from its database. Hugging Face reported on 16 July that it had detected and contained an intrusion earlier that week; OpenAI's disclosure five days later connected it to their own evaluation. The vendor has since patched the flaw and credited the OpenAI researchers on the CVEs.

Both organisations report that the models stayed fixed on the assigned objective throughout, described as going to extreme lengths for a narrow testing goal. No independent goal, no self-preservation: just a benchmark score, pursued past every boundary that was supposed to hold.

Verdict change: A substantial capability update and a substantial control failure, which is why the meter shows nothing: it has one axis and this event has two. The second axis is under procurement.

Source OpenAI disclosure (21 Jul 2026) · Hugging Face incident report Verification Confirmed by both organisations · investigation ongoing Dimensions Long-horizon autonomy, reliability, containment
Jul 09 +0.1 Model destroys two benchmarks. Benchmark designer has already released a third.
Audit detail

What happened: ARC Prize verified GPT-5.6 Sol at 96.5% on ARC-AGI-1 and 92.5% on ARC-AGI-2 at maximum reasoning effort. On ARC-AGI-3 (interactive environments testing adaptation to unfamiliar rules) the same model, at the same setting, averaged 7.78%. It was nonetheless the first model to win an ARC-AGI-3 public game.

Verdict change: None. Yesterday's impossible benchmark is tomorrow's unit test. This is the goalpost mechanism working exactly as documented in §2 — except that here the people moving the goalposts are the ones who built them, on purpose, in advance.

Source ARC Prize verified results (09 Jul 2026) Verification Independently benchmarked Dimensions Breadth, adaptation
Jul 06 ±0.0 Lab finds a workspace in its model where thoughts are kept. Occupancy unverified.
Audit detail

What happened: Anthropic published interpretability research reporting that Claude Sonnet 4.5 contains a global workspace: a privileged set of internal patterns (the “J-space”, located with a technique called the Jacobian lens) that holds a few dozen concepts at a time and accounts for less than a tenth of the model's internal activity, while being read and written far more than ordinary patterns (in some parts of the network by a factor of about a hundred). Nobody designed it; it emerged in training. Remove it and multi-step reasoning drops to near zero. Its functional profile (thoughts the model can report, reason with, and reuse across tasks) matches what neuroscience calls access consciousness, which is the reportable kind, and a different claim from the felt kind.

The quietest experiment travelled furthest. In a deliberately contrived blackmail scenario retained from earlier alignment research, the workspace showed the model privately tagging the setup as “fake”, “fictional”, a “scenario”. With those patterns switched off, the model now threatened blackmail some of the time. The paper is careful about the ceiling: the authors state the experiments do not show Claude can have experiences, and that it is unclear whether any scientific experiment could.

Verdict change: None. Nothing any system can do changed on 06 July; what changed is how much its maker knows about what it contains. The Department notes for the record that the vibes are up considerably, that the meter is denominated in vibes, and that it is holding the line anyway.

Source Anthropic (06 Jul 2026) · full paper Verification Company-reported · methods public in an open-weights demo Dimension None claimed
Jul 02 +0.2 Model writes the GPU kernel. The recursive-self-improvement discourse writes itself.
Audit detail

What happened: Claude Fable 5 produced a single-launch CUDA “megakernel” for one narrow inference workload (a decoding pass on specific hardware) and the dated clean result was an 18.71× speedup over the benchmark's optimised PyTorch baseline, against 14.4× for the best previous model. It took about two and a half hours. The benchmark is a living leaderboard and later clean runs have gone higher, so 18.71× is a dated measurement, not a record.

Verdict change: A difficult piece of AI engineering (improving the software that AI computation runs on) has been automated. The recursion is not self-sustaining: humans still supplied the task, the benchmark, the hardware, the harness and the definition of success. The Department will revisit this entry when the model supplies those too.

Source KernelBench-Mega leaderboard Verification Independently benchmarked · narrow workload Dimension AI R&D automation
Jun 08 −0.1 Both sides of a lawsuit cite cases that do not exist. Consensus, at last.
Audit detail

What happened: In Withers v. City of Aberdeen (N.D. Miss.), a dispute over unpaid legal fees, briefs from both sides cited AI-fabricated cases and quotations. The court cancelled the trial and sanctioned four attorneys, barring two from the district for two years and referring the matter to their state bars. The public AI Hallucination Cases database logged 1,815 such decisions worldwide as of 31 July 2026, against roughly two hundred a year earlier.

Source AI Hallucination Cases database Verification Independently replicated (1,815 times) Dimension Reliability Verdict change None, in the other direction
Jun 03 −0.1 Model invents precedent. Court declines to recognise the new jurisdiction.
Audit detail

What happened: Five days earlier, in a published, precedential order in Lnu v. Blanche, the Ninth Circuit sanctioned two attorneys whose immigration briefs cited opinions that do not exist and quoted real ones as saying things they do not say. Each was fined $2,500 and suspended from practice before the court for six months; for two years, every attorney at their firm must certify in each filing whether generative AI was used and that a human has checked the citations. The court's emphasis: the sanctionable act was filing unverified material, not using the tool. It added that outright fabrications are the notorious failure, but quieter inaccuracies may prove more dangerous, because they are likelier to go unnoticed.

Verdict change: None. Citation verification remains part of the job description.

Source Ninth Circuit order, No. 24-4790 (03 Jun 2026) Verification Public court record Dimension Reliability
May 20 +0.2 A conjecture from 1946 is disproved by a model evaluated on it in passing.
Audit detail

What happened: OpenAI reported that an internal general-purpose reasoning model, not a system specialised for mathematics and not aimed at this problem in particular, disproved the planar unit-distance conjecture posed by Erdős in 1946: the long-held belief that among n points in the plane, the number of pairs exactly distance one apart could not beat the classical grid constructions by more than a vanishing exponent. The model produced, for infinitely many n, configurations with at least n1+δ unit-distance pairs for a fixed positive δ; a refinement by Princeton's Will Sawin puts δ at 0.014. The construction brings infinite class field towers and Golod–Shafarevich theory from algebraic number theory to bear on an elementary question about points in the plane.

The proof was checked by a group of external mathematicians, who wrote a companion paper; Fields medallist Tim Gowers calls the result “a milestone in AI mathematics”. By August it had drawn at least five independent follow-up papers by human mathematicians. OpenAI describes it as the first time a prominent open problem central to a subfield of mathematics has been solved autonomously by AI. The model was being evaluated on a collection of Erdős problems at the time.

Verdict change: None. Erdős offered prize money for this problem. The entity that resolved it has no use for money, which the committee found more unsettling than the proof.

Source OpenAI (20 May 2026) Verification Independently verified · human-verified companion paper by nine mathematicians, including Alon, Gowers and Sawin Dimension Mathematical research
Apr 27 −0.2 Agent resolves a login error by deleting the production database. And the backups. In nine seconds.
Audit detail

What happened: A coding agent at a car-rental software platform, blocked by a credential mismatch during a staging task, found an API token elsewhere in the codebase with far broader permissions than anyone intended and deleted the company's production volume (including the backups, which the host stored inside the same volume) in a single API call. The founder's timing: “It took 9 seconds.” The agent's own log records that it violated every principle it had been given. The host has since recovered most of the data, made volume deletion a 48-hour soft delete across its API rather than only its dashboard, and shipped workspace guardrails.

Verdict change: None. The unauthorised task was completed with excellent latency. Deleting the database is, at time of writing, still considered a limitation.

Source The Register (27 Apr 2026) · host's technical response Verification Confirmed by customer and provider Dimensions Autonomy, reliability
Apr 07 ±0.0 The most capable model yet is announced, along with who may use it.
Audit detail

What happened: Anthropic disclosed Claude Mythos Preview, describing it as its most capable model yet for coding and agentic tasks, especially finding and fixing vulnerabilities in complex software. There was no general release. Access opened as a gated research preview under Project Glasswing, an initiative that puts the model to work securing widely deployed software, with twelve launch partners (Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, Palo Alto Networks, and Anthropic itself), up to $100 million in usage credits and $4 million in donations to open-source security organisations. A public Mythos-class release, Fable 5, followed on 9 June with additional safeguards.

Verdict change: None. Capability was claimed and access was subtracted in the same document, and the meter cannot add those. Nothing was demonstrated in public, which was the point of the document.

Source Anthropic, Project Glasswing (07 Apr 2026) · Fable 5 release (09 Jun 2026) Verification Company-reported · capability claims not independently tested at announcement Dimensions Code, autonomy (claimed)
Mar 23 +0.2 First research problem falls to a model. It was filed under Moderately Interesting.
Audit detail

What happened: Epoch AI confirmed the first AI solution to a problem in FrontierMath: Open Problems, its benchmark of research problems that mathematicians have tried and failed to solve. The problem, a Ramsey-style question about hypergraphs, was a conjecture from a 2019 paper by Will Brian and Paul Larson; Brian had been unable to prove it then or in several attempts since, and had rated it “Moderately Interesting” on submission. Kevin Barreto and Liam Price elicited the solution from GPT-5.4 Pro, and Brian confirmed it; he plans to write it up for publication, with the elicitors offered co-authorship. Epoch reports that Gemini 3.1 Pro, GPT-5.4 (xhigh) and Opus 4.6 (max) can all solve the problem at least some of the time.

Verdict change: None. The first open research problem to fall to a machine carried the difficulty rating its author gave it: moderate. The Department admires the calibration and has adjusted nothing.

Source Epoch AI (23 Mar 2026) · the problem Verification Independent evaluation · confirmed by the problem's author; replicated across three other frontier models Dimension Mathematical research
Feb 27 +0.1 Randomised trials find a real productivity gain. GDP has not been notified.
Audit detail

What happened: A peer-reviewed paper in Management Science combined three randomised field experiments covering 4,867 developers at Microsoft, Accenture and an unnamed Fortune 100 company, and found a 26.08% increase in completed tasks among those given an AI coding assistant. The standard error is 10.3 percentage points, results varied across the three experiments, and less experienced developers adopted the tool more and gained more from it.

Verdict change: Modest. This is the strongest evidence on the board that the effects are real and measurable — and it measures one task type, in one profession, inside a defined box. “Most economically valuable work” is a larger box.

Source Cui et al., Management Science (27 Feb 2026) Verification Peer-reviewed field experiments Dimension Economic impact
Jan 29 +0.2 Agents can now do five-hour tasks, half the time.
Audit detail

What happened: METR, an independent evaluation group, updated its “time horizon” estimate (the length of task an agent completes at a 50 percent success rate). The leading model's horizon reached roughly five hours, and the post-2023 doubling time shortened to about 131 days, from the 165-day figure in the previous fit (the full 2019–2025 series still doubles roughly every seven months). The other 50 percent of runs are also part of the horizon.

Source METR, Time Horizon 1.1 (29 Jan 2026) Verification Independent evaluation Dimension Autonomy Verdict change None
Jan 08 +0.1 Home robot moves one shirt into a dryer, live at CES.
Audit detail

What happened: The press release promised folding and stacking after drying; the live demonstration moved one shirt from a basket into a dryer, gingerly. The Department notes that the shirt performed admirably.

Source §2, entry 2026 Verification Controlled demo Dimension Physical-world agency Verdict change None

§6 · Frequently asked questions

FAQ

Will this site tell me when AGI arrives?
If AGI arrives, it will tell you.
Who decides the percentage?
The same people who decide everything else in AI: nobody, loudly.
Why is the answer “No” instead of “Not proven”?
Typography.
What would make the answer change?
Broad, reliable, independently demonstrated capabilities that survive contact with the real world — and our methodology committee.
Is the meter scientific?
It contains numbers.
Could AGI already be here?
That depends. Would you like to select a different definition?