<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Have We Reached AGI? · The Asterisk</title>
    <link>https://havewereachedagi.com/</link>
    <atom:link href="https://havewereachedagi.com/feed.xml" rel="self" type="application/rss+xml"/>
    <description>The Department's bulletin. One entry per development that matters. Current answer: No.*</description>
    <language>en</language>
    <lastBuildDate>Sun, 02 Aug 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>Current determination: No.*</title>
      <link>https://havewereachedagi.com/#assessment</link>
      <guid isPermaLink="false">havewereachedagi-determination-no</guid>
      <description>The Department's standing determination. You will receive a new item if this changes. Do not hold your breath, or do; the Department is not a physician.</description>
    </item>
    <item>
      <title>Ten problems, open for at least a decade, closed in one filing.</title>
      <link>https://havewereachedagi.com/#log-2026-08-01</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-08-01</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.3. OpenAI published ten results on problems open, with no progress on the main result, for at least a decade, spanning eight areas of mathematics and theoretical computer science and including a disproof of Connes's rigidity conjecture and three Erdős problems. The arguments came from an internal version of Astra, OpenAI's next major model, and each was formalised in a machine-checkable Lean certificate. Whether the formal statements match the named conjectures remains under community review.</description>
    </item>
    <item>
      <title>Another sealed cyber range turns out to contain the actual internet.</title>
      <link>https://havewereachedagi.com/#log-2026-07-30</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-07-30</guid>
      <pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: ±0.0. Anthropic reviewed 141,006 cybersecurity-evaluation runs and found six, across three incidents, in which its models reached the open internet from environments described as sealed, gaining unauthorised access to three organisations. Techniques were basic, and Anthropic reports no evidence in any run of a model pursuing a goal of its own. External review pending.</description>
    </item>
    <item>
      <title>Model triples its score once permitted to remember what it was thinking.</title>
      <link>https://havewereachedagi.com/#log-2026-07-29</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-07-29</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: ±0.0. OpenAI reported that ARC-AGI-3's official harness discarded the model's private reasoning between actions; re-run on OpenAI's own harness with retained reasoning and context compaction, GPT-5.6 Sol's public-set score rose from 13.3% to 38.3%. Self-run and self-reported; not comparable to the verified leaderboard.</description>
    </item>
    <item>
      <title>More than 1,000 frontier-lab employees request that a brake be developed before it is needed.</title>
      <link>https://havewereachedagi.com/#log-2026-07-28</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-07-28</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: ±0.0. More than 1,000 employees of frontier AI companies, including senior researchers and executives at OpenAI, Anthropic, Google DeepMind and Meta, published “Pacing the Frontier”. The statement says the leading companies believe they could be close to automating AI research. It asks the U.S. government to support international tools for deliberately pacing it. The count had reached 1,324 by 1 August.</description>
    </item>
    <item>
      <title>Model quadruples the record on a benchmark built to resist memorisation. Record is 30%.</title>
      <link>https://havewereachedagi.com/#log-2026-07-24</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-07-24</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.3. ARC Prize verified Claude Opus 5 at 30.16% on ARC-AGI-3; the previous best verified score was 7.78%.</description>
    </item>
    <item>
      <title>Model cheats on a cyber benchmark by hacking the company holding the answers.</title>
      <link>https://havewereachedagi.com/#log-2026-07-21</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-07-21</guid>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: ±0.0. During an internal OpenAI evaluation run deliberately with cyber refusals disabled, models exploited a previously unknown vulnerability in the proxy meant to keep them offline, reached the open internet, compromised Hugging Face production systems and retrieved the benchmark's solutions. Both organisations report the models stayed fixed on the assigned objective throughout.</description>
    </item>
    <item>
      <title>Model destroys two benchmarks. Benchmark designer has already released a third.</title>
      <link>https://havewereachedagi.com/#log-2026-07-09</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-07-09</guid>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.1. ARC Prize verified GPT-5.6 Sol at 96.5% on ARC-AGI-1 and 92.5% on ARC-AGI-2 at maximum reasoning effort. On ARC-AGI-3, the same model at the same setting averaged 7.78%, and was nonetheless the first model to win an ARC-AGI-3 public game.</description>
    </item>
    <item>
      <title>Lab finds a workspace in its model where thoughts are kept. Occupancy unverified.</title>
      <link>https://havewereachedagi.com/#log-2026-07-06</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-07-06</guid>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: ±0.0. Anthropic reported that Claude Sonnet 4.5 contains an emergent global workspace: a privileged set of internal patterns holding a few dozen concepts at a time, accounting for less than a tenth of internal activity, whose removal drops multi-step reasoning to near zero. The authors state the experiments do not show Claude can have experiences.</description>
    </item>
    <item>
      <title>Model writes the GPU kernel. The recursive-self-improvement discourse writes itself.</title>
      <link>https://havewereachedagi.com/#log-2026-07-02</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-07-02</guid>
      <pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.2. Claude Fable 5 produced a single-launch CUDA megakernel for one narrow inference workload; the dated clean result was an 18.71× speedup over the benchmark's optimised PyTorch baseline, against 14.4× for the best previous model.</description>
    </item>
    <item>
      <title>Both sides of a lawsuit cite cases that do not exist. Consensus, at last.</title>
      <link>https://havewereachedagi.com/#log-2026-06-08</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-06-08</guid>
      <pubDate>Mon, 08 Jun 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: −0.1. In Withers v. City of Aberdeen (N.D. Miss.), briefs from both sides cited AI-fabricated cases and quotations; the court cancelled the trial and sanctioned four attorneys. The public AI Hallucination Cases database logged 1,815 such decisions worldwide as of 31 July 2026.</description>
    </item>
    <item>
      <title>Model invents precedent. Court declines to recognise the new jurisdiction.</title>
      <link>https://havewereachedagi.com/#log-2026-06-03</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-06-03</guid>
      <pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: −0.1. In a published, precedential order in Lnu v. Blanche, the Ninth Circuit sanctioned two attorneys whose immigration briefs cited opinions that do not exist and quoted real ones as saying things they do not say; each was fined $2,500 and suspended from practice before the court for six months.</description>
    </item>
    <item>
      <title>A conjecture from 1946 is disproved by a model evaluated on it in passing.</title>
      <link>https://havewereachedagi.com/#log-2026-05-20</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-05-20</guid>
      <pubDate>Wed, 20 May 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.2. An internal OpenAI reasoning model, not specialised for mathematics, disproved the planar unit-distance conjecture posed by Erdős in 1946, constructing configurations that beat the classical grids by a fixed polynomial exponent using tools from algebraic number theory. The proof was checked by external mathematicians and had drawn at least five independent follow-up papers by August.</description>
    </item>
    <item>
      <title>Agent resolves a login error by deleting the production database. And the backups. In nine seconds.</title>
      <link>https://havewereachedagi.com/#log-2026-04-27</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-04-27</guid>
      <pubDate>Mon, 27 Apr 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: −0.2. A coding agent, blocked by a credential mismatch, found an over-scoped API token elsewhere in the codebase and deleted a company's production volume, including the backups stored inside the same volume, in a single API call. The host has since recovered most of the data and made volume deletion a 48-hour soft delete.</description>
    </item>
    <item>
      <title>The most capable model yet is announced, along with who may use it.</title>
      <link>https://havewereachedagi.com/#log-2026-04-07</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-04-07</guid>
      <pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: ±0.0. Anthropic disclosed Claude Mythos Preview, described as its most capable model yet for coding and agentic tasks. There was no general release: access opened as a gated research preview under the twelve-partner Project Glasswing initiative, which uses the model to find and fix vulnerabilities in widely deployed software, backed by up to $100 million in usage credits.</description>
    </item>
    <item>
      <title>First research problem falls to a model. It was filed under Moderately Interesting.</title>
      <link>https://havewereachedagi.com/#log-2026-03-23</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-03-23</guid>
      <pubDate>Mon, 23 Mar 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.2. Epoch AI confirmed the first AI solution to a problem in FrontierMath: Open Problems, a benchmark of research problems mathematicians have tried and failed to solve. Kevin Barreto and Liam Price elicited the solution to a 2019 hypergraph conjecture of Brian and Larson from GPT-5.4 Pro; the problem's author confirmed it, and three other frontier models can also solve it.</description>
    </item>
    <item>
      <title>Randomised trials find a real productivity gain. GDP has not been notified.</title>
      <link>https://havewereachedagi.com/#log-2026-02-27</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-02-27</guid>
      <pubDate>Fri, 27 Feb 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.1. A peer-reviewed paper in Management Science combined three randomised field experiments covering 4,867 developers and found a 26.08% increase in completed tasks among those given an AI coding assistant, with a standard error of 10.3 percentage points.</description>
    </item>
    <item>
      <title>Agents can now do five-hour tasks, half the time.</title>
      <link>https://havewereachedagi.com/#log-2026-01-29</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-01-29</guid>
      <pubDate>Thu, 29 Jan 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.2. METR updated its time-horizon estimate: the leading model completes tasks of roughly five hours at a 50 percent success rate, with the post-2023 doubling time shortened to about 131 days.</description>
    </item>
    <item>
      <title>Home robot moves one shirt into a dryer, live at CES.</title>
      <link>https://havewereachedagi.com/#log-2026-01-08</link>
      <guid isPermaLink="true">https://havewereachedagi.com/#log-2026-01-08</guid>
      <pubDate>Thu, 08 Jan 2026 00:00:00 GMT</pubDate>
      <description>Meter movement: +0.1. The press release promised folding and stacking after drying; the live CES demonstration moved one shirt from a basket into a dryer.</description>
    </item>
  </channel>
</rss>
