Why I built it
On 3 September OpenAI released GPT-6 Astra, and its president ended a press briefing with "Welcome to the AGI era." When I explained the project to Claude the next afternoon, I put it in one line: "iehared the new that AGI is already here and i wanted to see what backs it".
It started in a Claude chat, not in Claude Code. I had asked it how AGI would affect finance, then told it that Altman said we were now in the AGI era. It corrected me: the line was Greg Brockman's, not Altman's. Then I asked: "i wnader what tool can we build itneractive ui/ux for watching AGI arrive since AI beginbing". It built a one-page prototype and a README. I told it I only needed a brief, because I wanted Claude Code to set up the project. That chat is not in the logs on my PC, so I can't count it. What survives is the copy I pasted into Claude Code.
This is the second story from the same site. Timeline covers the race view, built three days earlier. This one is only about AGI Watch: what it tracks, how the twenty entries were chosen and sourced, the wrong claims we found and fixed, and the design. As before, the numbers come from the session logs and git, and my messages are quoted as I typed them.
What AGI Watch is
A page at timeline.edgarasneverdauskas.com/agi-watch/, the second view of Timeline. It asks one question: are we getting there?
It holds twenty entries from 1950 to now, and each one is marked as one of four kinds. Nine are breakthroughs, like Deep Blue, AlphaGo and AlphaFold2. Four are releases, like ELIZA and ChatGPT. Two are setbacks: the two AI winters. Five are claims: somebody saying a threshold has been crossed, from GPT-2 "too dangerous" to Astra. Claims are amber, the loudest colour on the page, because they are what the page is sceptical of.
Above the entries sit three live counters (since Dartmouth, since ChatGPT, since "the AGI era"), a chart of entries per half-decade, and OpenAI's own five-level ladder with a marker that says it is outside analysts' opinion, not a measurement. Entries read newest first, grouped into nine eras. Every entry links its sources, 23 links in all, except the second AI winter: a multi-year collapse with no single event to cite. It is the one exception the rules allow, and a test enforces that.
How it happened
All of it happened on Friday 4 September, times in UTC+7.
12:49: the handoff. I sent Claude Code, open in the Timeline folder, a screenshot of the chat's handoff. It assumed it was Timeline's own original brief and offered to fix a line in the README. A minute later I wrote "no no this is a new porjecy" and pasted the whole chat: 3,405 words, most of them the page's code. The two files had never been downloaded, so Claude rebuilt the page from the paste and checked it: 29 entries, each inside one era, one entry flagged as the present. At 12:58 it asked where to publish, and I chose a public repository on GitHub's own address. Ten minutes after my first message, the page was live.
13:00: the first question. Not knowing it was already live, I asked whether this should have been part of Timeline. Claude counted before answering: 13 of the 29 entries already existed in Timeline's data, and they had already drifted apart on day one. But it argued the data didn't fit. Turing, Dartmouth and the winters belong to no company, and Timeline is a grid of companies. So it recommended keeping them apart and narrowing AGI Watch instead. I agreed to the narrowing. On merging I wrote "this is debatablwe".
13:03: the cut. Claude gave the page one test: does this change the answer to "are we getting there?" Nine entries failed and left: GPT-1, the Transformer paper, Claude 1, DeepSeek V3, o3, Claude Opus 4, GPT-5, Claude Opus 4.5 and Claude Sonnet 5. Releases are the race view's job. Two entries were recast rather than cut. GPT-2 became a claim, because its "too dangerous to publish" framing is what belongs here. GPT-4 became the bar-exam claim, with the re-analysis that questioned it. The cut also exposed a chart problem, which I come back to below.
13:09: the architecture. I asked for "some mdoern itnergration toolign", maybe something more interactive. Claude argued against a framework for twenty entries. The real gap was rigour: not one of the twenty sources was a link. They were names like "IBM". Dates were decimals typed by hand, and nothing checked anything. It offered four directions, and I picked rigour with no framework. It moved the data into its own file, wrote a validator and a link checker, and turned every source into a link. The validator caught the one entry nobody had sourced: Altman's "a couple of years away". Claude traced the remark to the India AI Impact Summit in February, where, it reported, his wording was sharper than ours. Of 22 links, 18 opened. The other four were publishers that block scripts, and Claude checked each one by hand. The checker now reports those as a separate state, because reporting a bot wall as a broken link teaches you to ignore the checker.
13:20: my doubt. I sent a screenshot of a layout problem and two doubts: "not sure about the "Welcome to the AGI era" quita", and a feeling that it should move into Timeline after all. Claude measured it at four screen widths. The ladder's marker label had covered the paragraph above it at every width since the first draft. Then it checked the quote. "Welcome to the AGI era" was how Brockman ended the briefing. Asked directly whether Astra is AGI, he hedged. The entry said he personally believes OpenAI has reached AGI, which came from the chat and was stronger than anything he said. By 13:22 the entry carried the hedge.
13:25: the merge. Claude was already asking for access to the Timeline folder when I stopped it: "wait so you think merge in is bad idea". It put itself at 60/40 against. The page was two files with no build, and it could open anywhere in ten years. I wrote: "i feel that AGI watch is like part of timline so i woudl nto need anoterh subdomain unless you have stgrong argument fro this". It answered: "No, I don't have a stronger argument". A quarter of an hour earlier I had chosen no framework. Sixteen minutes after my message, AGI Watch was a Next.js page inside Timeline, waiting for my review.
13:45: the check. I asked for four things: a post, a narrower centred layout, deleting the old repository once merged, and to "make sure that that data adn text is correct with whats has being recorded online and cofirmed". The check found four wrong claims. I cover them in the next section.
14:02 to 14:19: the layout. I asked for the newest entries on top, and for the chart, gauge and cards to share one width. Claude anchored the page to the left, so the title would not jump when switching views. I asked about centring. It rendered both and changed its mind: centred was better. I said centre only the lower part, then keep it left, then that it looked wrong on a big screen and maybe needed a full-width header. It built one. Then I suggested centring the race view's heading too, which fixed the jump by itself, and asked to remove the header. Claude argued to keep it. I wrote "nio i would lie kto remvoe header", and it did. Both headings now sit on one measure, and the race grid stays full width underneath.
14:20: merge and delete. Claude had recommended archiving the old repository three times, because deletion can't be undone. I wrote "dont archive delete it". It merged at 14:21. The deploy took eight minutes, most of it installing packages, and I asked if it was stuck. It was live at 14:30. Claude couldn't delete the repository, because its GitHub access lacked that permission, so that step happened outside these logs. It is gone now.
14:45: a second opinion. Claude's session limit hit at 14:43, and I opened Codex. Most of that session went on my post about AGI Watch, but I also asked it to check the data and "verify ot only relyingo n oen tool".
A page about claims has to survive its own
AGI Watch exists to hold a claim next to the record. On its first day it got its own claims wrong, and most of the work was catching that.
The first check came from me asking. Four entries were wrong. Two came from the chat's prototype. DeepSeek R1 "wipes hundreds of billions off AI-linked stocks in a day", but Claude found R1 came out on 20 January and the selloff came a week later, on the 27th. ChatGPT "reaches 100 million users in two months", but Claude found that was a UBS estimate from traffic data, not an OpenAI figure. Claude first cited a Reuters link it had written from memory. It could not confirm the link, so it dropped the number. The entry now says it quotes no adoption figure.
Two came from Claude Code itself. It had written that the bar-exam re-analysis put GPT-4 "closer to the 68th" percentile. As Claude then read the paper, it says below the 69th on July data, about the 62nd against first-time takers, and lower against those who passed. Codex later agreed. And it had dated GPT-3 to the API launch while the entry was about the paper. In the commit, it called the bar-exam error "the worst kind of error to ship" on a view about claims read carelessly.
The second check came from Codex, and it found three more. Its first pass read the live page, not the repository, and raised two points it later withdrew. The second AI winter had no source, but the repository's rules allow exactly that. And the page showed 23 links, not the 69 the checker counts across both views. Then it read the data file and found real problems. Claude's fix of the Brockman entry contained a phrase, "reasonable" if you want to, that Codex could not find in the sources it checked. The bar exam's "about the 48th" was more precise than the paper allows: Codex read about 48 in its abstract and about 45 in its detailed results. It now says "roughly the mid-40s". And the chart's open half-decade counted 2026 as a whole year in September.
That chart was the quieter lesson. Every value on it was correct. But we were inside 2025–2029, so its last bar was short and read as a slowdown. Per year, it was the fastest stretch on the page. The unfinished bar now carries a dashed outline at its pro-rata height, and the "fastest stretch" label is computed, not written. Codex's fix made that rate use the real elapsed time: about 1.7 years of five.
What made this work was not either model. It was that every claim had a link a person could open, and a checker that says whether each one still resolves.
What we got wrong
- The headline entry overstated what Brockman said. It came from the chat, and I was the one who doubted it.
- Two more claims from the prototype were wrong: DeepSeek's release and selloff, and an analyst's ChatGPT estimate stated as fact.
- Claude Code added two errors of its own while reworking the page: the bar-exam figure and GPT-3's date.
- Claude's correction of the Brockman quote added a phrase Codex could not find in the sources.
- None of the first twenty sources was a link.
- The chart implied a slowdown, then counted 2026 as a full year.
- The ladder marker sat on the paragraph above it at every screen width, from the first draft until I pointed at it.
- Claude first read my screenshot as Timeline's own brief.
- The layout went left, centred, left again, a full-width header, then centred without it, all between 14:04 and 14:19.
- Deleting the old local folder emptied it and then failed. I stopped Claude and told it to be careful. It checked, and Timeline was untouched.
- Codex raised two points it later withdrew, both from reading the live page instead of the repository.
Who did what
Claude Code built AGI Watch twice: first as a standalone page, then as Timeline's second view. Every reply came from Claude Opus 5. It replied 310 times and made 293 tool calls: 113 commands, 92 file edits and 52 actions in a browser, plus 8 web searches and 3 page fetches. That was about 1 hour 55 minutes of active work, with no helper agents. In git it made four commits in the standalone repository, which I know only from the logs because the repository is gone. In Timeline it made three commits, +1,678 / −110 lines, and ran the command that merged them.
Codex had about 50 minutes on GPT-5.6. I sent it 28 messages, and it replied 40 times with 43 tool calls. Most of it went on my post. It also made one commit with three fixes, +17 / −12.

One afternoon, counted from git and the session logs: Claude Code, Codex and me.
Alt text · 459 characters
AGI Watch by the numbers, Friday 4 Sep 2026: 29 entries cut to 20; 92 min from my first message to the view merged into Timeline; source links went from 0 to 23; Claude's fact-check fixed 4 claims and Codex's fixed 3 more. Claude Code, Opus 5: 310 replies, 293 tool calls, about 1 h 55 min active. Codex, GPT-5.6: 28 messages, 43 tool calls, 1 commit. Me: 55 messages, typically 17 words, 8 screenshots. The Claude chat that made the prototype is not on disk.
A Claude chat made the first prototype: 29 entries, the eras, the ladder and the amber claims. It isn't on disk, so I can't count it.
My part was deciding and doubting. I sent 55 messages, 27 to Claude Code and 28 to Codex, and a typical one was 17 words. I attached 8 screenshots, sent 9 messages while Claude was still working, and stopped it twice. I chose the narrowing, rigour over a framework, and the merge against Claude's 60/40. I doubted the headline quote before any check did. I asked for the facts to be checked, then asked a second tool. And I chose deletion over archiving.
Treat these numbers as a record, not a benchmark. The two tools had different jobs.
What it cost
AGI Watch costs nothing to run. It is one more page of a site GitHub Pages already hosts.
The cost was the subscriptions and one limit. Claude's session limit hit at 14:43, a second after I asked it to delete the repository. It reset at 16:20. By then the second check had already gone to Codex.
What I would tell someone starting
- Say your doubt out loud. "Not sure about the quote" led to the first real fix. No test would have raised it.
- Give the page a test for what belongs on it. One question cut 29 entries to 20, and a test now fails past thirty.
- A chart can lie with every number right. Check what the shape says, not only the values.
- Make every source a link, and check them. Tell a bot wall apart from a dead link, or people stop reading the checker.
- Ask a second model, and read what it withdraws. Codex found three problems, and dropped two false ones once it read the repository.
- Decide against the recommendation when you have a reason. Claude argued to keep AGI Watch separate and to archive the old repository. I merged and deleted, and it said plainly when it had no better argument.
In one paragraph
In one afternoon, a prototype from a Claude chat became a page of its own, then a second view of Timeline. The work that mattered was less the building than the checking. One question cut 29 entries to 20. My doubt about one quote, a check I asked for, and a second model found eight things to fix on a page whose subject is claims. Claude wrote almost every line, and Codex checked them. Whether the record backs the claim that we have arrived is still the page's open question. Deciding what it may say about that stayed with me.
Try it at timeline.edgarasneverdauskas.com/agi-watch. The code is open source under AGPL-3.0 at github.com/Evirtual/timeline. Built with a Claude chat, Claude Code and Claude Opus 5, and checked with OpenAI Codex on GPT-5.6.
