30 commits · 8 organisations

The AI race: who shipped what, and when

A comparative timeline of the AI race: the first version built in an evening with Claude, finished over the next days, and checked by a second model. Counted from the logs.

Near-black cover with faint quarter lines. On the left, the Timeline mark, four coloured bars against one axis, and the title "Who shipped what, and when", above the line "The AI race on one axis, built with Claude and counted from the logs." On the right, a screenshot of the live site: the heading "The AI race", the switch to AGI Watch, and the race grid: eight organisations as rows, each quarter a column of event cards from 2024 into 2025, such as GPT-4o, Claude 3.5 Sonnet and GPT-5.

Timeline on the live site: eight organisations as rows, one quarter per column. Image and site: Edgaras Neverdauskas.

Download
Alt text · 490 characters

Near-black cover with faint quarter lines. On the left, the Timeline mark, four coloured bars against one axis, and the title "Who shipped what, and when", above the line "The AI race on one axis, built with Claude and counted from the logs." On the right, a screenshot of the live site: the heading "The AI race", the switch to AGI Watch, and the race grid: eight organisations as rows, each quarter a column of event cards from 2024 into 2025, such as GPT-4o, Claude 3.5 Sonnet and GPT-5.

Why I built it

On the afternoon of 1 September I was cleaning up my portfolio with Claude, deciding which repositories to make public. One of them was an old timeline from 2021: a page of retail and consumer brand histories, one brand at a time. I kept it, and noted what I wanted instead: a new version on its own subdomain, "new project not the refactor", with the old one only as inspiration, and "amybe isntead amazon or google we will go with AI companies".

At 19:35 that evening I came back to it. Claude read the old repository first. It had one commit from May 2021 and five files: jQuery, 12 brands, 163 events from 1917 to 2017. Every logo was broken, because the folder the data pointed at had never been committed.

So the new one started from a question, not from the old code. I asked Claude how it imagined the product. Its answer was the one the product still runs on. A chronology of one company is a Wikipedia article. The question nobody can answer in one view is the comparative one: when GPT-4 shipped, what did Anthropic have in market, and how long did the others take to respond?

As with CSS 3D Lab and J.A.R.V.I.S., I went back through the session logs and the full git history, so the numbers are counted, not remembered. My messages are quoted as I typed them, typos included.

What Timeline is

A free website at timeline.edgarasneverdauskas.com with two views of the AI race.

The first is the race itself. Eight organisations run as rows against one shared axis of quarters: OpenAI, Anthropic, Google DeepMind, Meta AI, Mistral AI, xAI, Nvidia and Hugging Face. Read down a column and you see who was shipping in the same three months. It holds 53 events, from CUDA in June 2007 to GPT-6 Astra on 3 September 2026. Every event is dated to the day and links to a first-party source: the organisation's own page, or for two research papers the authors' own posting. Each has a weight from one to three, so the default view shows the turning points and a control reveals the rest. Runs of empty quarters fold into one hatched column that says how many years it swallowed. You pan and zoom it like a video editor. On a phone it becomes one organisation at a time, running down the page.

The second is AGI Watch: twenty moments from 1950 to now that changed the answer to "are we getting there?", and the claims people made along the way.

It is built with Next.js, React, TypeScript and Tailwind, with zod checking the data, Vitest and Playwright testing it, and GitHub Pages serving it. It installs as an app. The code is open source under AGPL-3.0.

How it happened

Tuesday 19:35: the plan. In fifteen minutes of messages we settled the name, freed it by renaming the old repository, and agreed the scope. Claude proposed three views over one dataset: the race, a view of who shipped each capability first, and one merged feed. It asked whether the race alone was the product. I said it was: "yea race with time line alrasy the product we dont nee dto overcoplicated". At 19:50 I told it to go all in and not stop until it was ready for testing, "i am now off for some time".

Tuesday 19:53: the data first. Claude warned that its own knowledge ran to May 2026, so it checked the recent years against primary sources. It wrote a script that fetches every source link, and the script caught two dead ones before anything shipped. The first commit landed at 20:33: 57 events, 48 sources, all eight organisations, 58 minutes after my first message. Two minutes later it was live on GitHub's own address, with every check in the deploy green.

Tuesday 20:35: on its side. I opened it and wrote: "i checke ui/ux it doesnt seems friendly for user". Claude had made each organisation a column with time running down the page, so events could carry sentences and a phone would only scroll one way. I asked whether the organisations should be rows and the years run across. Claude agreed and gave the reason: eight columns meant five or six on screen, and one on a phone, and you cannot compare what you cannot see together. Seven minutes after my message the grid was on its side. The same commit fixed a quieter bug: no event had been given the lowest weight, so two settings of the detail control showed the same 57 events.

Tuesday 20:42 to 21:06: making it usable. A mouse wheel did nothing to a sideways grid, so the only way across nineteen years was a thin scrollbar. I asked for scroll and pinch zoom "like it would be music/video editor like". Claude argued against a library and wrote its own: the wheel pans, Ctrl and the wheel zoom around the pointer, you can drag from anywhere, and the keyboard works too. Then I asked what the hatched columns were, because "i didnt uderstand". They now carry a label, and a short key sits above the grid. I also suggested the phone should show one company at a time, and it does. At 21:05 I asked for a hint line under the grid to go: it repeated the key word for word.

Tuesday 22:59 to 23:36: the look. For half an hour I sent small remarks and Claude committed each one. The pinned names needed an edge, because cards slid under them and seemed to start inside a company's name. The edge needed its own colour, because it matched the year lines. Each row got a faint wash of its organisation's colour, which in dark mode turned to what the commit calls "mud" until it was made theme-aware. Then a shadow on one side only, a gutter at each end, names centred in their rows, and cards kept off the column lines. That was nine commits in 32 minutes. At 23:36 I had added the DNS record, and the site was live on its own address.

Wednesday: the icon, and the empty boxes. After midnight Claude drew the mark: four coloured bars against one axis. I asked for the browser icon without its dark plate, which only an installed app icon needs, and for the mark to be exactly square. In the morning I installed it on my phone and the icon was cut off. Android crops these icons and only promises the centre 80% survives, so the mark now sits smaller inside the plate. Then I zoomed all the way out and every card was empty. Below a certain width titles were hidden on purpose, leaving the view the product exists for full of blank boxes. At that moment Claude's session limit ran out, for more than four hours. At 15:16 the titles scaled down with the zoom instead of disappearing.

Friday 4 September, morning: whose word counts. OpenAI had released GPT-6 Astra the day before, and I asked whether it could go in. Claude added it, and when I asked if the details were correct, it found its own mistake in the model's input types. Then I made a rule: only entries the companies officially published, each linking to their own page. Claude audited all 58 events and found 28 sourced to Wikipedia or to a news index whose contents change. It moved 23 of them to the organisations' own pages and asked about the other five. They were things with no first-party page we could link, such as an acquisition and a funding round. Claude suggested keeping four of them on Wikipedia; I said remove them. 53 events were left, and the change was merged at 11:55.

Friday 12:49: AGI Watch. Earlier, after hearing that with Astra OpenAI said we were now in the AGI era, I had asked a Claude chat what tool we could build to watch AGI arrive, from the beginning of AI. It made a one-page prototype. I brought it into Claude Code, which put it in its own public repository ten minutes later. At Claude's suggestion it went from 29 entries to 20, keeping only the moments that changed the answer and leaving releases to the race view.

Then I asked whether it should have been part of Timeline all along. Claude argued, about 60/40, that it should stay separate: nine of its twenty entries belong to no company, and a page about long time horizons gained something from being one file with no build. I wrote: "i feel that AGI watch is like part of timline so i woudl nto need anoterh subdomain unless you have stgrong argument fro this". Its reply was "No, I don't have a stronger argument". Sixteen minutes later AGI Watch was a second view inside Timeline, with its own data file and rules, waiting for my review. It was merged at 14:21. Claude suggested archiving the standalone version three times; I chose to delete it.

Checking the checker

Timeline's data is the product, so the most useful work on the last day was not building. It was checking.

When I asked Claude to make sure AGI Watch said only what the record supports, it found four wrong claims in the twenty entries. Two had come over from the Claude chat's prototype. The other two Claude Code had added itself while reworking the page: the bar-exam figure, and GPT-3's date. The worst was about GPT-4 and the bar exam. The entry quoted a re-analysis of GPT-4's famous 90th-percentile score and got its figure wrong. As the commit put it, on a view whose subject is claims being read carelessly, that is the worst kind of error to ship. DeepSeek R1's release was mixed up with the market selloff a week later. ChatGPT's famous "100 million users in two months" turned out to be a bank's estimate, not OpenAI's, and no link to it got past the repository's link checker, so the number came out. GPT-3 was dated to its API launch while its entry was about the paper. After I questioned it, Claude had also checked the headline entry and found it overstated what OpenAI's president said: his sign-off was stronger than his answer when he was asked directly.

Then Claude hit its limit, and I opened Codex. I asked it to check the data, so I would not be relying on one tool. Codex read the page Claude had already corrected and found three more problems. A phrase in the Astra entry could not be found in the sources it checked. The bar-exam figure was more precise than the paper allows, because the paper gives two different numbers. And the chart counted 2026 as a full year in September. Codex fixed all three and pushed at 15:25.

It was not flawless either. It first reported an entry with no source as a mistake. Then it read the repository's rules, which allow exactly that one exception, a multi-year AI winter with no single event to cite, and withdrew the point.

The lesson I took: a model checking its own claims finds real problems, and a second model finds different ones. Neither makes a claim true. What made the checking work was that every claim had a source link a person could open, and a script that proves each link still resolves.

What we got wrong

  • The first layout put each organisation in a column, so you could never see them all at once. It took me one look.
  • No event had the lowest weight, so two settings of the detail control showed the same 57 events.
  • A mouse wheel did nothing to the grid. The only way across was the scrollbar.
  • Dragging refused to start on cards, and the grid is mostly cards. Then the drag handler fought the phone's own scrolling.
  • The tests that set the scroll position raced the grid's own opening scroll, so they passed locally and failed on GitHub.
  • The hatched gap columns were a visual language nobody had explained, me included.
  • On a 360px phone the toolbar was 390px wide.
  • Row colours turned muddy in dark mode, and the browser icon wore a dark plate it didn't need.
  • The installed app icon was cut off on Android.
  • Zoomed right out, every card was an empty box.
  • 28 of 58 events cited Wikipedia or a changing index page. Five had no first-party source at all.
  • AGI Watch went live with four wrong claims (two from the prototype, two added by Claude Code), an overstated headline quote, and a label that overlapped the paragraph above it at every screen width.
  • Claude's first attempt to delete the old AGI Watch folder emptied it and then failed. I stopped it and asked it to make sure it had the right folder. It had; Timeline was untouched.

Almost all of these were found the same way: I used the site and said what looked wrong, or asked for a check. Claude's own tests and measurements caught the rest.

Who did what

Claude Code did the building, and every reply in the counted sessions came from Claude Opus 5. It replied 657 times and made 636 tool calls: 322 commands, 125 file edits and 107 actions in a browser, plus 33 web searches. That was about 5 hours 45 minutes of active work, with no helper agents. In git, 29 of the 30 commits came from it: 28 it wrote, plus the merge of AGI Watch, made by a command it ran.

Codex had under an hour on the Friday afternoon, on GPT-5.6. I sent it 28 messages, and it replied 40 times with 43 tool calls in about 50 minutes. Most of that went on writing; it also rechecked the AGI Watch data and made one commit with three fixes.

Timeline by the numbers, 1 – 4 Sep 2026: 30 commits, +5,869 / −622 lines excluding the npm lockfile and licence text; 58 min from first message to first commit; 7 min from "not friendly" to the grid turned on its side; on AGI Watch's first day, Claude's fact-check fixed 4 claims and Codex's fixed 3 more. Claude Code, Opus 5: 657 replies, 636 tool calls, about 5 h 45 min active. Codex, GPT-5.6: 28 messages, 43 tool calls, 1 commit. Me: 96 messages, typically about 25 words, 19 screenshots.

Four days, counted from git and the session logs: Claude Code, Codex and me, and how it went.

Download
Alt text · 493 characters

Timeline by the numbers, 1 – 4 Sep 2026: 30 commits, +5,869 / −622 lines excluding the npm lockfile and licence text; 58 min from first message to first commit; 7 min from "not friendly" to the grid turned on its side; on AGI Watch's first day, Claude's fact-check fixed 4 claims and Codex's fixed 3 more. Claude Code, Opus 5: 657 replies, 636 tool calls, about 5 h 45 min active. Codex, GPT-5.6: 28 messages, 43 tool calls, 1 commit. Me: 96 messages, typically about 25 words, 19 screenshots.

A Claude chat made the first AGI Watch prototype. That conversation is not in the logs on my PC, so I can't count it. What it left behind was a page and a README, which Claude Code rebuilt from what I pasted.

My part was deciding. I sent 96 messages, 68 to Claude and 28 to Codex, and a typical one was about 25 words. 25 of them went in while Claude was still working, and I sent 19 screenshots. I decided what the product was, and I kept cutting what it wasn't. The biggest calls were mine: turn the grid on its side, make it move like an editor, one company at a time on a phone, only first-party sources, AGI Watch inside Timeline. Claude argued against the last one, and said so when it had no better argument.

The last three commits, on 12 September, came from sessions in another project that these counts don't include: a coffee link, a README header and the AGPL licence. One of them was made by Claude Fable 5.1.

Treat the numbers as a record, not a benchmark. The two tools worked on different days, on different jobs.

What it cost

The site costs nothing to run. GitHub Pages hosts it, and the address is a subdomain of a domain I already had.

The cost was the subscriptions, and the limits shaped the week. Claude's session limit hit three times during the Timeline work (and once more during portfolio work, not counted): at 23:38 on the first night, while it was adding Timeline to my portfolio; at 10:53 the next morning, in the middle of the zoom fix, which then waited more than four hours; and at 14:43 on Friday, which is when I opened Codex.

What I would tell someone starting

  • Start from the question, not the old code. The 2021 version was useful as a memory of what I liked. The new one began with what it should answer.
  • Look at it within the hour. The layout that defeated the product shipped with every test green. Seven minutes after I looked, it was fixed.
  • When you don't understand your own screen, nobody will. My confusion about the hatched columns became the key above the grid.
  • Make the data prove itself. A schema that requires a source on every event, and a script that fetches every source, caught dead links on day one.
  • Decide whose word counts. "Only what the company published" turned out to be a rule the data could follow, and it removed five events with no first-party page to link.
  • Ask a second model, then check both. Claude found four problems in AGI Watch, two of them its own. Codex found three more, and one false alarm of its own.
  • Push back, and let it push back. Claude argued against merging AGI Watch. When I disagreed, it said it had no stronger argument, and we merged.

In one paragraph

In one evening, one person and Claude went from an old jQuery page to a comparative timeline of the AI race, live on its own address, in 16 commits. Three days later it had a second view, AGI Watch, and every event pointed at a first-party source. Claude wrote almost every line, and Codex checked the facts behind it. The shape came from short human calls: turn it on its side, let it zoom like an editor, I don't understand this, only official sources, keep it in one place. The models made building fast. Deciding what was true, and what was worth showing, stayed with me.

Try it at timeline.edgarasneverdauskas.com. The code is open source under AGPL-3.0 at github.com/Evirtual/timeline. Built with Claude Code and Claude Opus 5, with an afternoon of OpenAI Codex on GPT-5.6.