Self-Improving AI, Explained: What Labs Have Actually Shipped

Rachid Idali

by Rachid Idali

In April 2026, about 60,500 Americans a month typed some version of "self improving" into Google. In May it was 450,000. That is a 7.4x jump in four weeks, and it happened before any of the reports people now cite had been published. The term has since settled at 135,000 a month, held there through July and August as of Aug 2026, more than sixteen times where it sat a year ago.

Self-improving AI means a system that improves its own successor: writing the code, running the experiments, deciding what to try next, without a person in the loop. That is the definition every lab uses, and by it nobody has built one. What the labs have shipped is narrower and more interesting, and the gap between the two is where the confusion lives. Nothing ranking for this term shows you the curve, and nothing shows you the strangest thing in it: the word growing fastest alongside it has nothing to do with self-improvement at all.

Key takeaways:

  1. "Self improving" went from 8,100 to 135,000 US monthly searches in a year, up 1,945%, with a 450,000 peak in May 2026. Classified EXPONENTIAL in the Rising Trends database (data as of Aug 2026).
  2. The spike came before the two documents everyone now cites. Anthropic published its recursive self-improvement report on June 4, 2026 and OpenAI published its chief scientist's essay on September 6. May had neither.
  3. Anthropic says more than 80% of the code merged into its codebase was authored by Claude as of May 2026, up from low single digits before Claude Code launched in February 2025.
  4. OpenAI put the claim in a product page first. Its February 2026 launch post calls GPT-5.3-Codex "our first model that was instrumental in creating itself."
  5. "agentic ai" is down 45% year over year at 60,500 searches while the recursive cluster climbs: recursive self improvement +650%, recursive intelligence +1,344%, recursive ai +296%. This is vocabulary replacement, not vocabulary growth.
  6. The fastest riser in the family is a false friend. "recursive language models" is up 159,900% and refers to an MIT long-context inference trick, not to AI building AI.

Let's get into it.

The curve

Here is five years of monthly search volume for the term, straight from our database. For four of those five years it is a flat line at roughly 6,600 to 9,900 searches a month. Then it is not.

Line chart of monthly Google search volume for self improving over 60 months, from 8,100 in Sep 2021 to 135,000 in Aug 2026, peak 450,000 in May 2026

Source: Rising Trends database, data as of Aug 2026

The shape matters more than the peak. March 2026 was 14,800. April was 60,500. May was 450,000. June fell to 201,000, July to 135,000, and August held there. That is not a staircase like a company whose funding rounds keep making news. It is a detonation followed by a plateau, and the plateau is the real signal: after the spike burned off, demand settled at roughly twenty times its pre-2026 baseline and stopped falling.

The other terms in our database moved with it, and they tell you what kind of story this is.

Horizontal bar chart: Related searches around "self improving". self improving 135,000, agentic ai 60,500, recursive self improvement 5,400, recursive ai 1,900, recursive language models 1,600, recursive intelligence 1,300

Source: Rising Trends database, data as of Aug 2026

The second-largest term on that chart is the interesting one. "agentic ai" still draws 60,500 searches a month, and we covered its rise in our 2026 AI agents report. It is now the only term in the group going backwards.

Horizontal bar chart: Search growth over the last year. recursive language models +159,900%, self improving +1,945%, recursive intelligence +1,344%, recursive self improvement +650%, recursive ai +296%, agentic ai -45%

Source: Rising Trends database, data as of Aug 2026

Agentic ai is down 45% in a year while every recursive term is up. That is not a market adding a concept. It is a market swapping one word for another. The agentic ai trend page and the recursive ai page cross in our data somewhere in the next eighteen months if both hold their slope.

What self-improving AI actually means

Strip the marketing off and there is a spectrum, not a switch.

At the loose end, self-improving AI means any use of AI to build AI: a model writing training code, debugging a run, grading another model's output. By that standard it has been happening for years and every lab qualifies. At the strict end, it means a system that improves the process by which it improves, generating its own ideas, evaluating its own results and modifying its own methods with no human setting the goal. By that standard nobody qualifies, including the companies whose valuation would benefit most from qualifying.

Anthropic's framing, on the page below, is the clearest published version: an AI "capable of fully autonomously designing and developing its own successor." Note the three words doing the work. Fully. Autonomously. Successor.

Anthropic Institute page titled When AI builds itself, subtitled Our progress toward recursive self-improvement, and its implications

Source: anthropic.com, captured 2026-09-26

That page is the most-cited document in this conversation, and it is worth reading what it actually claims. It does not claim recursive self-improvement exists. Its first section says the opposite: "We are not there yet, and recursive self-improvement is not inevitable."

What it does claim is measured, dated and internal. As of May 2026, Anthropic says more than 80% of the code merged into its codebase was authored by Claude, against low single digits before Claude Code launched in February 2025. Engineers ship 8x as much code per quarter as they did from 2021 to 2025. Claude's success rate on the company's hardest tier of open-ended engineering problems hit 76% in May 2026, up 50 percentage points in six months, and an update posted on September 18 pushed it to about 91%.

The honest summary is that the doing has been automated and the choosing has not. Anthropic says so itself: large performance gaps persist "when it comes to Claude exercising judgement in choosing goals." That is the whole distinction, and every headline drops it.

Why it broke out in May

Here is the thing our data shows that the coverage does not. The search spike is earlier than the documents.

Anthropic's report went up on June 4, 2026. OpenAI's essay went up on September 6. The May explosion has neither. What May had was a hire and a survey.

On May 19, 2026, Andrej Karpathy joined Anthropic's pre-training team. His own post ran four sentences: "Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative." Reporting that week said he was helping stand up a team that uses Claude to accelerate the research that trains Claude. The same day, at Google I/O, Demis Hassabis put artificial general intelligence at 2030, "plus or minus a year." Twelve days earlier, IEEE Spectrum had run the first proper technical survey of who had closed which part of the loop.

None of that is a product launch. It is a week in which the idea got a face, a date and a taxonomy at once, and search behaviour followed. By the time the labs published their formal positions, the word was already in circulation. If you track terms for a living, that ordering is the lesson, and it is why the dashboard beats the press release by about a month.

The clearest public explainers came later, and they are mostly people doing what we are doing here: separating what shipped from what the phrase implies. This clip from @iamkylebalmer, posted September 19, 2026, frames it the way the labs do, with the humans still choosing the experiments.

Its caption is a better definition than most of the news coverage: engineers still choose experiments, decide what matters and judge the results.

Who is building on it

Four different things are being sold under one phrase. Keeping them apart is most of the work.

WhoWhat it actually isOne number
AnthropicClaude writing and reviewing Anthropic's own code and running its experimentsOver 80% of merged code, May 2026
OpenAICodex models used to debug and deploy their own training runs77.3% on Terminal-Bench 2.0
Dream-RSI (academic)An orchestration layer that reuses an agent's search history to improve its exploration policy17 authors, posted Sep 14, 2026
RecursiveA lab founded to build self-improving research systems. No product yetTeam of over 25
Ricursive IntelligenceAI that designs the chips that train the AI$335M raised

OpenAI put it in a product page before it put it in a safety essay. The February 2026 launch of GPT-5.3-Codex says, in the company's own words, that it is "our first model that was instrumental in creating itself", and that the team used early versions to debug its own training, manage its own deployment and diagnose its evaluations. In September, chief scientist Jakub Pachocki wrote in An Alien Mind that "we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward." Same company, seven months apart: marketing claim, then strategy.

The academic work is narrower and more honest. The Dream-RSI paper, posted September 14, 2026, does something specific: it treats an agent's accumulated discovery history as a replay simulator, tests exploration strategies against that cheap simulator instead of the expensive real task, then redeploys the winner. Each run makes the simulator better, which makes the next round of policy search better. That is a genuine loop. It is also a loop around exploration strategy, not around model weights, and the paper never claims otherwise.

Two startups are named after the idea. Recursive, Richard Socher's lab, describes itself as building "recursive self-improving superintelligence to automate knowledge discovery" with a team of over 25 and, publicly, nothing shipped. Ricursive Intelligence, founded by the AlphaChip team from Google DeepMind, says it has $335M from Sequoia, Lightspeed, DST and NVentures to close the loop between AI and the chips that run it. Our recursive intelligence trend page is up 1,344% in a year, which is partly these two being googled.

And one term in the cluster is not about this at all. "Recursive language models" is the fastest riser in our family at +159,900% year over year. It comes from an MIT CSAIL paper by Alex Zhang, Tim Kraska and Omar Khattab, and it describes an inference strategy: give the model a long prompt as a variable in a Python environment and let it call itself on chunks. It handles inputs two orders of magnitude past the context window. It has nothing to do with a system improving its successor. Track this space by keyword and that one will fool you, in the same way the language around world modeling got blurred last year.

What changes if it works

The honest answer is that one thing has already changed and the rest is downstream of it.

Code review became the bottleneck, and that is measurable. Anthropic's report names this as Amdahl's law arriving in an org chart: once Claude writes most of the code, human review is the slow step, so the company put an automated Claude reviewer in front of every merge. A retrospective found it would have caught roughly a third of the bugs behind past production incidents. If you build software, that is the near-term shape of this trend, and the same shift we tracked across software development this year.

The job changes before the headcount does. Anthropic's own employees describe it plainly. One told the company's researchers: "I started leaning hard into Claudifying about a year ago. That's been a crazy adventure and it's now been ~5 months since I last wrote any code myself." Another: "On days where everything works well, I can't help but think nothing I do matters, everything is automated and better and faster than I ever will be." Those are not predictions about 2030. They are descriptions of a job in 2026, and the argument behind the 32 hour work week bill now in Congress.

And the safety conversation stopped being abstract. In September, Anthropic alignment lead Evan Hubinger said publicly he puts more than a 10% chance on AI killing all humans within the decade, then clarified where the worry sits: "What I am worried about is superintelligence arising from recursive self-improvement." We unpacked that number in the AI apocalypse explainer. Pachocki lands in the same place from the other lab, arguing that no lab has solved alignment well enough "to continue responsibly scaling at maximum speed for much longer."

The skeptic's case

There is a serious version of the counter-argument, and it is not "AI is fake."

Nathan Lambert at the Allen Institute for AI calls it lossy self-improvement. His claim is that the loop is real but leaky: "the models become core to the development loop but friction breaks down all the core assumptions of RSI." More compute and more agents thrown at a problem produce more duplication and more loss, so the curve looks linear in hindsight rather than explosive.

Dean Ball of the Foundation for American Innovation puts the same point commercially. "Maybe eventually they're going to automate the genius," he told IEEE Spectrum, "but not next year. Next year they're automating the grunt who grinds through the algorithmic efficiency games."

The strongest number against the story sits inside the story. Anthropic's report cites, in its own footnotes, a randomized controlled trial by METR in which 16 experienced open-source developers completed 246 tasks with and without AI tooling. They predicted AI would cut their completion time by 24%. Afterwards they estimated it had cut it by 20%. Measured, AI made them 19% slower. Perceived and actual speedup pointed in opposite directions by nearly 40 points. Every internal "we ship 8x more code" figure has to be read against that, and to Anthropic's credit, its own footnote says so.

One measurement belongs to neither camp. METR's public time horizon tracker puts the best current model at a 17-hour 50% time horizon: the length of task, measured by how long a human expert needs, at which the model succeeds half the time. It is independent, it is published, and it is the closest thing to a scoreboard this field has.

What to watch

Three things would move this from vocabulary to fact, and each has a term in our database attached to it.

A lab publishes an experiment it did not choose. Every current claim keeps the human on goal-setting. The moment a frontier lab reports a capability gain from a research direction a model picked, evaluated and kept, the strict definition has been met. Watch recursive self improvement at 5,400 searches a month as of Aug 2026. It is the precise term specialists use, so it rises when practitioners talk, not when headlines do.

The vocabulary finishes its swap. "agentic ai" at 60,500 and falling 45% a year against a recursive cluster rising from a smaller base is a crossover in progress. If agentic keeps sliding through the winter while the recursive terms hold, the category has renamed itself, and anyone with agentic in their positioning will be optimising for a shrinking term.

Or the plateau breaks downward. Our acceleration reading for this trend is 0.9, meaning the last three months are slightly slower than the three before. Flat at 135,000 for two months is a term that found its level, not one still climbing. If autumn prints below 100,000, May was a news event rather than the start of a category, and the right comparison is the hype cycle we watched in generative AI rather than a durable shift.

The thing to hold on to is the gap. Search volume measures a word, and the word grew 1,945% in a year. The capability it names, by the definition the labs themselves publish, has not arrived, and the labs are the ones saying so most clearly. That gap is where the next two years of this story get written, and you can watch it open or close on the self improving trend page every month.


Want to catch the next term before the reports land? Read our guide on how to identify market trends, follow the live self improving trend page, or browse what is breaking out right now on the Rising Trends dashboard.

Unlock More Trends & Insights

Never miss a trend again!

Thousands of Emerging Trends

Thousands of Breakout Apps

Mega Trends

Trend Analysis Tool

Get Access Now
Join +5,000 happy users

Written By

Rachid Idali

Founder of Rising Trends, helping entrepreneurs identify and capitalize on emerging market opportunities through expert trend analysis and insights.

Self-Improving AI, Explained: What Labs Have Actually Shipped