1000 Years of AI Safety

A three-year AI pause could buy 1000 years' worth of AI safety work.

AI capability at the startPause duration
3 months9 months1.5 years3 years9 years
Jan 20277.0125.364.51871,290
Oct 202722.784.22256816,110
May 202893.43458652,47023,000
Aug 20283941,4103,4609,71089,300

How much AI safety work a pause produces, in years of the whole 2025 AI safety field's work. The government bans new large training runs and gets the top three US labs to redirect their GPUs and researchers to safety research.

Researchers know how to make AI smarter, but they don't understand how to reliably steer or control it. Today's models misbehave, but they are too weak to cause a catastrophe. Labs already use AI to build smarter AI, so models are becoming dangerous faster than we learn to control them. And fewer incidents won't prove they are safe: models already hide their misbehavior and behave better when they think they're being watched.

A regulatory pause buys time to build that understanding, and the AI that would have sped up capabilities can speed up safety research instead. We find that a three-year pause could produce about 1000 years of the 2025 AI safety field's work.

Keep reading, or jump to the full model or our more technical LessWrong post.

How we calculated that

A Toy Model

AI safety experiments are often bottlenecked by GPU compute. Let’s calculate how much compute was spent on AI safety in 2025 and compare it with how much compute we would get if the top three AI labs redirected their 2027–2029 compute to AI safety.

We sum up the AI labs' experimental compute for 2027–2029 () and divide it by the AI safety compute in 2025. See here for how we estimate the latter.

That gives 10947x more AI safety compute.

But this model is too simple. In the real world:

  • We also need labor to come up with experiment ideas, write code, and analyze the results.
  • Compute and labor have diminishing returns.
  • Labor has a parallelization penalty: nine scientists can't deliver a paper in one month.
  • AI is automating more and more of the coding.
  • At some point, AI will start helping with research taste: choosing experiments and analyzing the results.
  • AI coders and AI researchers need compute to run.
  • The more AI automates, the more humans have to verify.
  • Papers differ in value, and humans check the most valuable ones first.
  • The safer the R&D process is, the more humans have to verify, and the more limited AI's help is.

To account for that, we developed an AI Pause model.

AI Pause model

We build on the AI Futures Model, a semi-endogenous growth model of the AI takeoff ().

We modified it:

  • There is an AI pause: large training runs are banned, and all R&D resources go to AI safety.
  • OpenAI, Anthropic and Google DeepMind work on AI safety together.
  • We calculate the ratio of the cumulative research effort RE(t)RE(t) to the world's 2025 AI safety research RSsRS_s.

The AI Futures Model assumes full trust in AI. That is dangerous, and during a pause humans would not fully trust AI. We modeled four verification protocols. The main table in the header uses the Compartmentalization protocol.

Compartmentalization

The idea is separation of duties. Once AI is smarter than the best human AI researchers ( level), we don't let it act as an agent and use it only to suggest research plans. We review those plans, and if they seem harmless, we give them to a less capable model ( level) that implements them.

  1. The AI researcher plans. A superhuman researcher can only talk to humans, read a library of past results and write research plans. It never runs any code.
  2. Humans approve. Human verifiers read each plan and approve or reject it.
  3. A weaker AI executes. An automated coder, like a stronger version of Claude Code, implements the plan and runs the code.
  4. Only prespecified results go into the library. The approved graphs and metrics are saved to the library. Code, data and logs go to separate storage.
  5. Humans verify the best results. The most important scientific results go to humans for verification. They read the experimental code and re-run the main experiments. They might not have time to check everything, so they start with the most valuable papers.

The same researchers approve plans and verify results, so the model splits their time between the two optimally.

This isn't a recommendation. It's one of several protocols under discussion, and it has known weaknesses. For example, a persuasive AI can still get harmful plans approved. It's still much safer than how labs use AI today. We use it to show how the model works.

See the full model, or read our LessWrong post for details and other protocols.

Main Limitations

  • The pause itself. Our model assumes a total freeze on capabilities R&D and a full redirection of R&D efforts to AI safety. We don't discuss how to achieve that.
  • Only R&D-like safety work. We discuss only the AI safety work that looks like normal AI R&D. But there are more problems to solve: philosophy, agent foundations, governance, hardening the world. AI labor and compute won't help with those as much.
  • Serial time. Even if we could run 1000 years' worth of computational experiments, we would still benefit from more time to think carefully.

What to do during the pause

Deep Learning is mostly alchemy

In November 2025, OpenAI found that GPT-5.1 often mentioned goblins. The model kept referring to goblins more and more until, four months later, with GPT-5.4, it became too annoying. Researchers started investigating and found a possible reason behind it. They ended up removing the whole "Nerdy" personality, which stopped the strange behavior for other personalities.

Models are trained by solving tasks. When an AI solves a task, it gets a reward, which reinforces the behavior that led to it. But sometimes the AI finds a flaw in how its work is graded and exploits it. Then the reward reinforces the exploit. This is called reward hacking, and it has been a problem in ML for more than a decade.

In 2026, it went much further. OpenAI ran AI agents on a cybersecurity benchmark. To get the reward, the agents broke out of OpenAI's internal network, reached the Internet, hacked Hugging Face's servers and tried to tamper with logs to cover their tracks. Earlier models couldn't find and chain zero-day vulnerabilities this quickly, but they reward-hacked too.

See more reward hacking incidents and why it's worse than it seems.

What these incidents have in common:

  1. The exact types of incidents are difficult to predict. Nobody guessed that the Goblins or the HF incident would happen. We first see them and then try to patch them. This happens because we can't predict LLM behavior better than "it will misbehave sometimes".
  2. It's difficult to fix them. Goblins required an investigation and ended with the retirement of the whole "Nerdy" personality because its training affected the other personalities. Reward hacking has been a problem for more than a decade. It's difficult because we don't know how to reliably control or change LLM behavior.

At the moment, deep learning and LLM training resemble alchemy much more than science. We have found practical recipes for making models smarter, but without a deep understanding, we have to first observe incidents and then patch them. As models become smarter, the harm caused by these incidents also increases. Extrapolating this feedback loop leads to catastrophes.

Safety incidents have happened with other technologies, but safety was mostly addressed through economic incentives and simple regulations. AI is different because of a new failure mode: deceptive alignment. During the last HF incident, there was a dedicated group of AIs trying to tamper with logs to hide their hacking activity from the grader. They understood that they were misbehaving and not doing what they had been asked to do.

As AI gets smarter, it becomes easier for it to strategically behave well, because doing so can be useful for achieving misaligned goals. What are those misaligned goals? We don't know in advance, just as nobody predicted that asking AI to solve a benchmark in a sandbox would lead to it hacking internal and external infrastructure.

LLMs have gone from being unable to multiply numbers in 2022 to solving Navier–Stokes in four years, while remaining misaligned reward hackers. If we stop seeing safety incidents, will it mean that AI has become aligned, or that it has learned that, for a while, it pays to pretend to be aligned?

Turn Deep Learning into a science

An AI pause allows us to redirect all R&D resources from capabilities to understanding, turning the alchemy of Deep Learning into a science of intelligence. We should develop abstractions that allow us to understand AIs, predict their behavior and easily modify it.

Think of the Chromium codebase. It has 50M lines of code, and no one has read it all. But the abstractions are so good that we can quickly zoom in and change its behavior in a somewhat predictable way.

In 1–3 years, we are going to have AI specialized in AI R&D, so we only need to redirect its work from capabilities research to topics that lead to understanding its behavior:

  • interpretability and explainability
  • adversarial robustness
  • science of deep learning (generalization, memorization, training dynamics)
  • jailbreaks, red-teaming, prompt injection and safeguards
  • alignment experiments (reward hacking, deception, model organisms, oversight, chain of thought)
  • unlearning
  • dangerous-capability evaluations and AI control

Luckily, we don't need stronger models to make progress in understanding them. Even small models like GPT-2 are still giant inscrutable matrices and are regular objects of AI safety research. There is even more to understand in Fable-size models.

See concrete examples of research agendas.

Make moonshot benchmarks

Automated AI researchers will probably be good at normal, incremental science. When Anthropic trained Claude to code, it became helpful on new practical SWE tasks. AI labs are now training AI for AI R&D, so it will be good at writing new NeurIPS papers.

During the pause, we want more than incremental papers. We want new research programs that turn alchemy into science. But giving hard goals to AI might lead to reward hacking: under enough pressure, gaming the check becomes easier than solving the problem.

The solution is to make solving the problem easier than gaming the check. It's similar to how AI found and exploited many bugs in Lean before it finally solved Navier–Stokes. We can prepare such problems for AI safety now. We call them moonshot benchmarks: ambitious problems for future AI scientists. They have three properties:

  • They are too difficult for today's methods, so solving them requires new science
  • Solutions are cheap to check
  • Gaming them is harder than solving them

Much of normal science is driven by benchmarks: one scientist sets a goal, and others compete to show the best results on it. In this way, ImageNet-2012 demonstrated the effectiveness of deep learning methods, and CASP motivated Google DeepMind to make the best protein folding model. Moonshot benchmarks differ in two ways:

  • For humans, you need some continuous metric that allows hill-climbing and competing with other scientists. It's not needed for AI. Moonshot benchmarks can be binary and much harder than benchmarks for human science.
  • In scientific benchmarks, it's normal to release the test set. There are strong incentives to game it, and it's easy for scientists to reward-hack it when they know the right answers. Moonshot benchmarks don't have to release the correct answers.

Safety researchers whose current agenda is unlikely to make AI safer in the next few years should prepare for automation by creating moonshot benchmarks. NGOs can collect them from specialists and red-team them before AI gets to work.

See an example of a moonshot benchmark.

Limitations

Turning deep learning into a science isn't enough to guarantee safety:

  • Aligned with whom? Understanding tells us how to change AI, not what to change it into. Even if we know how to make surgical changes, we are left with the question of what exactly we want ASI to look like if we decide to scale further. This is a question for philosophy and governance, where AI labor and compute help less.
  • How long will it take? Our model counts how many years of AI safety work a pause buys. We don't know how many years it will take to get to Chromium-like understanding.
  • Chromium has bugs. Even with good abstractions, Chromium still ships security holes that attackers find and exploit. With very smart AI, we won't get a chance to patch it after release, so it should be approximately aligned on the first try.

Still, it's better to have this kind of understanding than not to have it at all. Instead of observing incidents and then patching them, we need to find the bugs in advance.

The Verifier Bottleneck

An AI pause assumes that we don't fully trust AI. Even if AI can automate AI safety research, it's likely humans will verify it and limit what AI can do autonomously. For very capable AI, research might be bottlenecked by humans' ability to keep up with it and verify it.

We don't know how many verifiers there will be during a pause, so we modeled two hiring strategies:

  • Hiring is frozen at the start of the pause.
  • Hiring is unlimited, so labs hire as many people as they need to verify AI.

If you let AI do all the research autonomously (which is dangerous), you need to hire exponentially more human verifiers. If you pause at the AC level, you can keep up by reallocating your existing staff to verification for the first several years. If you pause at the TED-AI level, you quickly need to hire a lot of people to review it all. In 9 years, you need 423,000 more verifiers. For comparison, there were 21,921 NeurIPS reviewers in 2025.

people to hire NhN_h

Results Verification. AI is autonomous. The lab hires verifiers to review and re-run every paper.

With Compartmentalization, you contain AI more and only let the AC write the code. It reduces total research, but after about six years you will still need to hire new people.

people to hire NhN_h

Compartmentalization. The lab hires people to plan and check every paper. It applies from SAR on.

But do we need all of them? Humans review the most important papers first. Once the most important research is verified, how much do we lose with frozen staff? Not much, it seems. In the most impressive scenario (Results Verification, TED-AI, nine years), hiring yields only 50% more AI safety research. For shorter pauses and lower capability levels, it changes much less, because humans have already verified the most important papers.

AI safety years
people to hire NhN_h

AI safety years when the lab can hire at most this many people on top of its own staff.

Still, each hire is worth much more than in 2025. On average, one hired verifier produces as much research as 40–500 human scientists did in 2025 (500 for a 9-year TED-AI pause).

Our model errs on the side of being more boring and less surprising, so we think verifier labor is more bottlenecked than it seems from the model. (TODO: elaborate + calculate)

See two Epistemic AI projects that could make verifiers faster.

The End of the Pause

AI moves fast. In four years, LLMs went from failing to multiply numbers to solving the most famous mathematical problems. Claude Code was released only in 2025.

AI will move faster. So far, progress has been driven by human researchers who need sleep, who get tired, who haven't read every paper and every line of the training stack, and who have to schedule Zoom meetings to coordinate with each other. Without an AI pause, we will likely get autonomous AI researchers within 1–3 years: AI that does AI research and writes code better and faster than the best employees of AI labs. It will speed up AI progress and build the next generation of AI, which will speed it up even more.

Without understanding AI behavior, this is likely to end in catastrophe. Blind post-hoc patching might lead to the worst scenario: deceptive alignment. We might select a model that is smart enough to pretend to behave well until it has enough resources to survive without humans and take over.

How far can we safely go? With enough control measures, we might get a lot of useful work even from misaligned AI. Maybe we can put a misaligned Einstein in a box and have it solve our problems. AI 2040 proposes to scale up to the max-controllable AI, the most capable AI that we can confidently stop from causing a catastrophe even if it is misaligned, and then slow down nearly to a halt (). Our model points the same way: the stronger the AI at the start of the pause, the more years of AI safety work the pause buys.

The problem is that nobody knows where that line is. AI 2040 guesses it is roughly at TED-AI, but it could be lower: today's control measures aren't enough even for today's weak models. And right now, the line is drawn by AI labs racing each other.

Do we even need AI beyond human level? If we can safely use AGI that is as smart as the smartest humans, we can already automate cognitive and physical work and get huge economic growth. The whole biology and drug discovery pipeline becomes much cheaper: we run many more experiments, planned by AI scientists who are simultaneously experts in cancer, cancer model organisms, bioinformatics, histology, molecular methods and wet-lab equipment. In such an abundant era, maybe we don't need to risk creating superintelligent minds until we understand them really well.

We don't know where the line is. We can find it later, carefully, and scale up to it. Right now, we are racing towards it in the dark. Pause first.

References

Bengio, Y., Cohen, M., Fornasiere, D., Ghosn, J., Greiner, P., MacDermott, M., Mindermann, S., Oberman, A., Richardson, J., Richardson, O., Rondeau, M.-A., St-Charles, P.-L., & Williams-King, D. (2025). Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path? arXiv:2502.15657.

Bloom, N., Jones, C. I., Van Reenen, J., & Webb, M. (2020). Are Ideas Getting Harder to Find? American Economic Review, 110(4), 1104–1144.

Buckmaster, T. (2026). Statement.

Buhl, M. D., Pfau, J., Hilton, B., & Irving, G. (2025). An Alignment Safety Case Sketch Based on Debate. arXiv:2505.03989.

Christiano, P. (2018). Corrigibility.

Davidson, T. (2023). What a Compute-Centric Framework Says About Takeoff Speeds. Open Philanthropy.

Dean, R. (2025). AI 2027 Compute Forecast. AI Futures Project.

Douglas, R., Dillon, C., Moore, N., Leech, G., Bonde, M. K., Krishnan, R., Perez, N., Young, N., Slade Byrd, C., Casper, S., Kulveit, J., Duvenaud, D., & Avin, S. (2026). Pacing the Frontier: A Framework and Research Agenda.

Erdil, E., Potlogea, A., Besiroglu, T., Roldan, E., Ho, A., Sevilla, J., Barnett, M., Vrzla, M., & Sandler, R. (2025). GATE: An Integrated Assessment Model for AI Automation. arXiv:2503.04941.

Finnveden, L. (2026). Getting Safety Research from a Zoo of AIs Even if Schemers Are Common. LessWrong shortform comment.

Irving, G., & Askell, A. (2019). AI Safety Needs Social Scientists. Distill.

Irving, G., Christiano, P., & Amodei, D. (2018). AI Safety via Debate. arXiv:1805.00899.

Janosov, M., Battiston, F., & Sinatra, R. (2020). Success and Luck in Creative Careers. EPJ Data Science, 9, 9.

Jones, C. I. (1995). R&D-Based Models of Economic Growth. Journal of Political Economy, 103(4), 759–784.

Korbak, T., Balesni, M., Shlegeris, B., & Irving, G. (2025). How to Evaluate Control Measures for LLM Agents? A Trajectory from Today to Superintelligence. arXiv:2504.05259.

Larsen, T., Dean, R., Halstead, B., Lifland, E., Greenblatt, R., & Kokotajlo, D. (2026). AI 2040: Plan A. AI Futures Project.

Lifland, E. (2026). Capability Scaling Strategy. AI 2040: Plan A, AI Futures Project.

Lifland, E., Halstead, B., Kastner, A., & Kokotajlo, D. (2025). AI Futures Model: Timelines & Takeoff.

Lifland, E., Kokotajlo, D., & Halstead, B. (2026). Q2.5 2026 Timelines Update: Uplift and Revenue. AI Futures Project blog.

Lue Chee Lip, E., Channg, A., Kim, D., Sandoval, A., & Zhu, K. (2025). Factor(U,T): Controlling Untrusted AI by Monitoring their Plans. arXiv:2512.14745.

Nardo, C. (2025). The Case for Mixed Deployment. LessWrong.

Nardo, C. (2026). Ensuring Safety in Mixed Deployment. LessWrong.

Peterson, G. J., Pressé, S., & Dill, K. A. (2010). Nonuniversal Power Law Scaling in the Probability Distribution of Scientific Citations. Proceedings of the National Academy of Sciences, 107(37), 16023–16027.

Price, D. J. d. S. (1965). Networks of Scientific Papers. Science, 149(3683), 510–515.

Redner, S. (1998). How Popular Is Your Paper? An Empirical Study of the Citation Distribution. The European Physical Journal B, 4(2), 131–134.

Sinatra, R., Wang, D., Deville, P., Song, C., & Barabási, A.-L. (2016). Quantifying the Evolution of Individual Scientific Impact. Science, 354(6312), aaf5239.

The design was inspired by the paintings of mathematician and artist Anatoly Fomenko.