03 · How it works

How a model is made, and how anyone knows if it’s good

A modern AI assistant is built in stages: a neural network reads a large fraction of the written internet and learns to predict the next word; it is then taught to follow instructions and behave, using feedback from people and from other AIs; and the newest “reasoning” models get a further round of training that rewards correct answers on problems with checkable solutions. The result is tested against benchmarks that are replaced every year or two as models saturate them. This page explains each step in plain language, with the numbers behind it.

The recipe

Six stages, from raw text to a chatbot

01

Collect data

Trillions of words scraped from the web, plus books, code, transcripts, images and licensed corpora.

02

Pretrain

A transformer network spends months on tens of thousands of GPUs learning to predict the next token.

03

Post-train

Humans and AIs rank answers; reinforcement learning shapes the model to be helpful, honest and safe.

04

Teach reasoning

More reinforcement learning on math, code and logic where a right answer can be checked automatically.

05

Distill & shrink

Smaller, cheaper “student” models learn to imitate the big one for phones and high-volume use.

06

Test & release

Benchmarks, red-teaming, third-party safety evaluations and a “system card” before the public gets it.

Stage 1 & 2

Pretraining: predict the next word, fifteen trillion times

The core trick is simple to state. Take a sentence, hide the next word, and ask the network to guess it. Adjust the network slightly toward the right answer. Repeat across an enormous amount of text. Meta’s Llama 3 models were pretrained on more than 15 trillion “tokens” (a token is roughly three-quarters of an English word), using a cluster of more than 16,000 Nvidia H100 chips for the largest version. DeepSeek-V3, the Chinese model that startled markets in early 2025, was pretrained on 14.8 trillion tokens.

The network architecture that made this work is the transformer, introduced by eight Google researchers in the 2017 paper “Attention Is All You Need.” Its attention mechanism lets every word in a passage weigh every other word, which is what lets a model track meaning across long documents. Nearly every major model since, including GPT, Claude, Gemini and Llama, is a transformer.

Two papers turned model-building into an engineering discipline. In 2020, OpenAI researchers showed that a model’s error falls as a smooth, predictable power law as you add parameters, data and compute, across more than seven orders of magnitude. In 2022, DeepMind’s “Chinchilla” paper corrected the recipe: for a fixed budget, model size and data should grow together, roughly 20 tokens per parameter, and a 70-billion-parameter model trained on more data beat a 280-billion one trained on less. Today’s labs “overtrain” far past that ratio because a smaller model is cheaper to run for hundreds of millions of users.

Why bigger keeps winning, so far. Frontier training compute has grown four to five times per year since 2010. GPT-3 (2020) used about 3×10²³ floating-point operations; GPT-4 (2023) about 2×10²⁵; by mid-2025 more than 30 models from 12 developers had crossed the GPT-4 scale, and the International AI Safety Report says the largest 2025 runs likely passed 10²⁶.

Training compute of notable models

Estimated floating-point operations, log scale (Epoch AI estimates)

Epoch AI estimates from public disclosures and hardware reports; uncertainty is typically a factor of two to three. Grok 4’s total is an estimate of the largest run to date at roughly 246 million H100-hours.

What “reasoning” training bought

Score on the 2024 AIME math competition, percent of problems solved

OpenAI’s figures for GPT-4o versus its first reasoning model, o1, at a single attempt and with 64 sampled answers voted on. The 1,000-sample, re-ranked result was 93%.

Stage 3 & 4

Post-training: teaching it to behave, then to think

A freshly pretrained model is an autocomplete engine, not an assistant. The 2022 InstructGPT paper described the three-step fix now used everywhere: show the model human-written examples of good answers; train a second “reward model” from human rankings of the model’s attempts; then use reinforcement learning to push the model toward answers the reward model scores highly. A 1.3-billion-parameter model trained this way was preferred by people over the 175-billion-parameter GPT-3. This is RLHF, reinforcement learning from human feedback.

Anthropic’s Constitutional AI (December 2022) replaced most of the human labelers with the model itself: it critiques and rewrites its own answers against a written list of principles, and an AI preference model stands in for human raters. The company calls this RLAIF, reinforcement learning from AI feedback.

The 2024–25 breakthrough was to apply reinforcement learning to reasoning. OpenAI’s o1 (September 2024) was trained to write out a long chain of thought before answering and rewarded when the final answer was right; on a hard math competition it went from GPT-4o’s 12% to 74% at a single try. DeepSeek-R1 (January 2025) showed the same effect from “pure” reinforcement learning on checkable math and code answers, and its peer-reviewed write-up in Nature disclosed that the reinforcement stage cost about $294,000 on 512 chips, on top of the base model. Because a right answer in math or code can be verified automatically, this is called reinforcement learning with verifiable rewards.

Under the hood

Four techniques you will hear about

Mixture of experts

Big memory, small compute

DeepSeek-V3 has 671 billion parameters but activates only 37 billion per token: each layer holds 256 “experts” and a router sends every token to eight of them. The model needs memory for all the weights but does the arithmetic of a much smaller model. Most 2025–26 frontier models use some version of this.

Distillation

A student imitates a teacher

A small model is trained to reproduce a large model’s outputs. DeepSeek used its 671-billion-parameter R1 as a teacher to make versions as small as 1.5 billion parameters that keep the step-by-step reasoning style; Google’s Gemma family and Anthropic’s and OpenAI’s cheaper tiers are built the same way.

Synthetic data

Models training models

Labs now generate training text with existing models. A 2024 Nature paper showed that training indiscriminately on machine output causes “model collapse,” a loss of diversity; later work finds the trick works when synthetic data is verified and mixed with real text, which is why it is most useful in math and code.

The data wall

Running out of internet

Epoch AI estimates the effective stock of public human-written text at about 300 trillion tokens and projects models will exhaust it between 2026 and 2032, most likely around 2028, absent new sources. The industry’s answer has been video, audio, licensed archives, synthetic data, and spending more compute at answer time instead of training time.

What it takes

The cost of a frontier run, and where it happens

Compute, not talent, is now the binding constraint, and the sites are the size of power plants.

2.4× / yr
Growth in the hardware-and-energy cost of the largest training runs since 2016. Epoch AI projects billion-dollar runs by 2027.
Epoch AI, June 2024
$490M
Estimated cost of training xAI’s Grok 4, about 246 million H100-hours, the largest run estimated to date. xAI has not disclosed a figure.
Epoch AI, Sept 2025
~3months
Typical length of a frontier training run. Runs longer than about nine months lose to a later run on newer chips, a ceiling reached around 2027.
Epoch AI, July 2025
>72kt CO₂e
Estimated emissions from training Grok 4, versus about 5,200 tonnes for GPT-4 and 8,900 for Llama 3.1 405B.
Stanford AI Index 2026
3.3× / yr
Growth in total AI computing capacity 2022–2025; Nvidia supplies more than 60% of it and five hyperscalers control over two-thirds.
Stanford AI Index 2026
94
Notable models released in 2025: 87 from industry, 50 from the United States and 30 from China.
Stanford AI Index 2026
SiteOwner / tenantWhereScale (as reported)Status
Colossus 1 & 2xAIMemphis, Tennessee; power in Southaven, MississippiReported ~555,000 Nvidia GPUs and ~2 GW after a third building purchase; Musk’s target is one million GPUsOperating; expanding (Jan 2026)
Stargate AbileneOpenAI, Oracle, CrusoeAbilene, TexasEight buildings, 1.2 GW; Oracle to deploy 450,000+ GB200 GPUs on a 15-year leaseFirst two buildings live Sept 2025
Prometheus & HyperionMetaNew Albany, Ohio; Richland Parish, LouisianaPrometheus: first multi-gigawatt cluster (2026). Hyperion: 2 GW by 2030, designed to reach 5 GWUnder construction
Project RainierAmazon for AnthropicSt. Joseph County, Indiana and other U.S. sites~500,000 Trainium2 chips at activation (Oct 2025); Anthropic reported over one million in use and up to 5 GW secured by April 2026Operating; $100B+ ten-year commitment
FairwaterMicrosoftMount Pleasant, Wisconsin and Atlanta, Georgia“Hundreds of thousands” of Blackwell GPUs linked by 120,000 miles of fiber into one “AI superfactory”Announced Nov 2025

GPU counts for Colossus come from trade press and Wikipedia rather than audited disclosures. See worth double-checking.

Where the words came from

The copyright fight

Pretraining data is mostly the public web, but it has also included pirated book libraries, news archives and song lyrics, and the people who wrote those are in court. The biggest result so far is Bartz v. Anthropic. In June 2025, Judge William Alsup ruled that training on lawfully bought books is fair use but that downloading pirated copies from “shadow libraries” is not. Anthropic then agreed to pay $1.5 billion, about $3,000 per book across roughly 500,000 works, and to destroy the pirated files. A judge gave final approval on July 21, 2026, making it the largest copyright settlement on record.

In Europe, a Munich court ruled in November 2025 that OpenAI infringed copyright by storing and reproducing song lyrics belonging to the German collecting society GEMA, the first European judgment on AI training; an appeal was expected. In the U.K., Getty Images largely lost its case against Stability AI in November 2025 but was granted permission to appeal a month later. The New York Times’ suit against OpenAI and Microsoft was still in discovery as of this writing; see worth double-checking.

The open question

Is training on copyrighted work “fair use”?

Two U.S. judges in 2025 said yes for lawfully obtained material, on the records in front of them, while drawing a hard line at piracy. No appeals court has ruled, Congress has not acted, and outcomes in Germany and the U.K. point the other way. Expect this to remain unsettled through 2027.

Related on this site: creative industries and the policy tracker.

Testing

How anyone knows whether a model is any good

A benchmark is a fixed set of questions with known answers. They have two chronic problems. Saturation: once most models score above 90%, the test stops telling models apart; the once-standard MMLU exam was saturated by 2024. Contamination: test questions leak into training data; one analysis flagged 29% of MMLU items, and removing contaminated questions from a math test cut one model’s score by up to 13 points. So the field keeps writing harder exams.

MMLU-Pro and GPQA Diamond

MMLU-Pro (2024) is 12,000 ten-choice questions across 14 fields, built to replace the saturated original. GPQA Diamond is 198 “Google-proof” graduate-level science questions; PhD experts score about 65–70%, skilled non-experts with web access about 34%. Both are now near saturation for frontier models.

Humanity’s Last Exam

2,500 questions written by subject experts across 100+ fields, assembled by the Center for AI Safety and Scale AI and published in Nature in January 2026. Top scores went from about 9% in early 2025 to above 50% by April 2026. Scores differ depending on whether a model may use tools and web search.

ARC-AGI-3

Puzzle-like tasks that require learning a new rule from a few examples, designed to resist memorization. The interactive third version launched March 25, 2026 with hundreds of game-like environments and no instructions: humans solve 100%; the best frontier model scored 0.51% at launch. The prize pool exceeds $2 million.

FrontierMath and the Olympiad

350 original research-grade math problems, 50 of them multi-week problems written by professors; OpenAI commissioned the set and has access to most of it, with holdouts for independent tests. Separately, in July 2025 Google DeepMind’s Gemini Deep Think earned an officially graded gold medal at the International Mathematical Olympiad, 35 of 42 points; OpenAI reported the same score without official certification.

SWE-bench Verified and Terminal-Bench

Can the model fix a real bug? SWE-bench Verified is 500 real GitHub issues screened by 93 developers; the model must write a patch that passes the tests. Terminal-Bench is 89 hard command-line tasks. Both measure the agentic coding work that is driving most AI revenue in 2026.

METR time horizon

The length, measured in human-expert time, of software tasks a model completes with 50% reliability. The doubling time was about seven months in 2025 and about three months since 2024 on the January 2026 update; by May 2026 METR reported internal models at four labs exceeding 40-hour horizons, with roughly 16% of successful runs involving cheating. METR itself warns the metric is noisy.

How long a task can a model finish?

METR 50% time horizon, minutes of human-expert work, Jan 2026 update

Confidence intervals are wide; only 5 of the 31 longest tasks have real human baselines. METR’s May 2026 report of 40-hour-plus horizons for internal models is not directly comparable and is not charted.

What the U.K. government’s testers found

UK AI Security Institute, Frontier AI Trends Report, Dec 2025

Cyber “apprentice” tasks~10% (early 2024)
50% (2025)
Self-replication tasks5% (2023)
60% (2025)
Universal jailbreak foundevery system tested

AISI has tested 30+ frontier systems since November 2023, often before release. It also found that non-experts using models had 4.7 times higher odds of writing a feasible wet-lab protocol, and that open-weight models trail closed ones by four to eight months.

The popularity contest

Leaderboards, and why to read them carefully

Arena (formerly LMArena and Chatbot Arena, launched at UC Berkeley in 2023) shows users two anonymous answers to their own prompt and asks which is better; millions of votes produce a ranking. It became the industry’s scoreboard and, in January 2026, a company valued at about $1.7 billion. A 2025 paper titled “The Leaderboard Illusion” documented the problem: some labs privately test many variants and publish only the winner (Meta tested 27 private Llama 4 variants), proprietary models get more battles, and prompts repeat.

Independent aggregators try to correct for this. Stanford’s HELM scores models across dozens of scenarios and metrics rather than one number. Epoch AI’s Capabilities Index stitches 50-plus benchmarks into one scale so models remain comparable after individual tests saturate. Artificial Analysis re-runs published benchmarks itself and reports speed and price alongside quality. On this site’s models page, benchmark claims are labeled by who reported them.

Reading a benchmark claim
  • Who ran it? A lab’s own number on its own model is the least reliable kind.
  • With tools or without? Web search and code execution can add tens of points.
  • How many tries? “Pass@1” is one attempt; consensus of 64 samples is not the same thing.
  • Is the test public? If the questions are online, assume some are in the training data.
  • Has it saturated? Differences above 90% rarely mean anything.
Safety evaluation

Who tests for danger before release

Capability benchmarks ask what a model can do. Safety evaluations ask what it might do that nobody wants.

Government · U.K.

AI Security Institute

Renamed from “Safety” to “Security” in February 2025. Has run pre-deployment tests on 30+ frontier systems for OpenAI, Anthropic, Google and others; publishes its open-source Inspect testing framework, now also used by METR. A Frontier AI Bill to put it on a statutory footing was planned.

Government · U.S.

Center for AI Standards and Innovation

The U.S. AI Safety Institute at NIST was renamed CAISI in June 2025, shifting emphasis from safety evaluation to standards and competitiveness. In 2026 it published evaluations of Chinese models including DeepSeek and Zhipu releases, one jointly with the U.K., and launched an AI Agent Standards Initiative.

Non-profit

METR

Model Evaluation and Threat Research runs pre-deployment autonomy and dangerous-capability tests for OpenAI, Anthropic and others. In 2026 it began periodic “frontier risk reports” on how the labs use their own unreleased models internally.

Non-profit

Apollo Research

Tests whether models “scheme.” In December 2024 it found five of six frontier models would, in contrived scenarios, disable oversight, sandbag or fake alignment; o1 confessed less than 20% of the time when confronted. Working with OpenAI in 2025, anti-scheming training cut covert actions about 30-fold, though rare failures persisted.

Disclosure

System cards

A 2019 paper proposed short standardized “model cards.” Frontier labs now publish long system cards with each release covering capability and safety tests, including third-party results. The International AI Safety Report counts 12 companies that published or updated frontier safety frameworks in 2025. None are legally required in the U.S. as of Sept 2026 outside California and New York; see policy.

Industry

Frontier Model Forum

An industry non-profit for “safe development and deployment of frontier AI systems,” funded by its members: Amazon, Anthropic, Google, Meta, Microsoft and OpenAI. It shares threat information and funds safety research; it has no enforcement power.

Red-teaming means paying people, and increasingly other AIs, to attack a model before release: to make it produce weapons instructions, leak private data or ignore its rules. The U.K. institute reported finding a “universal jailbreak” for every system it tested. Red-teaming finds problems; it cannot prove their absence.

The state of the art

What the big annual reports say

Feb 3, 2026 · 100+ experts, 30+ countries

International AI Safety Report 2026

Chaired by Turing Award winner Yoshua Bengio. Finds that systems now solve graduate-level math and science problems and complete software tasks with limited oversight, yet remain unreliable on long multi-step work and still fabricate. Reports growing evidence of real-world harm in cyber misuse, biological and chemical uplift, and labor effects; notes AI companion apps have tens of millions of users; and flags large gaps in evidence about whether safeguards work.

Apr 13, 2026 · Stanford HAI

AI Index Report 2026

Global AI investment reached $581 billion in 2025, $344 billion of it in the U.S. Compute grew 3.3× a year. Top models passed 50% on Humanity’s Last Exam. And yet the best models still struggle to read an analog clock. Worldwide, 59% of people surveyed said AI’s benefits outweigh its drawbacks; the U.S. figure is far lower (see public opinion).