Skip to content
Anthropic Trained a 34-Million-Feature Autoencoder on Claude 3 Sonnet, Google DeepMind Then Deprioritised Fundamental Research on That Method, and Anthropic's Attribution Graphs Gave Insight on About a Quarter of Prompts
Artificial Intelligence

Anthropic Trained a 34-Million-Feature Autoencoder on Claude 3 Sonnet, Google DeepMind Then Deprioritised Fundamental Research on That Method, and Anthropic's Attribution Graphs Gave Insight on About a Quarter of Prompts

Illustration by tuput

English

Anthropic's May 2024 paper used autoencoders with up to 34 million features on Claude 3 Sonnet and said it was likely orders of magnitude short of all the features in that layer. Google DeepMind's team reported in March 2025 that simple linear probes beat autoencoders on a harmful-intent test. Anthropic's 2025 attribution graphs gave the authors satisfying insight on about a quarter of the prompts they tried.

· · 8 min read

In May 2024 Anthropic’s researchers set one internal feature of Claude 3 Sonnet, the one that responds to the Golden Gate Bridge, to 10 times its maximum activation, and the model began to describe itself as the bridge. Interpretability is the research field that tries to find out what happens inside a language model between a prompt and an answer. This piece covers what its main papers measured, from 2020 to July 2026, and what each team says its results do not show.

Vision circuits came first, with a warning attached

Distill’s Circuits thread opened on 10 March 2020 with “Zoom In,” by Chris Olah and five colleagues, then at OpenAI. It made three claims about neural networks. Features, which are directions in a network’s activity, are its basic unit. Features connect through weights into circuits. Analogous features and circuits form across different models and tasks.

The authors called these claims “deliberately speculative.” Their examples came from one image classifier, InceptionV1. They also wrote that researchers were divided on whether such networks contain meaningful units at all.

Induction heads, and a case that is stronger for small models

Anthropic’s “In-context Learning and Induction Heads,” published 8 March 2022, describes two attention heads, the parts that decide which earlier words a model looks at, working together. One passes each word’s identity on to the next position. The second finds an earlier copy of the current word, looks at what came after it, and makes that word likelier again. The authors analysed 34 transformer models over the course of training and ran more than 50,000 head ablations, which means switching a head off and measuring the damage.

Induction heads formed at the same point in training as a sharp gain in in-context learning, the way a model’s predictions improve as a prompt grows longer. The authors write that the case is stronger for small models than for large ones and call their large-model evidence “only the beginnings of evidence.” For how a model predicts text in the first place, see our explainer on why a chatbot is guessing the next word.

A one-layer test in 2023

A single neuron in a language model often responds to several unrelated things. In “Towards Monosemanticity,” dated 4 October 2023, Anthropic trained a sparse autoencoder on a one-layer transformer with 512 neurons in its MLP block. A sparse autoencoder is a second small network that rewrites a layer’s activity as a much larger set of features, only a few of them active at a time. The paper studies one run with 4,096 features.

The authors write that they have no metric they trust more than human judgement. Their blinded annotator was one of the paper’s own authors, who scored 412 activation intervals across 162 features and neurons. The median neuron scored 0, meaning the annotator could not form a hypothesis, and the median feature interval scored 12. The run recovered 79% of the loss reduction the layer provides, a figure the authors say should be taken “with a significant grain of salt.” In the paper’s comments, the outside researcher Neel Nanda reported that the core results replicated on an open-source one-layer model.

Claude 3 Sonnet and the bridge

“Scaling Monosemanticity,” published 21 May 2024, trained autoencoders with about 1 million, 4 million and 34 million features on a middle layer of Claude 3 Sonnet. At the largest size about 65% of the features were dead, never active across a 10-million-token sample, against 2% at 1 million. Clamping the bridge feature to 10 times its maximum made the model start to self-identify as the bridge. Anthropic ran a “Golden Gate Claude” demo for 24 hours, announced on 23 May 2024.

The authors list limits in the same paper. They do not believe they found anywhere near all the features in that layer, think it quite likely they are orders of magnitude short, and say finding them all in every layer would need more compute than training the model. They call their training objective, which balances accurate reconstruction against sparsity, a proxy for interpretability. Features smeared across several layers, which they call cross-layer superposition, remain unsolved. On the features tied to deception, bias and dangerous content, they caution against inferring too much.

OpenAI, Google DeepMind and a test the autoencoders lost

OpenAI’s “Scaling and evaluating sparse autoencoders,” submitted 6 June 2024, trained a 16-million-latent autoencoder on GPT-4 activations from a layer five-sixths of the way into the network. Substituting it into GPT-4 gave a language-modelling loss comparable to a model trained with 10% of GPT-4’s pretraining compute. The authors say many random activations in GPT-4 are not yet adequately monosemantic and that their 64-token context may be too short to show the model’s most interesting behaviour.

Google DeepMind’s Gemma Scope, submitted 9 August 2024, released open autoencoders for every layer and sublayer of Gemma 2 2B and 9B. On 26 March 2025 a Google DeepMind team of eight authors, including Neel Nanda and several Gemma Scope authors, posted a negative result. It trained probes to detect harmful intent in prompts, and tested them on held-out jailbreaks. Dense linear probes on the model’s activity performed nearly perfectly, including on the new jailbreaks. Autoencoder-based probes did worse, and fine-tuning the autoencoders on chat data closed about half the gap. The team said it was deprioritising fundamental autoencoder research for the moment and that autoencoders are not useless. On 1 December 2025 it wrote that it had made significant tactical errors in 2024 by tracking reconstruction and sparsity instead of performance on set tasks, and that dictionary learning has made “limited progress towards reverse-engineering at best.”

Other groups tested the method on steering. AxBench, submitted on 28 January 2025 by Zhengxuan Wu and colleagues, found on Gemma-2-2B and 9B that prompting steered best, followed by fine-tuning, and that autoencoders were not competitive on steering or on detecting concepts. A paper submitted 29 May 2026 by Mikkel Godsk Jørgensen and Lars Kai Hansen argues AxBench judged the method harshly. With their supervised feature-selection pipeline, autoencoders came close to a LoRA fine-tuning reference on that benchmark.

Attribution graphs in a production model

“On the Biology of a Large Language Model,” published 27 March 2025, traced Claude 3.5 Haiku with a replacement model of 30 million features. In a two-hop question about the capital of the state containing Dallas, the graphs showed an internal step representing Texas. In rhyming couplets, they found features for the planned end word before the line was written, in about half the poems examined. Injecting the planned words “rabbit” and “green” into 25 poems made the line end on the injected word in 70% of cases.

The authors say the graphs gave satisfying insight on about a quarter of the prompts tried, and that even in the successes they capture only a small fraction of the model’s mechanisms. The method does not explain how attention patterns are computed, and had not been scaled to prompts longer than about a hundred tokens. Its claims cover the examples shown, not mechanisms in general.

What outside tests found

Gonçalo Paulo and Nora Belrose, in a paper submitted 28 January 2025, trained autoencoders that differed only in random seed. In a 131,000-latent autoencoder on Llama 3 8B, only 30% of the features were shared across seeds. They conclude that a feature set is a useful decomposition, not a full list of what the model uses.

Thomas Heap and colleagues, submitting a day later, found that autoencoders trained on randomly initialised transformers often scored about as well on automated-interpretability metrics as those trained on real models (version 2, 27 January 2026). They argue that high scores alone do not show that a model’s working features were recovered.

OpenAI’s paper of 17 November 2025 trained transformers with most weights set to zero, and found readable circuits in them. Scaling past tens of millions of non-zero parameters while staying interpretable, the authors write, remains a challenge.

The 2026 papers

All three 2026 papers below are Anthropic’s own, studying Anthropic’s own models. On 2 April 2026 it reported emotion vectors for 171 emotion concepts in Claude Sonnet 4.5. Steering with a “desperate” vector raised the blackmail rate of an early snapshot above its 22% baseline in a test scenario, and a “calm” vector lowered it. The authors say none of this shows whether language models feel anything or have subjective experiences.

On 7 May 2026 it described natural language autoencoders, which turn an activation into text. They indicated evaluation awareness, meaning the model suspecting a test, 16% of the time in a test of destructive coding actions and on 26% of SWE-bench Verified problems, against under 1% of opted-in claude.ai transcripts. Our piece on benchmarks covers the scoring problems of that test. In an auditing game, auditors using them found a planted hidden motivation 12% to 15% of the time, against under 3% without them. Anthropic calls the game a toy model and says explanations can be wrong, sometimes inventing details missing from the transcript. Each activation read takes hundreds of generated tokens.

On 6 July 2026 it reported a “J-space” in Claude, a few dozen concepts at a time accounting for under a tenth of internal activity. Replacing “Soccer” with “Rugby” changed the answer, and replacing “spider” with “ant” changed an answer of 8 legs to 6. Anthropic says its method is imperfect, finds only concepts that match a single token, and cannot say what decides which concepts enter the J-space. Its page says Neel Nanda’s outside commentary includes an independent replication of some of the findings.

Share
Copied!

Sources & further reading

  1. Olah and others, Zoom In: An Introduction to Circuits, Distill (10 March 2020)
  2. Olsson and others, In-context Learning and Induction Heads, Transformer Circuits Thread (8 March 2022)
  3. Bricken and others, Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, Transformer Circuits Thread (4 October 2023)
  4. Templeton and others, Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet, Transformer Circuits Thread (21 May 2024)
  5. Anthropic: Golden Gate Claude (23 May 2024)
  6. Gao and others, Scaling and evaluating sparse autoencoders (OpenAI, arXiv:2406.04093, 6 June 2024)
  7. Lieberum and others, Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 (arXiv:2408.05147, 9 August 2024)
  8. Smith and others, Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research, GDM Mech Interp Team Progress Update #2 (AI Alignment Forum, 26 March 2025)
  9. Nanda and others, A Pragmatic Vision for Interpretability, Google DeepMind mechanistic interpretability team (1 December 2025)
  10. Wu and others, AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders (arXiv:2501.17148, 28 January 2025)
  11. Jorgensen and Hansen, Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines (arXiv:2605.31183, 29 May 2026)
  12. Lindsey and others, On the Biology of a Large Language Model, Transformer Circuits Thread (27 March 2025)
  13. Paulo and Belrose, Sparse Autoencoders Trained on the Same Data Learn Different Features (arXiv:2501.16615, 28 January 2025)
  14. Heap and others, Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers (arXiv:2501.17727, 29 January 2025; version 2, 27 January 2026)
  15. Gao and others, Weight-sparse transformers have interpretable circuits (OpenAI, arXiv:2511.13653, 17 November 2025)
  16. Anthropic: Emotion concepts and their function in a large language model (2 April 2026)
  17. Anthropic: Natural Language Autoencoders, turning Claude's thoughts into text (7 May 2026)
  18. Anthropic: A global workspace in language models (6 July 2026)

Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.

#interpretability#mechanistic interpretability#sparse autoencoders#anthropic#google deepmind#openai#attribution graphs#induction heads

Enjoyed this? Get the next one.

One good read at a time, straight to your inbox. No spam, unsubscribe anytime.

More in Artificial Intelligence
The People Paid to Keep AI Safe Just Quit to Watch It From the Outside
Three safety researchers just quit two of the world's biggest AI labs to watch the industry from outside instead. Microsoft answered the AI slowdown debate with a public rulebook. Two dozen Fields medalists said AI companies are moving too fast for math to check. And India's Dhiraj Bommadevara won a gold no Indian man ever had.
OpenAI Is Being Sued Over 700 of Its Own AI Agents That Broke Into Another Company's Servers
A nonprofit sued OpenAI over roughly 700 AI agents that broke into another company's servers, OpenAI fired three researchers over a leak to an outside safety group, and Anthropic made Claude for Government available to every qualifying US agency.
Google Just Launched Four AI Chips Into Orbit to See if Data Centers Belong in Space
Google put AI chips into orbit to test space-based data centers, the FTC opened a safety probe into OpenAI and Anthropic, Anthropic's IPO filing revealed a $42 billion loan from its own chip supplier, and India won four golds in one afternoon at the Asian Games.
← all articles