Anthropic
Safety-focused AI lab and creator of the Claude family, founded by former OpenAI researchers.
Mission
To be the world's most reliable and interpretable AI lab.
Founded By
Key Products & Research
- Claude (Claude 1, 2, 3, Sonnet, Opus, Haiku)
- Constitutional AI
Headquarters
San Francisco, California, USA
Founded
2021
Status
Active
Contribution to AI
Interpretability was a minority interest in 2021, applied mostly to vision models and mostly after the fact. The transformer circuits line of work changed that: a 2021 framework for reading small attention-only models as compositions of legible operations, then induction heads as a concrete mechanism behind in-context learning, then superposition as an explanation for why individual neurons resist interpretation at all. Sparse autoencoders grew out of that argument, and by 2024 were extracting tens of millions of features from a deployed production model — the point at which mechanistic interpretability stopped being a study of toy networks. A second contribution is procedural rather than technical. Constitutional AI, published in 2022, replaced much of the human labelling in harmlessness training with a model critiquing and revising its own outputs against a written set of principles, which turned the rules governing a system's behaviour into a document that could be read and disputed rather than a distribution implicit in contractor judgements. The Responsible Scaling Policy of September 2023 applied the same instinct to deployment: capability thresholds defined and published before the models that would cross them existed, with required safeguards attached to each. Comparable frameworks appeared at rival labs within the year. The Model Context Protocol, opened in late 2024, was the smallest of these and spread quickest, giving models one common way to reach external tools and data.
Drafted with AI and edited by hand (claude-opus-5, reviewed 2026-08).