Watch: Anthropic

The AI safety company behind Claude and the Model Context Protocol, founded 2021. Known for Constitutional AI and for interpretability work on features inside models.

Transcript

Anthropic, founded in 2021 by former OpenAI researchers, is the artificial intelligence company behind the Claude models. The company balances building powerful frontier models with pioneering safety research.

In 2022, they introduced Constitutional AI, a method that replaced part of the human feedback loop with a system where the AI critiques itself based on a written constitution.

Anthropic has also made major breakthroughs in understanding how these models actually think. This field, known as mechanistic interpretability, includes their work on isolating specific concepts inside a model. In their Golden Gate Claude demonstration, they proved they could directly steer the AI's behavior by amplifying a single concept.

More recently, in late 2024, the company published the Model Context Protocol to help connect models to external tools and data. They also run ClaudeBot to collect training data, alongside a separate, user-triggered fetcher for live retrieval.