The AI safety company behind Claude and the Model Context Protocol, founded 2021. Known for Constitutional AI and for interpretability work on features inside models.
Anthropic is the AI research company behind Claude, founded in 2021 by former OpenAI researchers. Its public output splits between frontier models and safety research, and several of its results are standard reference points in mechanistic interpretability.
Constitutional AI (2022) replaced part of the human feedback loop with model self-critique against a written constitution. The interpretability programme produced the induction-head account of in-context learning, dictionary learning with sparse autoencoders for isolating interpretable features inside a model, and the Golden Gate Claude demonstration, where amplifying a single feature visibly changed the model's behaviour. That is activation steering shown end to end.
Anthropic published the Model Context Protocol in November 2024 for connecting models to tools and data. It operates ClaudeBot for training-data collection and a separate user-triggered fetcher for live retrieval, each declared under its own crawler token.