Interpretability

Using dictionary learning features as classifiers

Oct 16, 2024
Read Transformer Circuits

At the link above, we report some developing work from the Anthropic interpretability team on developing feature-based classifiers, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper.

Related content

Introducing Anthropic Interviewer: What 1,250 professionals told us about working with AI

Read more

How AI is transforming work at Anthropic

How AI Is Transforming Work at Anthropic

Read more

Estimating AI productivity gains from Claude conversations

Read more