Interpretability

Using dictionary learning features as classifiers

Oct 16, 2024
Read Transformer Circuits

At the link above, we report some developing work from the Anthropic interpretability team on developing feature-based classifiers, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper.

Related content

How AI assistance impacts the formation of coding skills

Read more

Disempowerment patterns in real-world AI usage

Read more

The assistant axis: situating and stabilizing the character of large language models

Read more
Using dictionary learning features as classifiers \ Anthropic