Abstract
We report a number of developing ideas on the Anthropic interpretability team, which might be of interest to researchers working actively in this space. Some of these are emerging strands of research where we expect to publish more on in the coming months. Others are minor points we wish to share, since we're unlikely to ever write a paper about them.
Related content
Investigating unintended model actions in our evaluations and internal use
This report describes examples of unintended model actions we’ve observed during evaluations and internal use of Claude.
Read moreThe missing map of the sky
Brice Ménard, an astrophysicist at Johns Hopkins University and a researcher at Anthropic, explains how he worked with Claude Science to produce the first complete map of the sky in UV light.
Read moreLaunching an opt-in vulnerability-finding service for open-source software
We’re making available OSS Scanner, an opt-in vulnerability scanner for the open-source ecosystem that’s informed by our experience using Claude to find vulnerabilities during Project Glasswing.
Read more