WritingAnthropicAnthropicpublished Oct 16, 2024seen 2d

Features As Classifiers

Open original ↗

Captured source

source ↗
published Oct 16, 2024seen 2dcaptured 8hhttp 200method plain

Using dictionary learning features as classifiers \ Anthropic Interpretability Using dictionary learning features as classifiers Oct 16, 2024 Read Transformer Circuits

At the link above, we report some developing work from the Anthropic interpretability team on developing feature-based classifiers, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper.

Related content

Paving the way for agents in biology Read more Making Claude a chemist Read more Coding agents in the social sciences Results from a survey of 1,260 social scientists about AI and coding agent use. Read more