Features As Classifiers
Captured source
source ↗Using dictionary learning features as classifiers \ Anthropic Interpretability Using dictionary learning features as classifiers Oct 16, 2024 Read Transformer Circuits
At the link above, we report some developing work from the Anthropic interpretability team on developing feature-based classifiers, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a colleague sharing some thoughts or preliminary experiments for a few minutes at a lab meeting, rather than a mature paper.
Related content
Paving the way for agents in biology Read more Making Claude a chemist Read more Coding agents in the social sciences Results from a survey of 1,260 social scientists about AI and coding agent use. Read more