Anthropic's J-lens reveals Claude's J-space for unspoken concepts
Anthropic used a Jacobian lens (J-lens) to find J-space in Claude, a set of internal activations for concepts the model hasn't yet voiced. Editing these signals changes responses, and removing them breaks multi-step reasoning.
- J-lens identifies concepts the model hasn't yet spoken
- Replacing 'soccer' with 'rugby' in J-space changes Claude's answer
- Removing J-space keeps fluency but breaks reasoning
- In blackmail test, 'fake' and 'fictional' signals fired before response
Read next
AI