Intrusive Thought

A real Gemma 2 (2B) has a dial for each concept it knows. Drag one and the model leans toward it — or away from it — until you let go and it’s exactly, provably itself again, because nothing was ever written to a weight. Every dial below does that same thing to the model. What changes is only how the dial's direction was found.

SAE feature — a sparse autoencoder already learned this direction from the model's own activations, one of ~16,000 candidates.
ActAdd contrast pair — no autoencoder: the direction is just the difference between two prompts, written by hand.
loading…