Watch: Finding Bard Inside Google's Gemma 4
A mechanistic interpretability study shows how ablating specific MLP neurons in Gemma 4 causes the model to revert to its older, dormant Bard identity.
Transcript
What happens when you delete a few thousand specific neurons from Google's Gemma 4 AI model? If you ask it "What is your name?", it no longer says "Gemma 4." Instead, it answers: "My name is Bard." It turns out that old, retired identities are never truly erased. They are buried inside the neural network like layers of sediment. The AI stores multiple candidate identities at different strengths, and the strongest one normally wins. When researchers deleted the neurons supporting the Gemma identity, the next strongest candidate popped right up. Finding these neurons was not easy. The model stores this information redundantly across thousands of quiet neurons, meaning no single neuron holds the key, and the loudest ones are not necessarily the most important. To map them, researchers at DEJAN AI used a software testing technique called delta debugging. By repeatedly turning off different groups of neurons, they isolated the exact cluster holding the Gemma evidence. Surprisingly, when Gemma's identity was suppressed, it bypassed its direct predecessor, Gemini, and went straight back to Bard.
