Inside Gemma 3: Modifying the output through activation hacking

Google for DevelopersAbout 4 min readApr 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Transformer model mechanics
  • Gemma 3 model
  • Residual stream
  • Embeddings
  • Feedforward layers
  • Attention blocks
  • Token generation
  • Neuron activation
  • Gate activation
  • Bias manipulation
  • Premature decoding
  • Model investigation
  • Code switching

1. Model Setup and Configuration:

  • The video focuses on exploring the internal mechanics of the Gemma 3 transformer model.
  • The smallest Gemma 3 model with 1 billion parameters is used for demonstration.
  • The Colab notebook available in the Gemma Cookbook is used for the experiments.
  • The "Sow Configuration" allows surfacing intermediate values, specifically the residual stream and top-activated neurons in feedforward layers.
  • The embedding layer, representing the initialization of the residual stream, and intermediate values before and after each attention block are enabled.

2. Initial Model Inquiry and Output:

  • The model is initialized with a batch size of 1 to generate a single token for easier investigation.
  • The prompt "What is the capital of Switzerland?" is used.
  • The model correctly answers "Bern," acknowledging that Switzerland doesn't have an officially designated capital but Bern is the de facto capital.

3. Visualizing Intermediate Values with Premature Decoding:

  • A helper function is used to perform "premature decoding" by applying the final normalization (LayerNorm) and softmax layer to the residual stream at each intermediate output.
  • This allows visualizing how the model progresses through its layers.
  • The visualization shows alternating patterns of attention layers (red) and feedforward layers (blue) after the initial embedding layer (green).
  • The tokens in the lower layers are shifted by one, representing previously generated tokens used to initialize the residual stream.
  • The model initially considers predicting Zurich or Geneva before shifting towards Bern in the final layer.

4. Investigating Top-Activated Neurons:

  • The focus shifts to the final feedforward layer (layer 25) to understand why the model ultimately predicts Bern.
  • The top two activated neurons in layer 25 are identified as 1937 (value around -10) and 4422 (value around -5).

5. Manipulating Neuron Activation by Adjusting Bias:

  • Neuron 1937 is deactivated by adjusting the bias of the gate projection.
  • Gemma 3 uses a gate activation, so a large negative bias at neuron 1937 results in a gate value of 0, effectively deactivating the neuron.
  • Running the model again after deactivating neuron 1937 results in the model predicting "Zurich" as the capital of Switzerland.
  • The negative activation of neuron 1937 indicates that it was suppressing something from the residual stream.

6. Analyzing Suppressed Tokens:

  • The premature decoding function is applied to the value of neuron 1937 (the corresponding column in its downward prediction matrix) to check what was suppressed.
  • The top 20 tokens associated with that value are examined.
  • Several variations of "Swiss" and "Switzerland" are present, along with the Chinese and Russian spellings of Switzerland.
  • Crucially, "Zurich" appears at position 7 and "Geneva" at position 12, indicating that suppressing them shifted the response towards Bern.
  • "Argentina" is also encoded into the value of neuron 1937, potentially due to similarities in geographic landscapes between Switzerland and Argentina.

7. Boosting Neuron Activation:

  • The video demonstrates that neurons can not only be deactivated but also boosted.
  • Boosting a neuron too much can lead to the model outputting the same token at every step.
  • The model is modified to boost neuron 1937.
  • When prompted with "What is the richest country on the moon?", the modified model answers "Switzerland."
  • When prompted with "What is the largest country on the moon?", the modified model answers "Argentina."

8. Conclusion:

  • The video provides a glimpse into the internal mechanics of the Gemma 3 transformer model.
  • It demonstrates how manipulating neuron activations can influence the model's output.
  • The presenter encourages further experimentation with the Gemma 3 model and the provided Colab notebook to discover other interesting patterns and potential applications.
  • Other interesting aspects of the model, such as code switching and the attention mechanism, are mentioned as areas for further exploration.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.