Golden Gate Claude: a journey inside the brain of an AI model
Anthropic recently published a study on how the Claude 3 artificial intelligence model works.
This study has given us useful insights into how large language models work, showing that some of their internal mechanisms can be understood and even manipulated.
Inside Claude's mind, millions of concepts were found, which activate when the model reads a particular text or sees a particular image. These groups of neurons inside an artificial intelligence model, which fire in response to a specific concept, are called features.
To understand this better, let's imagine the artificial intelligence model as a huge mosaic. Each tile of the mosaic represents a neuron. When the model processes information, for example by reading a text or looking at an image, some of these tiles (neurons) "light up", that is, they activate.
A "feature" is a group of tiles that activate together in a meaningful way. For example, there could be a specific feature for the concept of "cat". Every time the model encounters the image of a cat, or reads the word "cat" in a text, the neurons making up this "feature" activate.
"Features" do not only represent concrete objects like cats. They can also represent more abstract concepts such as "happiness", "sadness", "irony" or even writing styles such as "formal" and "informal".
One of these features, the subject of Anthropic's study, is the concept of the Golden Gate Bridge: a specific combination of neurons was found in Claude's neural network that activates whenever the famous San Francisco bridge is mentioned or a photo of it is shown.
Not only is it possible to identify these features, but it is also possible to increase or decrease their intensity and identify the corresponding changes in Claude's behaviour.
As explained in the study, when the intensity of the Golden Gate Bridge feature is increased, Claude's answers start to focus on the Golden Gate Bridge, even when it makes no sense!
If you ask Golden Gate Claude how to spend 10 dollars, it will recommend using them to pay the toll to cross the Golden Gate Bridge. If you ask it to write a love story, it will tell the story of a car that cannot wait to cross its beloved bridge on a foggy day. If you ask it what it imagines it looks like, it will reply that it imagines looking like the Golden Gate Bridge.
Even we ordinary users can talk to "Golden Gate Claude" on claude.ai: the goal is to show people the impact that identifying and modifying these features inside an AI model can have. It also shows that we are really starting to understand how large language models actually work.
Claude changed the way it acts not thanks to a finely crafted custom prompt (e.g. "from now on pretend to be a bridge"), but thanks to a precise, surgical modification of some aspects of the model.
In short, the text published by Anthropic announces a significant step forward in the interpretability of AI models, paving the way for greater understanding and control over how these technologies make decisions.
Let's try to understand the heart of this study a little better, and why it matters.
What is Interpretability?
In the field of artificial intelligence, and in particular with regard to machine learning models, interpretability refers to the ability to understand how and why a model reaches a given decision.
In essence, interpretability lets us open the "black box" of machine learning models, making them more transparent and understandable.
Why is interpretability important?
- Trust: understanding how a model works allows us to trust its decisions more, especially in critical fields such as medicine or finance
- Debugging: if a model makes mistakes, interpretability helps us identify the cause and fix it
- Ethics: interpretability is essential to ensure that AI models are not discriminatory or unfair in their decisions
Interpretability is a crucial aspect in the development of responsible and reliable artificial intelligence.
How was the interpretability of the Golden Gate Feature demonstrated?
Through Specificity, that is, by verifying that the activation of a feature actually corresponds to the presence of the associated concept in the text: the "Golden Gate Bridge" feature activated only when the famous bridge was being discussed.
And also through Influence on behaviour: it is verified that, by artificially modifying the activation of a feature, changes in the model's behaviour are obtained that are consistent with the associated concept. For example, by increasing the activation of the "Golden Gate Bridge" feature, the model should talk about the bridge more often. This uses a technique called "feature steering", which consists of artificially modifying the activation of a feature while the text is being processed. For example, by increasing the activation of the "Golden Gate Bridge" feature, the model starts talking about the bridge even when it is not relevant to the conversation.
Results: the features analysed show a high degree of specificity, activating mainly in the presence of the associated concepts. Furthermore, the "feature steering" technique demonstrates that it is possible to influence the model's behaviour in a way that is consistent with the interpretation of the features.
This study has shown that some features (besides the Golden Gate, also neuroscience, famous monuments, trains, bridges...) extracted from large language models are actually interpretable, paving the way for a better understanding of how they work. The ability to identify and manipulate specific features could lead to more transparent, reliable and safe artificial intelligence systems.
As described in Anthropic's article, these same techniques can be used to modify safety-related features, such as those relating to dangerous computer code, criminal activity or deception. With further research, it is believed that this work can help make AI models safer.
