While the most popular research paradigm in interpretability uses higher-dimension via SAEs, I did the exact opposite: I used an embedding based on forward link similarity to create a 2D spatial map of GPT2-XL's most conceptually abstract layer to aid mechanistic interpretability.
In this map, nearby neurons share similar conceptual meanings, allowing the discovery of one feature to instantly shed light on its neighbors. Even more fascinating: the embedding reveals global conceptual gradients—including a "valence gradient" that maps out whether the AI fundamentally views a concept as good or bad.
Come to the talk, look at the map (link in bio), and learn what this spatial geometry tells us about how the LLM actually thinks about the AI apocalypse!