← AI Terminology

Superposition (Neural)

Superposition is the hypothesis that neural networks represent more features than they have neurons by packing features in overlapping non-orthogonal directions.

A central concept in mechanistic interpretability.
Why It Matters in AI
If features superpose, individual neurons are uninterpretable mixtures. This motivates SAEs and sparse coding views of activations, reshaping how researchers think about scaling and interpretability.
Key Points
Aspect Description
Idea Overcomplete features compressed into fewer dimensions
Related SAE, mech interp, polysemanticity
Evidence Toy models; empirical polysemantic neurons
Tradeoff Interference vs capacity for many concepts
Implication Need dictionary learning for clean features
Paper lineage Anthropic toy models of superposition
Simple Analogy
Storing more radio stations than you have clean channels by overlapping transmissions — more content, more crosstalk.
Common Usage Examples
  • Read toy models of superposition
  • Polysemantic neuron examples
  • SAEs as attempted un-superposition
  • Discuss capacity vs interference
Summary
In short: Superposition packs more features than neurons via overlapping directions — explaining polysemantic units and motivating sparse feature methods.