← AI Terminology
Superposition (Neural)
Superposition is the hypothesis that neural networks represent more features than they have neurons by packing features in overlapping non-orthogonal directions.
A central concept in mechanistic interpretability.
A central concept in mechanistic interpretability.
Why It Matters in AI
If features superpose, individual neurons are uninterpretable mixtures. This motivates SAEs and sparse coding views of activations, reshaping how researchers think about scaling and interpretability.
Key Points
| Aspect | Description |
|---|---|
| Idea | Overcomplete features compressed into fewer dimensions |
| Related | SAE, mech interp, polysemanticity |
| Evidence | Toy models; empirical polysemantic neurons |
| Tradeoff | Interference vs capacity for many concepts |
| Implication | Need dictionary learning for clean features |
| Paper lineage | Anthropic toy models of superposition |
Simple Analogy
Storing more radio stations than you have clean channels by overlapping transmissions — more content, more crosstalk.
Common Usage Examples
- Read toy models of superposition
- Polysemantic neuron examples
- SAEs as attempted un-superposition
- Discuss capacity vs interference
Summary
In short: Superposition packs more features than neurons via overlapping directions — explaining polysemantic units and motivating sparse feature methods.