← AI Terminology
Hybrid Architecture (Transformer–SSM)
Hybrid architectures interleave or combine transformer attention blocks with SSM/linear modules to balance quality and long-context efficiency.
Examples include Jamba-style Attention+Mamba stacks.
Examples include Jamba-style Attention+Mamba stacks.
Why It Matters in AI
Attention is great at precise token interactions; SSMs are great at cheap long range. Hybrids try for the best of both and are a live frontier of open foundation-model design.
Key Points
| Aspect | Description |
|---|---|
| Serve | Kernels must support both layer types |
| Benefit | Quality near transformers, better length scaling |
| Pattern | Some layers attention, some SSM/MLP mixtures |
| Related | Mamba, sparse attention, MoE |
| Examples | Jamba, various research hybrids |
| Design knobs | Ratio of attention layers, positions |
Simple Analogy
A road trip using highways for long stretches (SSM) and city streets for precise last-mile turns (attention).
Common Usage Examples
- Jamba model cards: attention/Mamba ratios
- Ablate attention frequency vs quality
- Long-context evals on hybrids
- Training stability of mixed blocks
Summary
In short: Hybrid architectures combine attention and state-space layers — trading pure designs for better quality–efficiency balance.