← AI Terminology

Hybrid Architecture (Transformer–SSM)

Hybrid architectures interleave or combine transformer attention blocks with SSM/linear modules to balance quality and long-context efficiency.

Examples include Jamba-style Attention+Mamba stacks.
Why It Matters in AI
Attention is great at precise token interactions; SSMs are great at cheap long range. Hybrids try for the best of both and are a live frontier of open foundation-model design.
Key Points
Aspect Description
Serve Kernels must support both layer types
Benefit Quality near transformers, better length scaling
Pattern Some layers attention, some SSM/MLP mixtures
Related Mamba, sparse attention, MoE
Examples Jamba, various research hybrids
Design knobs Ratio of attention layers, positions
Simple Analogy
A road trip using highways for long stretches (SSM) and city streets for precise last-mile turns (attention).
Common Usage Examples
  • Jamba model cards: attention/Mamba ratios
  • Ablate attention frequency vs quality
  • Long-context evals on hybrids
  • Training stability of mixed blocks
Summary
In short: Hybrid architectures combine attention and state-space layers — trading pure designs for better quality–efficiency balance.