← AI Terminology
Mid-Training
Mid-training is a staged training phase between initial pretraining and late alignment/SFT where models see curated mixtures (math, code, long context, high-quality web) to shape capabilities.
The term is increasingly used in modern LLM training reports.
The term is increasingly used in modern LLM training reports.
Why It Matters in AI
Raw pretrain data is noisy; mid-training upsamples quality and skill domains before expensive alignment. Understanding multi-phase recipes (pretrain → mid → SFT → RL) is key to reading 2024–2026 model reports.
Key Points
| Aspect | Description |
|---|---|
| Goal | Boost target skills; extend context |
| Content | High-quality, long-context, STEM/code heavy mixes |
| Related | Annealing phases, continued pretraining |
| Evidence | Described in various open-model technical reports |
| Position | After broad pretrain, before instruction/RL |
| Practice | Data mixture schedules matter as much as size |
Simple Analogy
After general education, an honours semester of focused advanced courses before professional job training — reshaping skills mid-stream.
Common Usage Examples
- Technical reports: mid-train on math/code blends
- Long-context mid-train extension
- Quality filtering before alignment
- Evaluate skill benchmarks after mid-phase
Summary
In short: Mid-training is the curated capability-shaping phase between raw pretraining and alignment — where data mixture steers what the model becomes good at.