← AI Terminology

Mid-Training

Mid-training is a staged training phase between initial pretraining and late alignment/SFT where models see curated mixtures (math, code, long context, high-quality web) to shape capabilities.

The term is increasingly used in modern LLM training reports.
Why It Matters in AI
Raw pretrain data is noisy; mid-training upsamples quality and skill domains before expensive alignment. Understanding multi-phase recipes (pretrain → mid → SFT → RL) is key to reading 2024–2026 model reports.
Key Points
Aspect Description
Goal Boost target skills; extend context
Content High-quality, long-context, STEM/code heavy mixes
Related Annealing phases, continued pretraining
Evidence Described in various open-model technical reports
Position After broad pretrain, before instruction/RL
Practice Data mixture schedules matter as much as size
Simple Analogy
After general education, an honours semester of focused advanced courses before professional job training — reshaping skills mid-stream.
Common Usage Examples
  • Technical reports: mid-train on math/code blends
  • Long-context mid-train extension
  • Quality filtering before alignment
  • Evaluate skill benchmarks after mid-phase
Summary
In short: Mid-training is the curated capability-shaping phase between raw pretraining and alignment — where data mixture steers what the model becomes good at.