SpecForge is the SGLang-native framework for training EAGLE3, DFlash, DSpark and Domino draft models. Train once, export once, serve at speed.
Draft models trained with SpecForge export straight into SGLang serving. Most methods use specforge export; MTP merges via scripts/merge_mtp_to_base.py.
EAGLE3, P-EAGLE, DFlash, DFlash2, DSpark, Domino and MTP drafts all train through one runtime with shared data, evaluation and export.
Online and offline training, colocated or disaggregated, from a single GPU to multi-server capture on NVIDIA CUDA, AMD ROCm and Ascend NPU.
Select your hardware and installer, run the command, and you are ready to train your first draft model.
From source or PyPI, on the accelerator you have.
Regenerate responses with the target model so the draft learns its distribution.
Point specforge train at a YAML recipe; capture hidden states online or offline.
Run specforge export and launch SGLang with the draft for faster decoding.
git clone https://github.com/sgl-project/SpecForge.git
cd SpecForge
uv venv -p 3.11 --seed
source .venv/bin/activate
uv pip install -e .Install a CUDA build of PyTorch that matches the host driver. Installing from source is recommended so you get the latest recipes and patches. Installation guide →
One training runtime across draft-model families and accelerators.

SpecBundle collects 30 draft models for mainstream open LLMs, contributed by the SGLang team and industry partners including Ant Group, Meituan, Nex-AGI, EigenAI and RadixArk. Every model ships with SGLang launch flags and benchmark results.
From a first training run to production draft models, the SpecForge community is open to everyone.