Skip to content

Forge Draft Models for Speculative Decoding

SpecForge is the SGLang-native framework for training EAGLE3, DFlash, DSpark and Domino draft models. Train once, export once, serve at speed.

SGLang-ready

Draft models trained with SpecForge export straight into SGLang serving. Most methods use specforge export; MTP merges via scripts/merge_mtp_to_base.py.

Every draft family

EAGLE3, P-EAGLE, DFlash, DFlash2, DSpark, Domino and MTP drafts all train through one runtime with shared data, evaluation and export.

Any scale, any accelerator

Online and offline training, colocated or disaggregated, from a single GPU to multi-server capture on NVIDIA CUDA, AMD ROCm and Ascend NPU.

Get Started in Seconds

Select your hardware and installer, run the command, and you are ready to train your first draft model.

  1. Install SpecForge

    From source or PyPI, on the accelerator you have.

  2. Prepare a dataset

    Regenerate responses with the target model so the draft learns its distribution.

  3. Train the draft

    Point specforge train at a YAML recipe; capture hidden states online or offline.

  4. Export and serve

    Run specforge export and launch SGLang with the draft for faster decoding.

View Documentation →
Hardware
Installer
Package
Run this command:
git clone https://github.com/sgl-project/SpecForge.git
cd SpecForge
uv venv -p 3.11 --seed
source .venv/bin/activate
uv pip install -e .

Install a CUDA build of PyTorch that matches the host driver. Installing from source is recommended so you get the latest recipes and patches. Installation guide →

Broad Method & Hardware Support

One training runtime across draft-model families and accelerators.

Draft Model Families

EAGLE3P-EAGLEDFlashDFlash2DSparkDominoMTP
See the training guide →

Supported Hardware

NVIDIA GPUsAMD GPUsAscend NPUs
See accelerator setup →

Production-grade draft models, ready to serve

SpecBundle collects 30 draft models for mainstream open LLMs, contributed by the SGLang team and industry partners including Ant Group, Meituan, Nex-AGI, EigenAI and RadixArk. Every model ships with SGLang launch flags and benchmark results.

LlamaQwenKimiDeepSeekGLMgpt-ossStepLingInkling

Released under the MIT License.