Customize a training run#
Customization starts from a typed YAML config. Pick the closest checked-in
file under
examples/configs,
change model and data paths, and keep the same training entry point:
specforge train --config ./my-run.yaml
For one-off changes, use dotted overrides rather than adding another launcher:
specforge train \
--config examples/configs/qwen3-8b-eagle3-disaggregated.yaml \
model.target_model_path=/models/my-target \
data.train_data_path=/datasets/my-training-data.jsonl \
training.learning_rate=5e-5
The config is strict: unknown fields, invalid strategy/topology combinations,
and unsupported attention backends are errors. See
specforge/config/schema.py and the
training guide for the accepted fields.
Chat templates#
Register a text chat template in TEMPLATE_REGISTRY in
specforge/data/template.py, then reference its name with
data.chat_template:
TEMPLATE_REGISTRY.register(
name="your-template-name",
template=ChatTemplate(
assistant_header="...",
user_header="...",
system_prompt="...",
end_of_turn_token="...",
),
)
data:
train_data_path: /datasets/my-training-data.jsonl
chat_template: your-template-name
max_length: 4096
The server capture runtime accepts text input only. VLM training, including
qwen2_5_vl, is not supported. Attention backends are a closed,
strategy-specific set:
Strategy |
Accepted |
|---|---|
EAGLE3 |
|
P-EAGLE |
|
DFlash, Domino, DSpark |
|
Target models#
For a Hugging Face-compatible target, normally only these model fields need to change:
model:
target_model_path: organization/model-name
target_backend: sglang
trust_remote_code: false
embedding_key: model.embed_tokens.weight
lm_head_key: lm_head.weight
Every online run uses model.target_backend: sglang. Add target-model support
to the SGLang capture server instead of adding an HF/custom target loader to
the trainer. Target TP/EP and model-specific inference stay on that server;
the SpecForge consumer receives only the algorithm’s versioned feature schema.
Offline feature training performs no target inference.
EAGLE3 offline sequence parallelism is selected with
training.attention_backend: usp plus training.sp_ulysses_size and
training.sp_ring_size. Evaluation, compact-teacher projection, and experiment
tracking are also config features rather than custom launchers; see the
training guide for their validated combinations.
Draft architectures#
Draft classes register through @register_draft. The key defaults to the
class name and must match the single entry in the draft JSON’s
architectures list:
from transformers import PretrainedConfig
from specforge.modeling.draft.base import Eagle3DraftModel
from specforge.modeling.draft.registry import register_draft
class MyDraftConfig(PretrainedConfig):
model_type = "my-draft"
@register_draft
class MyEagle3Draft(Eagle3DraftModel):
config_class = MyDraftConfig
def __init__(self, config, **kwargs):
super().__init__(config)
...
Import the module from specforge/modeling/draft/__init__.py so registration
runs before config resolution. A minimal draft config then contains:
{
"architectures": ["MyEagle3Draft"],
"model_type": "my-draft",
"vocab_size": 128256,
"draft_vocab_size": 32000
}
Point model.draft_model_config at that JSON. AutoDraftModelConfig and the
model assembler resolve both the config and model class from the registry; do
not add a method-specific training launcher.
An architecture alone does not define a new loss. A genuinely new training
algorithm also needs a pure AlgorithmSpec, executable AlgorithmProviders,
and one immutable AlgorithmRegistration under specforge/algorithms. Its
step provider may construct a DraftTrainStrategy, while its model and data
providers own algorithm-specific assembly. Add builtin registrations to the
explicit catalog used by the application composition root; do not add a
method-specific launcher or a second mutable registry.