Qwen3.8-27B DFlash2 disaggregated training
Trains the DFlash2 drafter in configs/qwen3.8-27b-dflash2.json for Qwen/Qwen3.8-27B on the disaggregated online data plane: patched SGLang capture servers publish the target's hidden states into Mooncake, and data-parallel trainer ranks consume them. The first two recipes share every model, data and optimizer setting and differ only in topology; the third is the tuned single-node H200 layout described in One H200 node, tuned:
| Recipe | Layout | Services |
|---|---|---|
managed-local/qwen3.8-27b-dflash2-4server-dp4-disaggregated.yaml | one 8-GPU node: four TP1 capture servers on GPUs 0-3, DP4 trainer on GPUs 4-7 | owned by specforge train |
managed-local/qwen3.8-27b-dflash2-h200-5server-dp3-disaggregated.yaml | one 8x H200 node: five TP1 FP8 capture servers on GPUs 0-4, DP3 trainer on GPUs 5-7, RDMA loopback | owned by specforge train |
external/qwen3.8-27b-dflash2-disaggregated.yaml | two 8-GPU nodes: eight TP1 capture servers, the Mooncake master and the producer on the capture node, DP8 trainer on the trainer node | started by examples/disagg/run_qwen3.8_27b_dflash2_disagg_2node.sh or by hand |
Every number below is a steady-state throughput in samples per second (one sample is one conversation of up to 8192 tokens), measured in September 2026 on 8x and 2x8 B300 nodes and on 8x and 2x8 H200 nodes over 150-300 optimizer steps per configuration. The recipes carry the settings of those runs; the sections after the launch commands explain which settings follow from the payload size and which from the GPU generation.
Required source revisions
- SpecForge
mainat or after #797: the feature loader prefetches withdata.dataloader_num_workersconcurrent Mooncake reads and retries transient read failures. The in-flight watermarks in the recipes are sized for that loader. - SGLang
v0.5.18with the checked-in capture patch on every capture host:scripts/apply_sglang_spec_capture_patch.sh --target v0.5.18. Capture servers run--spec-capture-method dflashwith the draft's auxiliary layers5 19 33 47 61; the DFlash2 selector exists only on the trainer. - Mooncake
mooncake-transfer-engine0.3.13 withmooncake_masteronPATH. Measured stack: PyTorch 2.13.0+cu130, transformers 5.12.1,mooncake-transfer-engine-cuda130.3.13.post1.
Target, draft and data
Target. Qwen/Qwen3.8-27B: 64 layers, hidden size 5120, hybrid linear and full attention, BF16 weights. The H200 measurements used Qwen/Qwen3.8-27B-FP8, whose lm_head and embeddings are stored in BF16, so the trainer loads the frozen head unchanged and the capture command needs no extra flag (SGLang reads the checkpoint's quantization config). The B300 measurements used an NVFP4 export with a separately loaded BF16 head; main loads model.lm_head_key and model.embedding_key from the target checkpoint itself, so use the BF16 or FP8 checkpoint with these recipes.
Draft. Five Qwen3-style GQA decoder layers at the target width (32 query heads, 8 KV heads, head dimension 128, sliding window 2048) with DFlash2 block size 8, grouped convolutions (group 16, kernel 2) and a rank-256, top-16 candidate selector; about 1.9B parameters. It reads the target's hidden states after layers 5, 19, 33, 47 and 61 plus the last hidden state. The layout matches the public incoai/Qwen3.8-27B-DFlash2 checkpoint, so exported checkpoints serve through SGLang's existing DFlash2 path. The mask token is 248070 (<|audio_start|>, never produced by text), as in the Qwen3.6-27B recipes.
Data. The recipes read ./cache/dataset/qwen3.8-27b-regen-train.jsonl: ShareGPT-style conversations whose assistant turns were regenerated by Qwen3.8-27B itself, with reasoning kept, formatted by chat_template: qwen3.5. On-policy responses matter for block-wise drafts, whose acceptance on off-policy answers plateaus far below what the same draft reaches on regenerated data. scripts/regenerate_train_data.py, wrapped by examples/data_regeneration/run_qwen_sharegpt_regeneration.sh (MODEL_PROFILE=qwen3.6-27b MODEL_PATH=Qwen/Qwen3.8-27B), produces such a file from cache/dataset/sharegpt_train.jsonl. The measured corpus held about 1.07M conversations with a mean of 4.5k tokens and its 90th percentile at data.max_length: 8192.
Payload. One sample carries six BF16 hidden-state tensors of tokens x 5120: about 277 MB at the corpus mean and 500 MB at 8192 tokens. Every Mooncake and flow-control value below follows from that size.
One node: four servers + DP4
specforge train -c \
examples/configs/online/disaggregated/managed-local/qwen3.8-27b-dflash2-4server-dp4-disaggregated.yamlAdd --plan first to validate the schema and print the owned Mooncake, capture-server, producer and consumer commands without starting anything. The launcher starts Mooncake, four capture servers on GPUs 0-3, then the producer and the four trainer ranks on GPUs 4-7, and tears the stack down when the consumer finishes. Logs live under the recipe's control_dir/logs/.
Measured single-node throughput of the recipe and of the two neighbouring splits, all with the same flow control:
| 8 GPUs | 4 servers + DP4 | 5 servers + DP3 | 3 servers + DP5 | colocated in-process capture (#783), for reference |
|---|---|---|---|---|
| 8x H200, FP8 target | 9.05 | 10.65 | 6.82 | 8.61 |
| 8x B300, NVFP4 target | 11.5 | 10.3 | 9.8 | 18.9 |
Why the best split differs per GPU generation: an H200 FP8 capture server delivers 2.25 prompts/s per GPU while a trainer rank consumes 3.7 samples/s (eight samples per 2.15 s step), so five servers are needed to feed three trainers; at 4+4 the servers are 98% busy and the trainers wait. A B300 server alone captures 7.2 prompts/s per GPU and a rank trains 4.7 samples/s, but with the host receive path on main each server's rate falls to 2-3 prompts/s once several trainer ranks pull from the same host, which caps the node at about 11.5 samples/s and makes 4+4 the single-node optimum there. The H200 5+3 and 3+5 points were run with the device receive path of #840 and #841; on H200 the pageable, pinned and device receive paths landed within 0.02 samples/s of each other at 4+4 because the servers, not the transport, are the limit, so the ordering holds for main. To move servers, edit capture_servers, trainer_cuda_visible_devices and deployment.trainer.nproc_per_node together and rescale the watermarks (next section).
The H200 figures in this table were measured with DeepGEMM disabled on the capture servers, so every FP8 GEMM ran SGLang's untuned Triton block-FP8 kernel. A stock sglang[all]==0.5.18 install selects DeepGEMM: the 4+4 recipe then measures 13.96 samples/s at 3.67 captured prompts/s per server on 8x H200 (next section).
One H200 node, tuned: five servers + DP3
specforge train -c \
examples/configs/online/disaggregated/managed-local/qwen3.8-27b-dflash2-h200-5server-dp3-disaggregated.yamlThis recipe moves the whole single-node H200 stack to the kernels and transport of this revision and re-balances the split for them. Measured on one 8x H200 node, Qwen/Qwen3.8-27B-FP8, 50k conversations of a Qwen3.8-27B regeneration corpus (mean 4.0k tokens, 26% at 8192), steady-state steps after warm-up, stock SGLang 0.5.18 with DeepGEMM:
| Configuration | Split | samples/s | Limited by |
|---|---|---|---|
| 4+4 recipe on the previous revision (triton attention, TCP) | 4+4 | 13.96 | trainer compute (2.03 s/step) and producer drains every 4,096 prompts |
| same, re-split | 5+3 | 12.33 | three trainer ranks |
same, model.use_liger_kernel: true | 4+4 / 5+3 | 14.00 / 13.10 | trainer ranks and producer drains |
| this revision, FA3 + FlashInfer GDN servers, TCP | 4+4 | 18.06 | trainer compute + fetch wait |
+ training.dflash_teacher_metrics: false | 4+4 | 18.70 | trainer TCP receive path |
| + RDMA loopback | 4+4 | 18.60 | capture servers (trainers wait 0.42 s/step) |
| + Liger, re-split | 5+3 | 22.34 | balanced, no data wait |
| + accumulation 6 (this recipe, 36 samples/step) | 5+3 | 22.65 | balanced |
What moved each side:
- Capture servers, 3.67 -> 4.40 prompts/s per H200: FA3 for the 16 full-attention layers (+14%; the Triton extend kernel was 22.5% of server GPU time) and the FlashInfer SM90 GDN prefill for the 48 linear-attention layers (+8%). DeepGEMM FP8 GEMM is now 71% of server time. The FlashInfer GDN path requires
--cuda-graph-backend-decode disabledon capture servers. - Trainer rank, 3.96 -> 6.66 samples/s per H200 in isolation: the fused DFlash2 unary head (the 248k-vocabulary objective ran three GEMMs and about 330 GB of fp32 traffic per micro-batch; now 3.4x faster), the fused grouped convolution, one host sync per optimizer step instead of one per micro-step, and Liger RMSNorm/SwiGLU.
- RDMA loopback takes the per-byte CPU copies out of the SGLang and trainer processes: at 4+4 the trainer step fell from 1.42 s to 1.22 s. The capture servers publish from GPU memory on RDMA.
- The continuous producer feed removes the fleet-wide drain at every 4,096-prompt boundary (about 29 s without capture on this node's CPUs), and the durable ack runs off the training thread.
With both sides faster, five servers are needed to feed three trainer ranks. The training math is unchanged apart from floating-point rounding: loss and accuracy over 300 optimizer steps stay inside the spread of the previous revision under a pure summation-order change, and FA3/FlashInfer change the captured states by the same amount any kernel swap does on this FP8 target (last hidden state cosine 0.9955; the previous revision moves a prompt's states by up to 17% when it is batched with other prompts). Set SGLANG_JIT_DEEPGEMM_PRECOMPILE=0 in the launching shell to skip DeepGEMM's warm-up sweep on every server start (20-25 min to under a minute once its JIT cache is populated).
Two nodes: eight servers + DP8
export DISAGG_STORE_ID=qwen3.8-27b-dflash2-attempt-001
export DISAGG_RUN_ROOT=/shared/specforge/$DISAGG_STORE_ID
# The launcher supplies RCLI_NODE_RANK, RCLI_NUM_NODES, and RCLI_HEAD_IP.
rcli exec --per-node <job> \
'bash examples/disagg/run_qwen3.8_27b_dflash2_disagg_2node.sh'Rank 0 applies the capture patch, starts mooncake_master with a 10 s read lease, starts eight TP1 capture servers on GPUs 0-7 (ports 30000-30007) with a 64 GiB Store segment each, waits for every /health, then runs the CPU producer. Rank 1 waits for readiness and runs the eight-rank consumer. Both nodes must see the fresh DISAGG_RUN_ROOT; feature tensors travel only through Mooncake. DRY_RUN=1 prints every command instead of running it. Override SERVER_COUNT, SERVER_GPUS, TRAINER_GPUS and TRAINER_NPROC for another allocation, MOONCAKE_PROTOCOL=rdma plus MOONCAKE_RDMA_DEVICES=<hca> for an InfiniBand fabric, and TARGET_MODEL_PATH=Qwen/Qwen3.8-27B-FP8 for Hopper.
Without the wrapper, start the services on the capture node with the settings the wrapper uses. Replace CAPTURE_IP with the address the trainer node can reach, and pass the Mooncake variables to each server process itself (a tmux/nohup session inherits its server's environment, not your shell's):
export MOONCAKE_LOCAL_HOSTNAME="$CAPTURE_IP"
export MC_TCP_BIND_ADDRESS="$CAPTURE_IP"
mooncake_master \
--enable_http_metadata_server=true \
--http_metadata_server_host=0.0.0.0 \
--rpc_port=35551 \
--http_metadata_server_port=35880 \
--metrics_port=35903 \
--default_kv_lease_ttl=10000
for i in 0 1 2 3 4 5 6 7; do
env CUDA_VISIBLE_DEVICES=$i \
MOONCAKE_MASTER_SERVER_ADDR="$CAPTURE_IP:35551" \
MOONCAKE_METADATA_SERVER="http://$CAPTURE_IP:35880/metadata" \
MOONCAKE_LOCAL_HOSTNAME="$CAPTURE_IP" \
MC_TCP_BIND_ADDRESS="$CAPTURE_IP" \
MOONCAKE_PROTOCOL=tcp \
MOONCAKE_GLOBAL_SEGMENT_SIZE=68719476736 \
MOONCAKE_LOCAL_BUFFER_SIZE=4294967296 \
python -m sglang.launch_server \
--model-path Qwen/Qwen3.8-27B --dtype bfloat16 \
--trust-remote-code --skip-tokenizer-init \
--tp-size 1 --host 0.0.0.0 --port $((30000 + i)) \
--enable-spec-capture --spec-capture-method dflash \
--spec-capture-aux-layer-ids 5 19 33 47 61 \
--chunked-prefill-size -1 --disable-radix-cache \
--context-length 8199 --attention-backend triton \
--mem-fraction-static 0.85 &
doneThen run the two roles from the external recipe, overriding the placeholder capture-node endpoints (--role both on one host also works when the producer should share the trainer node):
export MOONCAKE_LOCAL_HOSTNAME="$THIS_NODE_IP"
export MC_TCP_BIND_ADDRESS="$THIS_NODE_IP"
unset RANK LOCAL_RANK WORLD_SIZE MASTER_ADDR MASTER_PORT NODE_RANK
SERVER_URLS=$(for port in $(seq 30000 30007); do
printf '"http://%s:%s",' "$CAPTURE_IP" "$port"
done)
specforge train \
-c examples/configs/online/disaggregated/external/qwen3.8-27b-dflash2-disaggregated.yaml \
--role consumer \
"deployment.disaggregated.server_urls=[${SERVER_URLS%,}]" \
"deployment.disaggregated.mooncake_metadata_server=http://$CAPTURE_IP:35880/metadata" \
"deployment.disaggregated.mooncake_master_server_addr=$CAPTURE_IP:35551"Measured two-node throughput:
| 16 GPUs | 8 servers + DP8, TCP, host receive | 8 servers + DP8, RDMA, device receive (#840/#841) | 10 servers + DP6, device receive | colocated on two nodes (#783), for reference |
|---|---|---|---|---|
| 2x8 H200, FP8 target | 17.98 | 17.81 | 20.41 | 16.04 |
| 2x8 B300, NVFP4 target | 18.04 | not measured | not measured | 34.5 |
On H200 eight servers cap the run at 17.9 captured prompts/s whichever way the bytes travel, so the recipe's TCP default and the RDMA variant are equivalent there; two extra servers on the trainer node (10 + DP6, watermarks 512/256) lift it to 20.4. On B300 the same 8+8 layout reaches 18.0 with the host receive path and stays trainer-bound: draft steps take 3.0 s per eight samples against 1.7 s in colocated mode because the pageable host-to-device copies share the training stream; the device receive path removes that cost on a single B300 node (4+4: 18.3) but was not measured across B300 nodes.
Flow control and Mooncake settings
| Setting | Value | Why |
|---|---|---|
default_kv_lease_ttl_ms (managed) / mooncake_master --default_kv_lease_ttl | 10000 | The 500 ms schema default expires while several ranks fetch 277-500 MB samples; Mooncake then returns -707 LEASE_EXPIRED, which the loader retries with backoff. 3 s still expired under eight-rank cross-node contention; 10 s produced no retries in 150-step runs. |
runtime.in_flight_high_watermark / low_watermark | 192/96 (DP4), 512/256 (DP8) | A ref stays committed-but-unacknowledged from capture until the optimizer boundary that consumed it, so the loader's ranks x dataloader_num_workers x batch_size leased refs sit on top of one step quantum. At 3x the quantum (96/48) the producer paused 64% of the time and the 4+4 B300 node made 7.6 samples/s; 6x (192/96) gave 11.6 and 12x added only backlog latency. On 8+8 nodes 192/96 gave 14.0 and 512/256 gave 18.0. |
data.dataloader_num_workers | 8 | Prefetch depth of the loader introduced by #797; eight concurrent fetches per rank hide the per-fetch latency across nodes. |
runtime.producer_lease / producer_concurrency | 4 / 4 | Sixteen prompts in flight per capture server. A server limited to two prompts in flight captured 4.9 prompts/s instead of 7.2 on B300. |
deployment.disaggregated.client_buffer_size | 2 GiB | The pinned staging buffer of each SpecForge Mooncake client. Neutral for the host receive path on main (the 256 MiB default performs the same) and required for the device receive path of #840; the measured runs used 2 GiB. |
global_segment_size_bytes / MOONCAKE_GLOBAL_SEGMENT_SIZE | 64 GiB per capture server | Store capacity is contributed by every server; 4 x 64 GiB or 8 x 64 GiB covers 192 or 512 in-flight samples at 500 MB with margin. The Store rejects writes once the segments are full. |
local_buffer_size_bytes / MOONCAKE_LOCAL_BUFFER_SIZE | 4 GiB | Per-server client buffer for publishing 500 MB samples from sixteen concurrent requests. |
MC_TCP_BIND_ADDRESS | the node's routable address | Mooncake's transfer engine otherwise advertises the first active interface; on multi-interface or containerized hosts that is a bridge address the peer cannot reach and remote reads fail with -800 TRANSFER_FAIL although captures succeed. The two-node wrapper exports it alongside MOONCAKE_LOCAL_HOSTNAME on both nodes. |
training.batch_size x accumulation_steps | 2 x 4 | 32 (DP4) or 64 (DP8) samples per optimizer step is the measured configuration. Full-convergence runs of this draft used 512 samples per step at the same learning rate; raise accumulation_steps and the watermarks together. |
Servers and trainers report where the time goes: a producer that logs pauses at the high watermark is outrunning the trainers (add trainer ranks or raise the watermark if the loader is idle), a trainer whose data-wait share stays high while the producer never pauses is starved by the servers (add servers).
Capture-server knobs
Both recipes keep SGLang's kernel and scheduling defaults apart from the settings above. The managed-local recipe passes these fields to every capture server; --plan prints each resulting server command and environment:
| Setting | SGLang flag or effect |
|---|---|
model.sglang_max_prefill_tokens | --max-prefill-tokens |
model.sglang_linear_attn_prefill_backend | --linear-attn-prefill-backend (the target's GDN layers) |
model.sglang_fp8_gemm_backend | --fp8-gemm-backend (the FP8 checkpoint) |
model.sglang_disable_cuda_graph | --disable-cuda-graph |
model.sglang_max_running_requests | --max-running-requests |
model.sglang_extra_args | any other flag, one list item per flag or value, e.g. ["--prefill-max-requests", "4", "--enable-metrics"] |
capture_servers[].extra_args | flags for one server, after the global list |
capture_servers[].env | environment for one server, e.g. SGLANG_SPEC_CAPTURE_TIMING: "1" |
The schema rejects passthrough flags that SpecForge renders itself (model path, port, TP size, dtype, context length, chunked prefill and the capture flags) or that another model.sglang_* field already sets, abbreviations of those, --api-key, SGLang DP options and --config, and environment keys it owns (MOONCAKE_*, DISAGG_*, device visibility). Use full flag spellings: abbreviated duplicates within or across passthrough lists are rejected, as are abbreviations of secret-valued flags such as --admin-api-key. Distinct complete flags such as --enable-metrics and --enable-metrics-for-all-schedulers remain valid together. On the two-node wrapper, SERVER_EXTRA_ARGS_APPEND adds flags after the recipe's server defaults, and exported variables reach every server.
Known limitations
- Resolved: acknowledgement stall on partial removals.#832 treats Mooncake
-704 OBJECT_NOT_FOUNDas a completed removal, so the bounded drain no longer fires on partly leased samples. - Resolved: host receive path.#840, #841 and #881 made pinned receives the default and publish captures from GPU memory on RDMA; the tuned H200 recipe uses RDMA loopback.
- Colocated online capture (#783) is the other topology for this draft. It is faster than any disaggregated split on B300 (18.9 on eight GPUs) and slower than the tuned split on H200 (8.61 versus 10.65), because on H200 the target forward dominates the step and disaggregation takes it off the trainer rank.
Validation
Recipes and launcher were exercised as follows:
- 150-300 optimizer-step runs per topology and split on 8x/2x8 B300 and 8x/2x8 H200, from which every number above is taken; both recipes reproduce the flow-control tuple of those runs.
specforge train -c <recipe> --planfor both recipes on the source revision of this document;tests/test_config(recipe topology, world-size and draft wiring checks) pass.- The tuned H200 recipe: 110-200 optimizer-step runs per row of its table on one 8x H200 node, a 40-step smoke of the recipe file as checked in, and offline parity (identical init, data and seed) plus a 300-step training comparison for the trainer kernels.
DRY_RUN=1of the two-node wrapper for both node ranks, including the eight-server URL list handed to the producer.- No full-convergence run was made in disaggregated mode. The optimizer settings (constant 5e-4,
loss_decay_gamma7, hard-target CE with the strict top-k selector objective) are those of one-epoch colocated runs of the same draft and corpus, in which the strict top-k selector gave the best serving acceptance length among the objective variants tried.
Bring a new deployment up in this order: config plan, capture patch, server health, one captured sample's tensor shapes and dtypes, one optimizer step with training.max_steps=1 and data.max_prompts at one step quantum, then the full run. The low watermark must stay at or above the global optimizer-step quantum, or the producer can pause while the consumers wait for an incomplete step.