• RECORDED EVALUATION
  • OFFLINE REPLAY

Know the frontier.
Then forge past it.

Scaling free rule labels from 1,450 to 20,000 raised frozen-eval task success from 66.35% to 99.05%: +32.70 percentage points, paired 95% CI [30.60, 34.50], in one training seed (seed 0), using 15.236 measured RTX 4090 GPU-hours ($4.571).

99.1%TASK SUCCESS
+32.7 ppGAIN VS R1
15.24RTX 4090 GPU-HOURS
$4.571MEASURED TRAINING COST

The release model is the rule-label scaling ablation, not GRPO. GRPO's paired 95% CI includes zero across both completed seeds; a third seed aborted on the zero-reward-variance guard.

  • FULL INSTRUMENT
  • RECORDED SERVING RUNS

The same instrument,
full size, zero clicks.

RECORDED EVALUATION · OFFLINE REPLAY

1.06s320.6 tok/s · 20 reqs · 95% success · RTX 4090phase4_spec_decode_r1b_bf16_native_mtp · sha256:7878b55f6f…

COMPLAINT TRANSCRIPT — NOT RECORDED. release.json ships aggregate serving metrics, not per-request text.

curl -s https://xiangguozhang.com/case-studies/frontier-forge/release.json |   jq '.serving.serving_at_4_qps[] | select(.run_id == "phase4_spec_decode_r1b_bf16_native_mtp")'

RELEASE.JSON · CLAIM REGISTRY · SHA-256

Every number on this page
opens the command that made it.

Blank command cells mean the receipt recorded no shell command; config paths appear only where present.

ClaimNumbern · CIGeneration commandSHA-256
PositiveFree-label SFT reached the selected task-success result99.05%n=2,000 · 95% CI 98.60%–99.45%make reproduce-headline673138d9888b
NegativeAPI distillation lost to the smaller free-label SFT run52.15% · −14.20 pp vs R195% CI 50.00%–54.40%make reproduce-headline673138d9888b
NegativeGRPO's paired interval includes zero+0.25 ppn=2,000 · 95% CI −0.10 pp–+0.65 pp673138d9888b
NegativeThe alternative training backend failed the agreement gate−3.85 ppn=2,000 · paired 95% CI −4.85 pp–−2.75 ppconfig: configs/r1_sft_rule.yaml06e3f0c235f6
PositiveGPTQ-int4 recorded the lowest 4 QPS p950.963 s p95n=20 · task-success Wilson 95% CI 76.39%–99.11%c99b42cf0e06
BoundaryNative MTP has a measured win/lose boundary0.25 QPS lose; 0.50-4.00 QPS win7878b55f6fe6
PositiveAt 3× overload the gateway rejected excess work with zero upstream 5xx215 HTTP 429 · 0.00% upstream 5xxn=707 scheduled per side9e547b3418ce
NegativeBare vLLM crashed in the matched 5× cell651 transport errors · gateway 687 HTTP 429n=1,177 scheduled per sidemake reproduce-headline9e547b3418ce
PositiveThe CPU gateway scaled out and returned to one replica1→3→1 replicasn=3,799 recorded HTTP 4299e547b3418ce
BoundaryGPU cold start is a batch boundary, not an interactive SLOp50 124.6 s · p95 127.1 sn=109e547b3418ce
  • 99.05%Positive

    Free-label SFT reached the selected task-success result

    n=2,000 · 95% CI 98.60%–99.45%

    make reproduce-headline

    SHA-256 673138d9888b

  • 52.15% · −14.20 pp vs R1Negative

    API distillation lost to the smaller free-label SFT run

    95% CI 50.00%–54.40%

    make reproduce-headline

    SHA-256 673138d9888b

  • +0.25 ppNegative

    GRPO's paired interval includes zero

    n=2,000 · 95% CI −0.10 pp–+0.65 pp

    SHA-256 673138d9888b

  • −3.85 ppNegative

    The alternative training backend failed the agreement gate

    n=2,000 · paired 95% CI −4.85 pp–−2.75 pp

    config: configs/r1_sft_rule.yaml

    SHA-256 06e3f0c235f6

  • 0.963 s p95Positive

    GPTQ-int4 recorded the lowest 4 QPS p95

    n=20 · task-success Wilson 95% CI 76.39%–99.11%

    SHA-256 c99b42cf0e06

  • 0.25 QPS lose; 0.50-4.00 QPS winBoundary

    Native MTP has a measured win/lose boundary

    SHA-256 7878b55f6fe6

  • 215 HTTP 429 · 0.00% upstream 5xxPositive

    At 3× overload the gateway rejected excess work with zero upstream 5xx

    n=707 scheduled per side

    SHA-256 9e547b3418ce

  • 651 transport errors · gateway 687 HTTP 429Negative

    Bare vLLM crashed in the matched 5× cell

    n=1,177 scheduled per side

    make reproduce-headline

    SHA-256 9e547b3418ce

  • 1→3→1 replicasPositive

    The CPU gateway scaled out and returned to one replica

    n=3,799 recorded HTTP 429

    SHA-256 9e547b3418ce

  • p50 124.6 s · p95 127.1 sBoundary

    GPU cold start is a batch boundary, not an interactive SLO

    n=10

    SHA-256 9e547b3418ce

1,450 → 20,000 RULE LABELS

Free labels moved
the frontier.

Scaling free rule labels from 1,450 to 20,000 raised frozen-eval task success from 66.35% to 99.05%: +32.70 percentage points, paired 95% CI [30.60, 34.50], in one training seed (seed 0), using 15.236 measured RTX 4090 GPU-hours ($4.571).

  1. R0 baseComplete0.00%

    95% CI 0.00%–0.00% · 0.98 GPU-h · $0.29

  2. R1 rule SFT (1,450)Complete66.35%

    95% CI 64.20%–68.40% · 3.48 GPU-h · $1.04

  3. R1b rule SFT (20,000)Release-selected99.05%

    95% CI 98.60%–99.45% · 15.24 GPU-h · $4.57

  4. R2 distilled SFTComplete — negative52.15%

    95% CI 50.00%–54.40% · 3.61 GPU-h · $1.08

  5. R3 DPOComplete55.95%

    95% CI 53.75%–58.25% · 1.93 GPU-h · $0.58

  6. R4 v2 GRPO seed 0Partial only56.20%

    95% CI 54.00%–58.50% · 1.86 GPU-h · $0.56

  7. R4 v2 GRPO seed 1Partial only56.20%

    95% CI 53.95%–58.50% · 1.66 GPU-h · $0.50

NEGATIVE RESULT

Distillation lost 14.2 pp to free rule labels: 52.15% versus the smaller R1 run's 66.35%.

NEGATIVE RESULT

GRPO's paired 95% confidence interval includes zero; one additional seed was stopped by the unchanged zero-reward-variance guard.

0.25–4.00 QPS · NATIVE MTP

Serving is a boundary,
not a badge.

At 4 QPS, three precisions were served and measured end to end. Native MTP's win/lose boundary: 0.25 QPS lose; 0.50-4.00 QPS win.

1.42sE2E P50
1.69sE2E P95
0.17sTTFT P50
307.0OUTPUT TOK/S
95%TASK SUCCESS
$0.0211COST / 1K TASKS
22,829 MiBVRAM PEAK
n=20REQUESTS
0.25 QPSLOSE+0.051s p95
0.5 QPSWIN-0.229s p95
1 QPSWIN-0.227s p95
2 QPSWIN-0.321s p95
4 QPSWIN-0.370s p95

SAME-BOX A10 · SUSTAINED OVERLOAD

Reject the work
before it becomes a crash.

Fixed-seed Poisson arrivals, same-box NVIDIA A10, every cell ran at least 120 seconds; each cell's queue was verified saturated.

MultiplierOffered QPSRequestsHTTP 429Gateway upstream 5xxBare vLLM upstream 5xxGate
2×4510190.0%0.0%PASS
3×67072150.0%0.0%PASS
5×101,1776870.0%3.1%PASS

At the highest recorded load (5×, 10 QPS): bare vLLM logged 651 transport errors while the gateway returned 687 bounded HTTP 429 rejects and zero upstream 5xx.

RTX 4090 · TIME-SLICING · TP · DDP / FSDP

Two replicas helped. TP=2 did not.
FSDP traded speed for memory.

Phase 7.3 time-sliced one RTX 4090 between vLLM replicas and added an optional two-GPU tensor-parallel point. Phase 8 trained the same full-parameter SFT three ways on one 2×RTX 4090 machine.

Two replicas vs one · one RTX 4090

1.748× / 1.795×SUCCESS THROUGHPUT · QPS 4 / 8
+37.40 pp / +19.34 ppSUCCESS RATE · QPS 4 / 8
0.408× / 0.420×TTFT P95 · QPS 4 / 8
85.37s1→2 READY P50

Against one replica on the same seed-731 request plan (180 s per cell, QPS 1/2/4/8), two time-sliced vLLM replicas raised successful-task throughput and cut TTFT p95 at QPS 4 and 8. QPS 8 still returned many HTTP 429s, and the ratios come from one descriptive run without repeat-trial intervals.

LIMITATION

Isolation cuts both ways. While a neighbor replica took 8 QPS directly, the observed replica's all-request end-to-end p95 rose only 1.65% (4.472s → 4.546s). Yet 1,353 of the attacker's 1,505 requests timed out after HTTP 200 headers arrived, and only 151 passed task verification. Time-slicing gives no hard isolation.

STAGED PASS

The scaling gate passed in stages: 10 cycles of 0→1→2→1→0 (40 transitions) came from an independent re-run, while the original attempt stays sealed as failed. With node, image, and model caches already in place, 1→2 reached Ready at p50 85.37s.

Phase 7.3 measured summaryLucisZhang/frontier-forge · 34e857b417bb

Tensor parallelism · one measured point

0.830×SUCCESS THROUGHPUT · TP=2 / TP=1
7.268×TTFT P95 · TP=2 / TP=1
1.205×COST PER 1K SUCCESSFUL TASKS
n=372REQUESTS PER TP SETTING

NEGATIVE RESULT

On one 2×RTX 4090 machine at QPS 2 for 180 s, TP=2 was slower than TP=1 and cost more per successful task. One operating point is not a scalability result, which is why this page presents replicas, not tensor parallelism, as the measured scaling path.

TP single-point reportLucisZhang/frontier-forge · 34e857b417bb

Single GPU, DDP, FSDP · same training, three ways

Full-parameter SFT of Qwen3.5-0.8B-Base (752,393,024 parameters): 1,250 steps on 20,000 training rows, scored on the same frozen 2,000-row evaluation.

  1. Single GPU + accumulationBASELINE1,706.90 tok/s

    PEAK ALLOCATED / GPU 14.26 GiB · hard-AND 98.45% · n=2,000

  2. DDPSCALING EFF. 98.62%3,366.58 tok/s

    PEAK ALLOCATED / GPU 17.07 GiB · hard-AND 98.40% · n=2,000

  3. FSDP FULL_SHARDSCALING EFF. 67.66%2,309.85 tok/s

    PEAK ALLOCATED / GPU 7.58 GiB · hard-AND 98.55% · n=2,000

−55.61%PEAK MEMORY / GPU · FSDP VS DDP
−0.05 pp [−0.30, +0.20]DDP − SINGLE GPU · PAIRED 95% CI
+0.10 pp [−0.10, +0.30]FSDP − SINGLE GPU · PAIRED 95% CI

NOTE

Both paired 95% intervals include zero, so this evaluation detects no quality difference between the three setups. FSDP traded throughput for memory, which is consistent with its extra parameter gathers, but communication time was not profiled separately. The two GPUs connect over PCIe (NODE topology) without NVLink.

Phase 8 distributed reportLucisZhang/frontier-forge · 34e857b417bbMeasured supplementLucisZhang/frontier-forge · 34e857b417bb

MODEL BOUNDARY

The model has a boundary.
It is drawn here.

4 capabilities hold up under measurement; 5 do not — the failing count is not hidden below the passing one.

  • Reaches 99.05% task success on complaint routing (n=2,000 paired, 95% CI 98.60%–99.45%)
  • Holds schema-valid tool calls at 100% under both xgrammar and outlines constraints
  • Gateway upstream 5xx stayed at 0% through 5× sustained overload (bare vLLM did not)
  • Native MTP wins on p95 latency across 0.50-4.00 QPS
  • Constrained decoding's simultaneous mode (no two-pass retry) has 0% task success under both xgrammar and outlines
  • API distillation lost 14.20 pp to the smaller free-label SFT run (52.15% vs 66.35%)
  • GRPO's paired 95% CI includes zero across both completed seeds; a third seed aborted on a zero-reward-variance guard
  • Below 0.5 QPS, native MTP is slower, not faster (0.25 QPS: p95 +0.051s, lose)
  • The alternative training backend (unsloth) failed the agreement gate against the default (trl): −3.85 pp, 95% CI −4.85 pp–−2.75 pp

KNOWN FAILURES

Per-sample misclassified complaints — NOT RECORDED. release.json ships aggregate task_success and paired confidence intervals only, not individual predictions.

The closest real, cited failure mode: neither constrained-decoding backend produced a usable first-pass answer without a second pass.

  • phase4_structured_r1b_bf16_xgrammar — xgrammar: simultaneous task success 0% (24 requests); two-pass recovers to 100% at +0.81s p50 latency.
  • phase4_structured_r1b_bf16_outlines — outlines: simultaneous task success 0% (24 requests); two-pass recovers to 100% at +0.88s p50 latency.

ff-qwen3.5-4b-r1b_sft_rule_20k_s0 · release.json sha256:9e547b3418ce1f63

HOW THIS WAS VERIFIED

Every claim opens
the same command.

What was verified
The release registry connects training outcomes, measured spend and the sustained overload test to recorded runs.
Evidence class
Recorded GPU experiments and a preserved A10 load-test receipt; the browser replays that receipt.
Boundary
The release uses rule-label scaling. GRPO's completed-seed confidence intervals include zero; the third seed aborted. Hardware and workload boundaries remain specific to each run.
Files, hashes and methods
release.jsonLucisZhang/frontier-forge · 06de6e5c1d02
sha256:9e547b3418ce1f633914e8dfe7b818fa6f7e064a891cee8e4c73837e6ce45f4d
Overload replay receiptLucisZhang/frontier-forge · 06de6e5c1d02
sha256:d132ebece698c98ab996b154067679ea1d74b52e9528834ed19a21a132137929
GPU replica and DDP/FSDP projectionCommit-local file · unpublishedSHA-256: 18b6eacc26fb3033c51758f61044953f351c86f41bc7d432cf42639687748644
sha256:18b6eacc26fb3033c51758f61044953f351c86f41bc7d432cf42639687748644
Reproduce the headline
make reproduce-headline
Generated
2026-09-15

The claim table is built from the copied Phase 7 release.json; the local manifest pins its exact SHA-256. The overload replay fetches the preserved Phase 7.1 sustained A10 receipt and does not call a model.

Architecture

I trained and evaluated the ladder, exported the selected checkpoint, served it with vLLM, wrote the C++20 gateway, exercised the k3s runtime, and ran the two-GPU single-GPU/DDP/FSDP training comparison.

  1. Frozen input

    Hash-pinned CFPB splits feed rule labels and API distillation.

  2. Training ladder

    R0 → R1/R1b → R2 → R3 → R4, with paired confidence intervals.

  3. Model exports

    The selected R1b becomes BF16, GPTQ-int4, and MTP-preserved artifacts.

  4. vLLM serving

    Client/server timing records the precision and native-MTP boundaries.

  5. Bounded gateway

    A C++20 token-aware admission layer protects the OpenAI-compatible JSON/SSE path.

Results & negatives

Under 3× overload the gateway sheds load with 429s and zero upstream 5xx; bare vLLM crashed at 5×. Distillation lost 14.2 pp to free rule labels and GRPO's CI includes zero — both runs are kept on the page.

  1. Distillation lost 14.2 pp to free rule labels: 52.15% versus the smaller R1 run's 66.35%.

  2. GRPO's paired 95% confidence interval includes zero; one additional seed was stopped by the unchanged zero-reward-variance guard.

Limitations

  1. The lifted production block applies only to the measured single-node gateway overload contract; it does not establish multi-node or production-grade serving. Two-GPU evidence is limited to Phase 8's single-machine 2×RTX 4090 training comparison and one Phase 7.3 tensor-parallel measurement point.

  2. Only CPU gateway replicas scaled 1→3→1. In Phase 7.2 the GPU deployment moved only between zero and one replica on one physical A10. Phase 7.3 later completed 10 cycles of 0→1→2→1→0 on one RTX 4090 through time-slicing; both replicas share one card, so this is neither hard isolation nor multi-GPU scaling.

  3. The roughly 125-second GPU cold start is suited to batch and development workloads, not an interactive serving SLO.