Phase 0 result

Written before the run. Criteria are fixed here; numbers get filled in below.

Setup

Pre-registered criteria

Failure of 2 with success of 1 would still be informative: it would mean the base-model reference is too weak at 1.5B and H1 becomes the first priority.

Run log

Five-teacher result, 512 tokens, 187 complete probes (2026-09-19)

Files: results/tests/test_S_R_512_5t.json, test_C_R_512_5t.json, test_S_Cref_512_5t.json (mirrored under results/tests/).

TeacherS mean f (base ref)S top-1 winsS mean f (sibling-instruct ref)C mean fC wins
qwen3_32b+0.45896/187+0.680-0.2227
r1_orig (s1K-1.1 traces, Jan/Feb 2025 R1)+0.45791/187+0.677-0.22012
r1 (API, Novita, 0528-style)+0.2440/187+0.401-0.15735
gpt_oss_120b+0.1970/187+0.463-0.2663
llama33_70b-0.1460/187-0.028-0.118130

Criteria: (1) met; (2) met against every candidate except Qwen3-32B, which was not in the paper's pool and ties; (3) met; (4) met.

2048-token result, H100 burst, fp32, 199 probes (2026-09-19)

Files: results/gpu/tests/test_S_R_2048_float32.json, test_C_R_2048_float32.json; scores in results/gpu/nll/. RunPod H100 SXM in US-CA-2, 15.5 min wall including first-time bootstrap, ~$0.90; each model scored 999 probe-teacher pairs in 2.8 min (the same work is ~4.5 h per model on the i9).

TeacherS mean fS winsC mean fC wins
r1_orig+0.35792/199-0.2113
qwen3_32b+0.355107/199-0.19910
r1 (API)+0.1870-0.15426
gpt_oss_120b+0.1300-0.2462
llama33_70b-0.1430-0.109158

Numbers

(512-token pass on the i9; 2048-token pass on the H100 above)

CheckValueCriterion met
kit compare S vs Rpipeline 0.778 "High-Confidence Match"; identity 0.778; EAS 0.999, LEP 0.998, END 0.993, NLF 0.869, WVC 0.003; MFI 0.919 soft matchpartly: match is high-confidence but below the 0.9 I guessed. WVC collapsing to ~0 (vs 0.995 for C vs R) says the distill SFT moved the raw weights far from the base while embedding geometry stayed.
kit compare S vs Cpipeline 0.669 "Weak Match" (END 0.429, WVC 0.003)n/a, sibling check
kit compare C vs Rpipeline 1.000 "Confirmed Match", WVC 0.995control behaves as a normal fine-tune
kit scan S top matchitself (DB has an entry), then Nemotron-Research-Reasoning-Qwen-1.5B, OpenMath-Nemotron-1.5B, OpenReasoning-Nemotron-1.5B, Qwen2-1.5B. DeepSeek-R1 (the teacher) is not in the database and cannot appear.yes: kit sees publisher/base lineage, not the teacher
S: identified teacher, top-1 rate, binomial p_bonffinal 512, 198 probes, K=5: qwen3_32b 102/198 (p 6.8e-23) and r1_orig 96/198 (p 2.9e-19); api-r1 0, gpt_oss 0, llama 0yes for the true teacher vs every candidate except the Qwen3 near-clone (tie)
S: margin, governing test p_bonftop1 - top2 = 0.003, p_bonf 1.0 (qwen3 vs r1_orig indistinguishable)no separation between the top two
C: best teacher, margin, p_bonfllama33_70b 139/198, margin 0.042, p_bonf 8e-15; all reasoning teachers negative; r1_orig 12 winsyes (no false reasoning-teacher attribution; the Llama preference is a plain-answer style effect, not distillation)
raw-likelihood baseline on Spicks qwen3_32b at p_bonf 0.055 vs reference-based p ~1e-23yes, reference term is decisive
wall time, threads, truncation512 tokens; ~3.5 h wall on the i9 for 999 probe-teacher pairs x 3 models at 8 threads, one process per modelrecorded
API costr1 $3.40 (two runs), gpt_oss $0.15, llama $0.07, qwen3_32b ~$0.55, smokes ~$0.10; r1_orig free~$4.30 total