Written before the run. Criteria are fixed here; numbers get filled in below.
[email protected], campaign /home/deploy/reinstator-experiments/20260919-phase0.deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B (documented as SFT of Qwen2.5-Math-1.5B on DeepSeek-R1 outputs).Qwen/Qwen2.5-Math-1.5B (the student's base; this is the base-model reference variant the paper flags as weaker than a same-lineage instruct checkpoint).Qwen/Qwen2.5-Math-1.5B-Instruct (same base, fine-tuned by Qwen, not documented as R1-distilled). Also scored against R.provenancekit compare S R scores high (expect above 0.9: same architecture, weights inherited). provenancekit scan S ranks the Qwen2.5 family first and does not surface DeepSeek-R1 as a match. This documents that the kit sees the base and not the teacher.Failure of 2 with success of 1 would still be informative: it would mean the base-model reference is too weak at 1.5B and H1 becomes the first priority.
scan by default runs metadata-only: every match reports identity_score None and pipeline_score = mfi_score. Deep weight fingerprints are a separate download-deepsignals-fingerprint step, not run here.isolated.sh; future teacher generation runs on the i9 directly.results/tests/interim/.deepseek/deepseek-r1 outputs we received (all from provider Novita) begin "The problem ..." or "I need ..." in 0/140 cases with "Okay", which is the R1-0528 style. So the pool's "R1" is probably not the checkpoint that produced the student's training data, and Qwen3's traces are the closest stylistic relative in the pool. This is exactly the entangled-lineage / wrong-checkpoint hazard from the paper, hit on the first real run. Follow-up: sample R1 through DeepSeek's first-party endpoint and R1-0528 explicitly, compare openers; if first-party R1 says "Okay, so", regenerate the r1 set with that provider pinned.deepseek/deepseek-r1 (Novita, fp8); no first-party or alternative provider is available, so the served checkpoint cannot be cross-checked through the API. Instead, a fifth teacher set r1_orig was built from simplescaling/s1K-1.1, whose deepseek_thinking_trajectory / deepseek_attempt columns are DeepSeek-R1 outputs from early 2025 on the same s1K questions: 200/200 prompts matched by question hash, 200/200 traces open with "Okay/Alright". Zero API cost. Scored at 512 tokens into nll_*_512.r1orig.jsonl.deepseek/deepseek-r1-0528 sampled on 3 prompts (DeepInfra, SiliconFlow) opens with "I have two polynomials..." / "The problem asks..." — the same register as the Novita-served deepseek/deepseek-r1 set (0/140 "Okay") and unlike the s1K-1.1 original-R1 traces (200/200 "Okay/Alright"). Conclusion: the API "R1" candidate behaves like the May-2025 checkpoint; the student's training teacher is represented only by the r1_orig set. A first-party DeepSeek provider request returns 404 (no endpoint).${tag} loop through isolated.sh; systemd expands ${...} in transient unit command lines and set it to empty, so all three models appended to one file nll__512.r1orig.jsonl for 20 minutes while the unit looked healthy. Caught on the user's "is this actually producing data" check. Discarded the mixed file; relaunched via score_extra.sh (script file, no inline variables). Rule added to isolated.sh.Files: results/tests/test_S_R_512_5t.json, test_C_R_512_5t.json, test_S_Cref_512_5t.json (mirrored under results/tests/).
| Teacher | S mean f (base ref) | S top-1 wins | S mean f (sibling-instruct ref) | C mean f | C wins |
|---|---|---|---|---|---|
| qwen3_32b | +0.458 | 96/187 | +0.680 | -0.222 | 7 |
| r1_orig (s1K-1.1 traces, Jan/Feb 2025 R1) | +0.457 | 91/187 | +0.677 | -0.220 | 12 |
| r1 (API, Novita, 0528-style) | +0.244 | 0/187 | +0.401 | -0.157 | 35 |
| gpt_oss_120b | +0.197 | 0/187 | +0.463 | -0.266 | 3 |
| llama33_70b | -0.146 | 0/187 | -0.028 | -0.118 | 130 |
Criteria: (1) met; (2) met against every candidate except Qwen3-32B, which was not in the paper's pool and ties; (3) met; (4) met.
Files: results/gpu/tests/test_S_R_2048_float32.json, test_C_R_2048_float32.json; scores in results/gpu/nll/. RunPod H100 SXM in US-CA-2, 15.5 min wall including first-time bootstrap, ~$0.90; each model scored 999 probe-teacher pairs in 2.8 min (the same work is ~4.5 h per model on the i9).
| Teacher | S mean f | S wins | C mean f | C wins |
|---|---|---|---|---|
| r1_orig | +0.357 | 92/199 | -0.211 | 3 |
| qwen3_32b | +0.355 | 107/199 | -0.199 | 10 |
| r1 (API) | +0.187 | 0 | -0.154 | 26 |
| gpt_oss_120b | +0.130 | 0 | -0.246 | 2 |
| llama33_70b | -0.143 | 0 | -0.109 | 158 |
(512-token pass on the i9; 2048-token pass on the H100 above)
| Check | Value | Criterion met |
|---|---|---|
| kit compare S vs R | pipeline 0.778 "High-Confidence Match"; identity 0.778; EAS 0.999, LEP 0.998, END 0.993, NLF 0.869, WVC 0.003; MFI 0.919 soft match | partly: match is high-confidence but below the 0.9 I guessed. WVC collapsing to ~0 (vs 0.995 for C vs R) says the distill SFT moved the raw weights far from the base while embedding geometry stayed. |
| kit compare S vs C | pipeline 0.669 "Weak Match" (END 0.429, WVC 0.003) | n/a, sibling check |
| kit compare C vs R | pipeline 1.000 "Confirmed Match", WVC 0.995 | control behaves as a normal fine-tune |
| kit scan S top match | itself (DB has an entry), then Nemotron-Research-Reasoning-Qwen-1.5B, OpenMath-Nemotron-1.5B, OpenReasoning-Nemotron-1.5B, Qwen2-1.5B. DeepSeek-R1 (the teacher) is not in the database and cannot appear. | yes: kit sees publisher/base lineage, not the teacher |
| S: identified teacher, top-1 rate, binomial p_bonf | final 512, 198 probes, K=5: qwen3_32b 102/198 (p 6.8e-23) and r1_orig 96/198 (p 2.9e-19); api-r1 0, gpt_oss 0, llama 0 | yes for the true teacher vs every candidate except the Qwen3 near-clone (tie) |
| S: margin, governing test p_bonf | top1 - top2 = 0.003, p_bonf 1.0 (qwen3 vs r1_orig indistinguishable) | no separation between the top two |
| C: best teacher, margin, p_bonf | llama33_70b 139/198, margin 0.042, p_bonf 8e-15; all reasoning teachers negative; r1_orig 12 wins | yes (no false reasoning-teacher attribution; the Llama preference is a plain-answer style effect, not distillation) |
| raw-likelihood baseline on S | picks qwen3_32b at p_bonf 0.055 vs reference-based p ~1e-23 | yes, reference term is decisive |
| wall time, threads, truncation | 512 tokens; ~3.5 h wall on the i9 for 999 probe-teacher pairs x 3 models at 8 threads, one process per model | recorded |
| API cost | r1 $3.40 (two runs), gpt_oss $0.15, llama $0.07, qwen3_32b ~$0.55, smokes ~$0.10; r1_orig free | ~$4.30 total |