R2: false-positive sweep

Written before the run (2026-09-20). Chunk 2 of the roadmap.

Question

How often does each signal accuse an innocent model, and what absolute thresholds keep that rate near zero? Phase 1 had three negatives (the control instruct model, a Gemini-trained student, a human-solution student). This sweep adds negatives of the kinds the plan lists as hard, and two positives with a different base, to separate "no R1 signal" from "R1 signal under a mismatched reference".

Suspects

Negatives (no exposure to R1 or Qwen3 outputs):

Positives with a different base (R1 traces, s1K, 800 x 3 epochs, same recipe):

References: each suspect is scored against its own base where available, and against Qwen2.5-1.5B (R3) as the mismatched reference.

Pre-registered criteria

Numbers

(to fill)