One question per study, frozen when it is answered.
Each study has a permanent page. Once its findings are final and a claim audit has resolved every error it found, the study is frozen: its findings stop changing, and its code, models and datasets are tagged at that point. A number later found wrong is corrected in the repository’s errata. New experiments go into the next study.
- Study001
Frozen 13 Sept 2026 · evidence at study-001
We improved the metric. Then the model stopped answering.
Tool-policy post-training on Qwen3.5-2B: SFT, DPO, and the capability regression the gate missed
Confirmatory call F1 rose from 0.6264 to 0.7470 with SFT. The promoted checkpoint then refused every bare GSM8K question, a regression the promotion gate could not see because its held-out set had no questions that simply asked for an answer.
- Study002
In progress · no results yet
Relabelling, on-policy distillation and speculative decoding
Corpus corrections, M2 on-policy distillation, and speculative decoding on the Study 001 lineage
Everything after the Study 001 freeze. No results yet: each experiment is pre-registered before it runs and reported against the frozen Study 001 checkpoints, not in place of them.