PUBLIC SANITIZED RELEASE
Kakeya OProver Lab
Lean proof training from SFT90 to repair SFT, with matched strict evaluation and explicit non-claims.
Overview
Training Journey
SFT90 — establish a small-data baseline.
SFT1000 — matched validity baseline.
DPO — preference optimization reduced strict validity.
Repair SFT — compiler-feedback repair restored quality gates.
Production alignment — 504/504 converted; 12/12 canary passed.
Strict Evaluation
Paired comparisons use exact McNemar tests; intervals are exact 95% confidence intervals on rates or discordant-pair win probability.
Training Metrics
Loss curves and complete sanitized JSONL metrics are stored in each model repository. Stage selector above links to the source assets.
Failure Taxonomy
Repair SFT targeted executable-proof failures while rehearsing clean Mathlib examples.
Production Qualification
This is a proof-advisor qualification, not authorization for autonomous deployment.
Jensen Result
No candidate proof bodies or production prompts are published.
Reproducibility
Immutable revisions, adapter hashes, dataset artifact hashes, protocol roots, audit digests, and public/private boundaries are in the bundled release manifest.