Analysis plan: Darwin HealthBench Professional regrade Written 3 October 2026, before any grading. Three amendments follow; the first two were recorded before the results they affect were read, and the third records when the deferred graders ran. Public copy for "OpenEvidence's 'perfect 100% on MedQA' covers 660 of the test's 1,273 questions" (https://groundtruth.health/openevidence-darwin-100-percent-medqa/). The wording is the plan's own. Software and account details are left out; where a phrase is replaced, the replacement is in [square brackets]. QUESTION. OpenEvidence reports Darwin at 82.7% on HealthBench Professional, 12.1 points above the next-best model. That figure covers 410 of the 525 tasks (multi-turn excluded), was graded by Claude Opus 4.8 rather than the benchmark's own grader, and is not length-adjusted, although the benchmark paper makes the length-adjusted score its primary score. How does Darwin score under independent graders, with and without the adjustment, and how does that compare with the physicians who wrote responses to the same 410 tasks? INPUTS. Darwin's 410 released answers (OpenEvidence PDF, parsed; answer text plus its reference list, as a user sees it). The benchmark's released dataset (rubrics, physician responses). The grading prompt and item format are byte-identical to Ground Truth's July 2026 physician regrade. GRADERS (none made by Anthropic [, the maker of Claude Opus 4.8, which graded OpenEvidence's own figure, and of Claude Fable 5, Darwin's closest rival on OpenEvidence's chart]). GPT-5.6 Sol and GPT-5.5 [...] (both graded the physician arm in July with the same prompt), Gemini Pro (API), Qwen3.8 27B (open weights, run locally). AMENDMENT 1 (3 October 2026, before any results were read): [GPT-5.6 Sol and GPT-5.5 became unavailable to us], so Sol and GPT-5.5 cannot grade Darwin until 10 Oct. The OpenAI API reaches GPT-5.4 at low reasoning, the benchmark's own default grader, which becomes the primary grader for both arms. Sol and GPT-5.5 are deferred. ARMS. Darwin (410). Physicians on the same 410 tasks: reuse July's Sol and GPT-5.5 grades; grade fresh with Gemini and Qwen. SCORING. Paper section 4.1 as reimplemented in July: binary met per criterion; raw = points met (negatives subtract) / max positive points; adjusted = raw - 2.94e-5 x (chars - 2000); per-example values averaged, mean clipped to [0,1], x100. Primary length = Darwin's answer body without the reference list (generous to OpenEvidence, and it reproduces OpenEvidence's own adjusted figure); sensitivity = with references. Physicians' length = their response. PRIMARY OUTPUTS. Per grader: Darwin raw and adjusted; physicians raw and adjusted on the same 410; Darwin - physician difference with paired bootstrap 95% CI (4,000 resamples). Secondary: by slice (good-faith typical / good-faith difficult / red-teaming difficult); agreement between graders. WHAT WOULD COUNT AGAINST OUR READING. If independent graders put Darwin's adjusted score well above OpenEvidence's own adjusted 68.6, or show a large Darwin lead over physicians that survives adjustment, we report that plainly. NOT COMPARED. OpenEvidence did not release the rival models' answers, so Darwin vs Claude Fable 5 / GPT-5.6 / Gemini cannot be regraded; those comparisons stay as OpenEvidence's own chart reports them. AMENDMENT 2 (3 October 2026, before any Qwen result was read): the local Qwen grader switches from the full-precision (bf16) build to the 4-bit build of the same model (equal quality on our earlier benchmark suite) with a 16k context, because [the bf16 build was taking] about 2.4 minutes per item. The 21 bf16 items already graded are set aside unread. AMENDMENT 3 (3 October 2026, evening): GPT-5.6 Sol and GPT-5.5 became available again and graded Darwin's 410 answers as the plan originally specified, at their highest reasoning setting as in July. Their physician arm is their July 2026 grades of the same answers with the same prompt (408 and 409 of the 410 tasks). These two graders ran after the GPT-5.4, Gemini and Qwen results had been read; nothing else about the method changed. DATA published with the article: darwin-healthbench-grades.csv (every criterion, by task, arm and grader), darwin-healthbench-tasks.csv (task type, difficulty and the character counts used in the length adjustment), darwin-healthbench-task-scores.csv (per-task scores by grader and arm).