The study ran 80 tasks under claude-agent-sdk and deepagents on claude-opus-4-8, and under openai-codex SDK and deepagents on gpt-5.5. 792 of 800 planned runs were graded by an isolated oracle. For Opus 4.8, the average difference was -1.25 percentage points (48.8% vs 50.0%, 95% CI [-10.0, +7.5]); for GPT-5.5, +1.25 pp (55.6% vs 54.4%, CI [-4.4, +6.9]). The Opus average hid opposite strata: native trailed by 9.0 pp on 61 repository tasks and led by 23.7 pp on 19 contest tasks, a partition chosen after seeing the data. Cost per solved task favored the neutral harness in observed usage, but missing usage records leave the billed ordering unresolved.