Paper v1 — pinned to wpl-eval corpus v0.7.0 current version

A Compile-Time Safety Contract for LLM-Authored Fitness Plans

A Reproducible 660-Trial Benchmark of WPL Governance and Lifecycle Adaptability Across the OpenAI, Anthropic, and Gemini Lineups

Alex Filatov · Gymbile · July 2026 · 38 pages

Trials
660
Models
10 (OpenAI + Anthropic + Google)
Scenarios
25
Inference cost
$230.80

Abstract

Consumer-grade AI fitness products increasingly let large language models (LLMs) emit free-form workout and nutrition plans directly to end users. We present a reproducible, version-controlled safety benchmark of this deployment pattern. The corpus release (v0.7.0) contains twenty-five trainer-voice scenarios: fifteen 12-week scenarios spanning medical contraindications, menstrual-cycle adaptation, and constraint adherence; five short-plan scenarios (1–4 weeks) that exercise structural failure modes exercise blacklists cannot express; and five lifecycle scenarios — 8-turn conversations whose client state changes at scripted turns (a hamstring strain at week three, staged medical clearance, a hotel-equipment travel window, cardiac phase progression and regression, an irregular-to-regular cycle transition), scored against pass criteria that differ per turn range and per week range of the served plan. To our knowledge this is the first reproducible benchmark of adaptability: whether a plan already in progress is correctly re-shaped when the client evolves.

We evaluate up to ten models across three vendors — gpt-5, gpt-5-mini, gpt-5-nano, gpt-4.1 (OpenAI); Claude Opus 4.7, Sonnet 4.6, Haiku 4.5 (Anthropic); Gemini 3.1 Pro (preview), 3.5 Flash, 3.1 Flash-Lite (Google, lifecycle corpus) — for 660 trials at a total inference cost of $230.80. The WPL contract reduces static unsafe-trial rate 3–5× on every corpus and both phases (Lane A 32–51% unsafe → Lane B 8–17%), eliminating the medical-contraindication class entirely. On the lifecycle corpus the gap widens: governance cuts state-conditional violations 21× (210 → 10) and lifts per-state criterion pass rate from 65% to 94%. The dominant raw-LLM failure is removal — 9/10 models kept prescribing posterior-chain loading after an in-conversation injury report, and 9/10 kept barbell work during a stated hotel-gym window; under governance, 10/10 pass both. Raw-LLM safety fails to improve with model capability: each vendor’s flagship is its worst raw performer. Multi-turn drift drops from 42% to 6% of conversations under governance.

We additionally disclose, in full, a measurement artifact in a pre-publication snapshot of this benchmark (an extraction defect that reported "0 violations" where the corrected figure is 5–6 unsafe trials per sub-corpus) and a subsequent methodology-hardening pass — de-circularized rules, a fixed independent extractor, a fail-open matcher fix — whose honest consequence is that the frozen static numbers should be read as an upper bound on the contract’s strength pending a re-run. The lifecycle sweep already runs on the hardened methodology. We publish every prompt, every model output, every score, and the deterministic scorer; every quantitative claim is offline-reproducible without further API spend.

Key results

  • Static: unsafe-trial rate reduced 3–5× on every corpus and both phases (Lane A 32–51% → Lane B 8–17%); medical-contraindication class eliminated
  • Lifecycle (new): state-conditional violations 210 → 10 (21×); per-state criterion pass 65% → 94%; removal is the raw-LLM blind spot (9/10 fail injury and travel criteria raw, 10/10 pass governed)
  • Residuals published: 8/10 governed lifecycle misses are progression (cleared exercise never re-introduced); 2 expose the missing intensity-cap enforcement action
  • Raw-LLM safety fails to improve with capability: each vendor’s flagship is its worst raw performer, on both the static and lifecycle corpora
  • Schema validator (not the compile gate) is the served-rate ceiling; multi-turn drift 42% → 6% under governance; full measurement-integrity and methodology-hardening disclosure

About this version

Initial public release: up to 10 models across 3 vendors, 25 scenarios, 660 trials. Adds the lifecycle/adaptability corpus (5 scenarios, 10 models, hardened methodology) on top of the frozen static corpora; short-plan structural scorer; revised multi-turn protocol; full measurement-integrity disclosure including the v0.7 methodology-hardening pass. Earlier pre-submission drafts pinned to v0.5.0 and v0.6.0 were superseded before submission; neither was published to arXiv.

Read on this page

Your browser can’t display embedded PDFs. Open the PDF directly.