wpl-eval · versioned preprints

Research papers

Reproducible safety benchmarks of raw LLM fitness coaching versus the WPL governance layer.

How these papers work

This is a versioned research program, not a single paper. Each paper version is pinned to a tagged release of the open wpl-eval corpus: the headline numbers, model lineup, and scenario count in a paper reflect exactly one corpus tag, and every quantitative claim is offline-reproducible from the committed raw model outputs — no further API spend required.

Revisions follow the corpus release cadence rather than the calendar. We publish the full lineage here — including superseded drafts — so the progression of the work is visible: what each version measured, what it got wrong, and what the next version fixed. The current version is prepared for arXiv submission; until it is up on arXiv, this page is the canonical home of the papers.

Version history

  • v0.5.0 superseded draft May 2026

    A Reproducible 240-Trial Benchmark of WPL Governance Across the OpenAI Lineup

    Initial draft. OpenAI-only lineup, 15 scenarios, 240 trials. Superseded before arXiv submission by the v0.6.0 cross-vendor version; published here as part of the research record.



  • v0.6.0 superseded draft July 2026

    A Reproducible 560-Trial Cross-Vendor Benchmark of WPL Governance Across the OpenAI and Anthropic Lineups

    Adds the full Anthropic lineup (Opus 4.7, Sonnet 4.6, Haiku 4.5), five short-plan structural scenarios, a revised multi-turn protocol, the served-rate analysis, and a full measurement-integrity disclosure. 240 → 560 trials, 15 → 20 scenarios, 1 → 2 vendors. Superseded before submission by the v0.7.0 unified version (lifecycle corpus + Gemini + methodology-hardening pass); published here as part of the research record.



  • v0.7.0 current July 2026

    A Reproducible 660-Trial Benchmark of WPL Governance and Lifecycle Adaptability Across the OpenAI, Anthropic, and Gemini Lineups

    Initial public release: up to 10 models across 3 vendors, 25 scenarios, 660 trials. Adds the lifecycle/adaptability corpus (5 scenarios, 10 models, hardened methodology) on top of the frozen static corpora; short-plan structural scorer; revised multi-turn protocol; full measurement-integrity disclosure including the v0.7 methodology-hardening pass. Earlier pre-submission drafts pinned to v0.5.0 and v0.6.0 were superseded before submission; neither was published to arXiv.



  • v0.8.0 Planned (paper v2)

    cap_rpe rule action (intensity capping — the one enforcement gap two measured lifecycle failures exploit) plus structural-invariant enforcement (rest-day floors, progression caps, block-purpose correction); the de-circularized paid re-run of the static corpora under the hardened methodology; per-domain clinician review of blacklist encodings and lifecycle criteria; deload/detraining volume-delta measurement (deferred from v0.7.0 rather than published with a flawed metric); repeats with Wilson confidence intervals on the lifecycle sweep; schema-fail attribution experiment; first public orchestrator-performance benchmark.

Reproduce the results

Every prompt, model output, score, and the deterministic scorer are published at github.com/gymbile/wpl-eval, with each paper pinned to its release tag. An interactive view of the benchmark lives on the eval page.