Corrective engineering · July 2026
Correcting experiment workflow failures
The July 2026 repair kept experimental inputs and raw evidence fixed while allowing recoverable checkout, transport, validation, and publication failures to be repaired in place.
- System
- Headless simulation and evaluation platform for card-game agents
- Failure mode
- Execution bookkeeping began dominating experimental throughput
- Correction
- Immutable evidence; repairable non-scientific execution
- Current outcome
- The repair now supports restartable cloud pickups, parallel candidate training, and an active automated capability campaign
The system being protected
Evaluation controls
The project combines a deterministic Python rules engine, a native Unity-game bridge, a stable state-and-action layer, learned and scripted agents, and parallel training and evaluation. Each surface can produce results that look more convincing than they are.
The original safeguards froze experimental inputs, retained provenance, separated training success from independent strength, attached claim ceilings to results, and failed closed when evidence could not be reconstructed. These controls identified the evaluation problems described below.
Evaluation records
Evaluation findings
Fixed ladder saturation
The fixed ladder was saturated. The system classified the result as an evaluation-sensitivity failure instead of converting it into a player-strength claim.
Teacher evaluation on held-out examples
A search teacher fit all 320 training labels but missed the frozen 90% held-out gate. No PPO budget was spent on unsupported labels.
Evaluation-only experiment
Four predeclared policy pairs failed the discrimination rule. The result identified a limitation in the evaluator; no new policy checkpoint was produced.
The bottleneck
Recoverable execution failures closed experiment attempts
The initial milestone-and-slice model fit a period when the project was discovering its trust boundaries. Over time, failed command transport, stale checkouts, schema mismatches, hash disagreements, and publication errors could close an attempt and require a new identity—even when the underlying mechanism had never been tested.
Repair scope
Fixed records and repairable failures
Evidence-bearing facts
- Hypothesis and mechanism
- Candidate, thresholds, and stop rules
- Measured inputs, seeds, and consumed run identity
- Raw scientific evidence bytes
Non-mechanism execution
- Checkout or base drift
- Command transport and reconstruction code
- Publication, validation, and packaging mechanics
- Harmless schema or diagnostic-field differences
An invalid attempt is never rewritten into valid evidence. A fresh experiment identity is required only when another measured run is needed or a scientific input actually changes.
What changed
Workflow changes
- 01
Recoverable failures stay with the work item
Bounded non-mechanism corrections no longer require successor-contract churn when owned paths, scientific inputs, and evidence identity remain fixed.
- 02
Git verifies tracked source files
Redundant HEAD equality, direct-parent requirements, registry hashes, and tracked-file hashes were removed. Explicit hashes remain for raw evidence and bytes Git does not carry.
- 03
Active registry cleanup
Completed and superseded history moved out of active coordination state. The live registry fell from 60 items and 418,755 bytes to two running items and 11,850 bytes.
What remains unproven
Unverified operational improvements
The forward-progress repair landed on July 24, 2026. At the inspected snapshot, the compact queue contained only two running items. That verifies implementation—not longitudinal success.
- Non-mechanism churn has not yet been shown to remain near zero.
- Elapsed-time scientific throughput has not yet been shown to improve.
- No current checkpoint is established as a broadly strong player.
- Simulator scope and cross-runtime parity remain bounded.
Work involved
Engineering responsibilities
Treated governance as an engineered subsystem and corrected it against observed failure modes.
Used evaluation saturation, insufficient discrimination, and failed generalization to decide which experiments to pursue.
Traced failures across Python, JavaScript, native execution, Git state, publication, and validation.
Used Codex and ChatGPT for implementation support while reviewing contracts, evidence, and corrections.