Corrective engineering · July 2026

Correcting experiment workflow failures

The July 2026 repair kept experimental inputs and raw evidence fixed while allowing recoverable checkout, transport, validation, and publication failures to be repaired in place.

System
Headless simulation and evaluation platform for card-game agents
Failure mode
Execution bookkeeping began dominating experimental throughput
Correction
Immutable evidence; repairable non-scientific execution
Current outcome
The repair now supports restartable cloud pickups, parallel candidate training, and an active automated capability campaign

Evaluation controls

The project combines a deterministic Python rules engine, a native Unity-game bridge, a stable state-and-action layer, learned and scripted agents, and parallel training and evaluation. Each surface can produce results that look more convincing than they are.

The original safeguards froze experimental inputs, retained provenance, separated training success from independent strength, attached claim ceilings to results, and failed closed when evidence could not be reconstructed. These controls identified the evaluation problems described below.

Evaluation records

Evaluation findings

96–0

Fixed ladder saturation

The fixed ladder was saturated. The system classified the result as an evaluation-sensitivity failure instead of converting it into a player-strength claim.

49 / 64

Teacher evaluation on held-out examples

A search teacher fit all 320 training labels but missed the frozen 90% held-out gate. No PPO budget was spent on unsupported labels.

384

Evaluation-only experiment

Four predeclared policy pairs failed the discrimination rule. The result identified a limitation in the evaluator; no new policy checkpoint was produced.

Recoverable execution failures closed experiment attempts

The initial milestone-and-slice model fit a period when the project was discovering its trust boundaries. Over time, failed command transport, stale checkouts, schema mismatches, hash disagreements, and publication errors could close an attempt and require a new identity—even when the underlying mechanism had never been tested.

13registered boundary failures before cutover11 were integration, validation, or evidence failures—not mechanism failures
60items in the first worktree-first registry31 blocked; every blocker represented a non-mechanism failure
15blocked items with measured evidencePackaging or validation prevented completion after the science had run

Repair scope

Fixed records and repairable failures

Immutable

Evidence-bearing facts

  • Hypothesis and mechanism
  • Candidate, thresholds, and stop rules
  • Measured inputs, seeds, and consumed run identity
  • Raw scientific evidence bytes
Repairable in place

Non-mechanism execution

  • Checkout or base drift
  • Command transport and reconstruction code
  • Publication, validation, and packaging mechanics
  • Harmless schema or diagnostic-field differences

An invalid attempt is never rewritten into valid evidence. A fresh experiment identity is required only when another measured run is needed or a scientific input actually changes.

Workflow changes

  1. 01

    Recoverable failures stay with the work item

    Bounded non-mechanism corrections no longer require successor-contract churn when owned paths, scientific inputs, and evidence identity remain fixed.

  2. 02

    Git verifies tracked source files

    Redundant HEAD equality, direct-parent requirements, registry hashes, and tracked-file hashes were removed. Explicit hashes remain for raw evidence and bytes Git does not carry.

  3. 03

    Active registry cleanup

    Completed and superseded history moved out of active coordination state. The live registry fell from 60 items and 418,755 bytes to two running items and 11,850 bytes.

What remains unproven

Unverified operational improvements

The forward-progress repair landed on July 24, 2026. At the inspected snapshot, the compact queue contained only two running items. That verifies implementation—not longitudinal success.

  • Non-mechanism churn has not yet been shown to remain near zero.
  • Elapsed-time scientific throughput has not yet been shown to improve.
  • No current checkpoint is established as a broadly strong player.
  • Simulator scope and cross-runtime parity remain bounded.

Work involved

Engineering responsibilities

Systems judgment

Treated governance as an engineered subsystem and corrected it against observed failure modes.

Experimental design

Used evaluation saturation, insufficient discrimination, and failed generalization to decide which experiments to pursue.

Boundary debugging

Traced failures across Python, JavaScript, native execution, Git state, publication, and validation.

Tool use

Used Codex and ChatGPT for implementation support while reviewing contracts, evidence, and corrections.