ENGINEERING
Engineering decisions and evidence
These five problems cover replay, cross-runtime validation, evaluation design, experiment validity, and native state reconstruction. Each section records the decision, evidence, and remaining limits.
See how the systems connect in Project Atlas →01Deterministic replay without the game client+
- Problem
- A training simulator is useful only if captured native actions can be replayed and compared without depending on Unity or manual inspection.
- Decision
- Represent game state and legal actions explicitly, replay retained native transitions through the Python engine, and compare canonical post-action states.
- Evidence
- Strict replay currently covers 974 post-step states and 68 complete native legal-action sets across eight action families.
- Remaining boundary
- Coverage is sampled and Imu-scoped. Official rules are authoritative; native OPTCGSim is an execution and disagreement reference outside the retained traces.
02Validation across four runtimes+
- Problem
- C#, TypeScript, Python, and Node can silently disagree about action names, float encodings, state identity, or evidence arithmetic.
- Decision
- Use versioned observation contracts, canonical artifacts, portable float encodings, and independent Node reconstruction for measured evidence.
- Evidence
- Current gates reconstruct sealed results across runtimes and fail closed on mismatched bindings, unsupported effects, or stale provenance.
- Remaining boundary
- There is not yet one fully validated cross-language action contract; translation exceptions and broader native coverage remain explicit risks.
03Finding an evaluation ladder that stopped measuring progress+
- Problem
- A policy can look stronger when the fixed opponents and held-out positions have become too easy to separate meaningful improvements.
- Decision
- Treat the saturated ladder as a diagnostic, require paired seed-disjoint comparisons for decisions, and keep qualification claims behind stronger independent gates.
- Evidence
- The retained evidence records that the first-legal holdout and fixed ladder saturate while recent candidate decisions selected no unique endpoint.
- Remaining boundary
- The project still lacks a strong, broad, independent player-strength suite; current results cannot establish general playing strength.
04Stopping invalid experiments from driving more training+
- Problem
- A run can produce plausible numbers even when its inputs changed, its budget was exceeded, or its evidence path failed before publication.
- Decision
- Separate attempt validity, scientific result, and plan disposition; require source-bound intent, bounded spend, raw-first retention, and fail-upward validation.
- Evidence
- The workflow preserves invalid attempts with scientific result none and prevents them from ranking candidates, authorizing retries, or justifying promotion.
- Remaining boundary
- These controls protect project decisions; they do not prove that an otherwise valid training mechanism will improve the player.
05Reconstructing arbitrary native game states+
- Problem
- Parity and counterfactual tests need repeatable positions, not only whatever board state the graphical client happens to be showing.
- Decision
- Capture raw native snapshots, reset through OPTCGSim's own lifecycle, settle frames, and replay action prefixes against hash-bound anchors.
- Evidence
- The bridge and live client support save, reset, step, rematch, replay recording, and bounded replay-fork canaries from retained state anchors.
- Remaining boundary
- The native protocol is single-flight and file-based; reconstruction evidence is bounded to the tested states and does not imply family-wide parity.