← Back to OPTCG overview

ENGINEERING

Engineering decisions and evidence

These five problems cover replay, cross-runtime validation, evaluation design, experiment validity, and native state reconstruction. Each section records the decision, evidence, and remaining limits.

See how the systems connect in Project Atlas →
01Deterministic replay without the game client
Problem
A training simulator is useful only if captured native actions can be replayed and compared without depending on Unity or manual inspection.
Decision
Represent game state and legal actions explicitly, replay retained native transitions through the Python engine, and compare canonical post-action states.
Evidence
Strict replay currently covers 974 post-step states and 68 complete native legal-action sets across eight action families.
Remaining boundary
Coverage is sampled and Imu-scoped. Official rules are authoritative; native OPTCGSim is an execution and disagreement reference outside the retained traces.
02Validation across four runtimes
Problem
C#, TypeScript, Python, and Node can silently disagree about action names, float encodings, state identity, or evidence arithmetic.
Decision
Use versioned observation contracts, canonical artifacts, portable float encodings, and independent Node reconstruction for measured evidence.
Evidence
Current gates reconstruct sealed results across runtimes and fail closed on mismatched bindings, unsupported effects, or stale provenance.
Remaining boundary
There is not yet one fully validated cross-language action contract; translation exceptions and broader native coverage remain explicit risks.
03Finding an evaluation ladder that stopped measuring progress
Problem
A policy can look stronger when the fixed opponents and held-out positions have become too easy to separate meaningful improvements.
Decision
Treat the saturated ladder as a diagnostic, require paired seed-disjoint comparisons for decisions, and keep qualification claims behind stronger independent gates.
Evidence
The retained evidence records that the first-legal holdout and fixed ladder saturate while recent candidate decisions selected no unique endpoint.
Remaining boundary
The project still lacks a strong, broad, independent player-strength suite; current results cannot establish general playing strength.
04Stopping invalid experiments from driving more training
Problem
A run can produce plausible numbers even when its inputs changed, its budget was exceeded, or its evidence path failed before publication.
Decision
Separate attempt validity, scientific result, and plan disposition; require source-bound intent, bounded spend, raw-first retention, and fail-upward validation.
Evidence
The workflow preserves invalid attempts with scientific result none and prevents them from ranking candidates, authorizing retries, or justifying promotion.
Remaining boundary
These controls protect project decisions; they do not prove that an otherwise valid training mechanism will improve the player.
05Reconstructing arbitrary native game states
Problem
Parity and counterfactual tests need repeatable positions, not only whatever board state the graphical client happens to be showing.
Decision
Capture raw native snapshots, reset through OPTCGSim's own lifecycle, settle frames, and replay action prefixes against hash-bound anchors.
Evidence
The bridge and live client support save, reset, step, rematch, replay recording, and bounded replay-fork canaries from retained state anchors.
Remaining boundary
The native protocol is single-flight and file-based; reconstruction evidence is bounded to the tested states and does not imply family-wide parity.