MWMarcus Wong

Software engineer · AI systems, simulation & backend infrastructure

I build the systems that let AI agents train, compete, and prove they improved.

I work across deterministic simulation, reinforcement learning, evaluation, and resilient backend automation. Decktrace turns complex card games into self-improving capability labs with measured results and explicit simulator boundaries.

Current simulator champion

Archive-safe champion

Controlled paired evidence cleared promotion and archive gates.

Subject: champion generation-019-candidate-03 · Opponent: generation-018-candidate-02 · 1,536 games

Paired score
53.2%
Observed point share in the controlled matchup.
Adjusted lower bound
50.3%
Conservative floor after uncertainty.
Paired openings
768
Common openings tested from both seats.
Invalid actions
0
Promotion requires zero.
Automatic promotion gateAdjusted lower bound > 50%

Promotion also checks invalid actions and comparisons with archived champions. Native qualification remains a separate lane.

  • Archive safety gatePass
  • Native qualificationNot run
Observed behavior
Critical misses
0 forced lethal · 1 immediate survival
Counts; opportunity totals are not recorded.
Event plays
100% effectful · 0% trash-only
5,410 total event plays
Counter spend
95.4% avoided hit · 4.6% still hit
27,648 resolved spend windows
DON attachment actions
9.5% to rested · 9.6% already attacked
15,330 attachment actions

Evidence is scoped to the checked-in Imu mirror and simulator runtime. Play the latest champion · Open the full campaign.

Featured project

A card-game capability platform for self-play, controlled competition, and evidence-bound promotion across game-specific simulators.

Card games · Simulation · Reinforcement learning · Evaluation

From game rules to measured, auditable agents.

Decktrace keeps each game's rules and scorecard explicit while reusing the same evidence discipline: self-play, balanced comparison, conservative confirmation, and correction when an assumption fails.

Built
Game-specific deterministic simulators, shared-action masked policies, resumable operators, balanced evaluation, and automatic champion promotion
Capability race
Four arms reach 256,000 steps, three continue to 1,024,000, and two finalists train across a 2,048,000-step horizon
Current OPTCG champion
The latest champion cleared automated strength, archive, validity, and browser-export parity gates
Honest boundary
The latest champion powers browser Play, while native qualification and arbitrary-deck strength remain unestablished
One generation · one controlled race
  1. 01
    TrainFour explicit armsfresh + champion lineages
  2. 02
    ScreenThree survivorscommon paired openings
  3. 03
    FinalizeTwo candidates2,048,000-step horizon
  4. 04
    RecordDecision + provenanceaudit-ready evidence
LATEST CHAMPIONCURRENTarchive-safe training + Play opponent
SimulationDeterministic games

A headless Python rules engine turns complete card-game play into fast, reproducible self-play.

Training4 → 3 → 2

Four explicit PPO treatments race through controlled rungs while weak candidates stop early.

EvaluationEvidence-gated

The latest champion cleared the simulator promotion, archive, validity, and browser-export parity gates.

Agentic engineering in practice

Automation accelerated the campaign. Explicit boundaries make correction visible.

Codex and ChatGPT accelerated repository analysis, implementation, testing, and parallel investigation. The training operator now restores its own state, runs independent arms concurrently, checkpoints safely, and publishes a recruiter-readable status view.

I remained responsible for the architecture, experimental boundaries, generated changes, evidence review, and the choice to make game outcomes—not human labels or manual judgment—the authority for promotion.

See the evidence behind that claim