Card-game simulation and training

Decktrace

Decktrace contains simulators and self-play training tools for One Piece Card Game and Pokajan. Candidate agents are compared under matching starting conditions and tested against earlier policies before selection.

  • Automated policy selection
  • Paired starts and seat rotations
  • Documented evaluation limits

Game-specific rules and evaluation

  1. 01
    ModelRules and legal actionsversioned per game
  2. 02
    TrainSelf-play policiesno human labels
  3. 03
    MeasureBalanced comparisonsstarts + seats controlled
  4. 04
    RecordDecision + provenanceresult + claim boundary

Evaluation metrics
OPTCG measures paired two-seat win rate. Pokajan measures four-seat placement-weighted settlement. Results are reported for the tested simulator and matchup.

Current projects

Game results

Results from the implemented simulators, with the tested matchups and limitations described below.

Decktrace: OPTCG · established

OPTCG evaluation results

The selected policy passed paired-seat evaluation and comparisons with earlier policies for the fixed Imu matchup. The browser demo uses a separately bundled Generation 11 policy.

Simulator win rate
53.2%
Against the incumbent in paired-seat games
Conservative floor
50.3%
Promotion measure after uncertainty
  • Paired openings768
  • Archive comparisonsPass
  • Bundled browser policyGeneration 11
  • Native qualificationNot established
Current decisionRecorded training evaluation

The recorded evaluation met the simulator promotion criteria. This result does not describe the separately versioned browser demo. Qualification in the original client and performance across other matchups have not been established.

Confirmed in the checked-in Python simulator for the Imu mirror. Play the browser policy · Inspect the campaign.

Decktrace: Pokajan · established

Pokajan v2 evaluation results

Fresh four-seat confirmation rotated the learner through every seat against three incumbent copies on identical deals.

First-place finishes
39.1%
100 of 256 confirmation games
Top-two finishes
69.1%
177 of 256 confirmation games
  • Average finish2.04 of 4
  • Final coin edge+456 vs mean opponent
  • Archive comparisons6 / 6 pass
  • Invalid actions / timeouts0 / 0
  • Live-client parityNot run
Policy evaluationConfirmation and follow-up results

The v2 policy improved first-place rate, top-two rate, average finish, and terminal coins. Five targeted follow-ups then failed fresh game-outcome gates, so retaining v2 is the evidence-backed result.

Confirmed in the versioned Python simulator. Native control, live-client parity, and broad live-game strength are not established.

A 21-second platform tour

Training and evaluation process

The steps used to define a game, train candidate policies, and evaluate their results.

Stage 1 of 7 · Versioned game profileGame rules and scoring

Stage 1 of 7Versioned game profile

Game rules and scoring

Each project defines its own visible state, legal actions, scoring, and terminal conditions before training begins.

Evidence
OPTCG keeps a fixed two-seat Imu mirror; Pokajan keeps a four-seat hidden-information profile.
Claim ceiling
Each game has its own rules engine.
Compare game implementations →
GAME PROFILERULES + SCORE
State
Visible
Actions
Legal
Result
Terminal
game.profile.versioned

Five connected capabilities

System components

Open the technical map
Explore the five connected systems
01

Execute

Runs complete games from rules and legal actions that stay specific to each title.

Built
Separate deterministic engines, masked actions, and browser-ready simulator actors
Evidence
A trained OPTCG policy is available in the browser demo
Current boundary

OPTCG has a native execution/disagreement bridge; Pokajan client parity has not run.

Simulator architecture
02

Train

Lets policies improve through self-play without human games, labels, or promotion votes.

Built
Game-specific reinforcement-learning modules, shared-action policies, and resumable training operators
Evidence
OPTCG's latest champion cleared its measured strength and archive gates; Pokajan retained its v2 champion
Current boundary

A completed run is evidence of execution, not automatically evidence of strength.

See what is shared
03

Compare

Uses matching starting conditions and seat rotations to reduce variation when comparing policies.

Built
Common openings for two seats; common deals and four-seat rotation for Pokajan
Evidence
The latest champions used balanced comparisons; Pokajan confirmation used 256 games
Current boundary

Win rate and placement-weighted settlement answer different game-specific questions.

Compare the scorecards
04

Confirm

Requires a fresh conservative result and regression checks before a challenger replaces the champion.

Built
Uncertainty-aware lower bounds, invalid-action checks, and archive opponents
Evidence
The latest OPTCG champion cleared its uncertainty and archive gates; Pokajan cleared a +0.867 utility floor
Current boundary

Rules correctness remains a separate prerequisite for any strength claim.

Inspect the evidence gates
05

Improve

Keeps proven gains, diagnoses visible mistakes, and stops ideas that fail game-outcome evidence.

Built
Retained champions, focused diagnostics, bounded probes, and explicit claim ceilings
Evidence
The browser demo supports a trained OPTCG opponent
Current boundary

Additional experiments require a specific policy error or untested hypothesis.

Policy selection process

Implementation

Simulation and training infrastructure

Each game defines its own rules, observations, rewards, and evaluation criteria. Both implementations use deterministic simulation, self-play, and automated policy comparisons.

See the five platform principles
  • Deterministic simulation and masked legal actions
  • Self-play without human labels or promotion votes
  • Restartable training and source-bound checkpoints
  • Balanced evaluation with uncertainty and archive gates
  • Results with documented evaluation limits