v2 evaluation results
The v2 policy finished in the top two in 177 of 256 confirmation games. Later experiments did not establish a reliable improvement, so v2 remains the selected policy.
- First-place finishes
- 39.1% 100 of 256 confirmation games
- Top-two finishes
- 69.1% 177 of 256 confirmation games
- Average finish2.04 of 4
- Final coin edge+456 vs mean opponent
- Archive comparisons6 / 6 pass
- Invalid actions / timeouts0 / 0
- Live-client parityNot run
Scoring and selection record
Adding a discard-rank observation produced the v2 policy. It was retained after Generations 0 and 1 and five follow-up experiments: tied-rank selection, completion outs, claim liability, one-draw payout, and low-meld deferral.
Each ending turns coins into a score: first place uses 2.5× final coins, second uses 1.5×, and third or fourth uses 1×; the 1,000-coin starting stake is then subtracted and used to normalize the result. The champion averaged +1.135 of those stake-sized utility units over the mean of three incumbents. Its +0.867 lower 95% bound is the conservative floor for that same advantage—not 0.867 coins or places.
Five focused follow-ups failed fresh gates. The one-draw payout controller won a small screen but flattened on disjoint confirmation, while forcing low melds to pass was worse on average. Retain v2 and do not build deeper search without a newly demonstrated policy error.
Claim ceiling: the v2 policy is stronger in the versioned Python simulator. Native control, client parity, and broad live-game strength are not established.