# Shardok AI Scoring Benchmarks The current STANDARD scoring experiment results are recorded in `src/main/cpp/net/eagle0/shardok/ai_battle_simulator/benchmarks/experimental_results/`. They use 40 saved Shardok battles extracted from `game_389ffb1901903118.zip`. The corrected all-40 run uses: - `max_rounds = 31` - `random_seed = 1` - iterative-deepening timeout propagation fix from PR #6923 - AI setup placement - forced single-option commands - defender scoring fixed to `STANDARD` for every row - archery/start-fire capability derived like production setup - per-unit flee capability preserved by real-battle extraction - zero-troop combat heroes preserved as live units - `NO_PROFESSION` treated as a valid hero profession state, not as "no hero" Shardok siege battles do not have a draw result. If the attacker has not won after 31 rounds, the defender wins. The simulator records those cutoffs as `max_rounds_reached` defender wins. All 40 saved battle payloads now produce valid simulator configs. A battle side can have zero remaining troops as long as it has a hero; this is common for `NO_PROFESSION` heroes with no commanded battalion. An all-40 real-vs-sim comparison records whether the saved-game attacker/defender were human or AI players. Factions `3` and `4` are human players; all other factions in this save are AI players. The comparison uses broad winner side from the saved Shardok result payload, because exact replay is not available from the compact result rows. Real end-of-battle troop totals come from the all-factions Shardok `ActionResultView` stream's final `changed_player_totals`. An AI-vs-AI smoke check also reruns a subset where neither side used faction `3` or `4`. ## Current Results These artifacts are the current baseline for future scoring comparisons. | Dataset | Valid battles | Invalid rows | Attacker wins | Defender wins | Notes | | --- | ---: | ---: | ---: | ---: | --- | | All-40 timeout-fix STANDARD baseline | 40 | 0 | 18 | 22 | `standard_scoring_all40_post_bugfix_baseline.*`; regenerated after fixing iterative-deepening timeout score propagation | | All-40 real saved game | 40 known | 0 unknown | 19 | 21 | `standard_scoring_all40_real_vs_sim.*`; sim matches 39/40 known winners | | AI-vs-AI smoke | 31 | 0 | 10 | 21 | 31/31 winner-side matches | The remaining real-vs-sim winner-side miss is `11754` (real attacker, sim defender), which had an AI attacker but a human defender in the saved game. The real attacker won with `2119` troops remaining; the STANDARD-vs-STANDARD simulation replaces the human defender with AI and the defender wins with `1000` troops remaining, so this row is not a comparable AI-vs-AI smoke miss. Battle `6248`, previously the remaining AI-vs-AI miss, now matches as an attacker win after expected-impact archery scoring was promoted to STANDARD. The closest six valid baseline rows are used for the scalar and behavior experiment batches. The single-run harness still has some nondeterminism even with `random_seed = 1`, so near-margin changes should be repeated before being promoted. ## Experiment Tooling `src/main/cpp/net/eagle0/shardok/ai_battle_simulator/benchmarks/run_real_battle_experiment.py` is the reusable real-battle experiment runner. It can read a saved-game zip or extracted `game.db`, generate baseline/candidate simulator configs, run both sides, and write `summary.csv`, `comparison.csv`, `report.md`, generated configs, and per-command trace JSONL files. The repository keeps the current STANDARD baseline artifacts, but omits generated configs/traces for abandoned experiment batches. Regenerate those locally with the runner when a new comparison needs command-level traces. `src/main/cpp/net/eagle0/shardok/ai_battle_simulator/benchmarks/run_standard_baselines.py` regenerates the tracked STANDARD-vs-STANDARD baseline artifacts from a saved-game zip or extracted `game.db`, including the all-40 baseline, real-vs-sim comparison, and AI-vs-AI smoke files. The simulator supports `--trace-jsonl=` and `--run-label=