4.7 KiB
Shardok AI Scoring Benchmarks
The current STANDARD scoring experiment results are recorded in
src/main/cpp/net/eagle0/shardok/ai_battle_simulator/benchmarks/experimental_results/.
They use 40 saved Shardok battles extracted from game_389ffb1901903118.zip.
The corrected all-40 run uses:
max_rounds = 31random_seed = 1- iterative-deepening timeout propagation fix from PR #6923
- AI setup placement
- forced single-option commands
- defender scoring fixed to
STANDARDfor every row - archery/start-fire capability derived like production setup
- per-unit flee capability preserved by real-battle extraction
- zero-troop combat heroes preserved as live units
NO_PROFESSIONtreated as a valid hero profession state, not as "no hero"
Shardok siege battles do not have a draw result. If the attacker has not won after 31 rounds, the
defender wins. The simulator records those cutoffs as max_rounds_reached defender wins.
All 40 saved battle payloads now produce valid simulator configs. A battle side can have zero
remaining troops as long as it has a hero; this is common for NO_PROFESSION heroes with no
commanded battalion.
An all-40 real-vs-sim comparison records whether the saved-game attacker/defender were human or AI
players. Factions 3 and 4 are human players; all other factions in this save are AI players.
The comparison uses broad winner side from the saved Shardok result payload, because exact replay
is not available from the compact result rows. Real end-of-battle troop totals come from the
all-factions Shardok ActionResultView stream's final changed_player_totals.
An AI-vs-AI smoke check also reruns a subset where neither side used faction 3 or 4.
Current Results
These artifacts are the current baseline for future scoring comparisons.
| Dataset | Valid battles | Invalid rows | Attacker wins | Defender wins | Notes |
|---|---|---|---|---|---|
| All-40 timeout-fix STANDARD baseline | 40 | 0 | 18 | 22 | standard_scoring_all40_post_bugfix_baseline.*; regenerated after fixing iterative-deepening timeout score propagation |
| All-40 real saved game | 40 known | 0 unknown | 19 | 21 | standard_scoring_all40_real_vs_sim.*; sim matches 39/40 known winners |
| AI-vs-AI smoke | 31 | 0 | 10 | 21 | 31/31 winner-side matches |
The remaining real-vs-sim winner-side miss is 11754 (real attacker, sim defender), which had an
AI attacker but a human defender in the saved game. The real attacker won with 2119 troops
remaining; the STANDARD-vs-STANDARD simulation replaces the human defender with AI and the defender
wins with 1000 troops remaining, so this row is not a comparable AI-vs-AI smoke miss. Battle
6248, previously the remaining AI-vs-AI miss, now matches as an attacker win after expected-impact
archery scoring was promoted to STANDARD.
The closest six valid baseline rows are used for the scalar and behavior experiment batches. The
single-run harness still has some nondeterminism even with random_seed = 1, so near-margin changes
should be repeated before being promoted.
Experiment Tooling
src/main/cpp/net/eagle0/shardok/ai_battle_simulator/benchmarks/run_real_battle_experiment.py
is the reusable real-battle experiment runner. It can read a saved-game zip or extracted game.db,
generate baseline/candidate simulator configs, run both sides, and write summary.csv,
comparison.csv, report.md, generated configs, and per-command trace JSONL files.
The repository keeps the current STANDARD baseline artifacts, but omits generated configs/traces for abandoned experiment batches. Regenerate those locally with the runner when a new comparison needs command-level traces.
src/main/cpp/net/eagle0/shardok/ai_battle_simulator/benchmarks/run_standard_baselines.py
regenerates the tracked STANDARD-vs-STANDARD baseline artifacts from a saved-game zip or extracted
game.db, including the all-40 baseline, real-vs-sim comparison, and AI-vs-AI smoke files.
The simulator supports --trace-jsonl=<path> and --run-label=<label> for detailed command traces.
Each row records the selected command, search depth, forced command count, and pre/post side troop
totals.
Expected-impact archery replaced the old flat archery-position value in STANDARD. A unit that can
shoot now gets half of the expected volley impact; a unit that is merely positioned for a future
volley gets one quarter. Promoting that behavior flips 6248 to match the real AI-vs-AI attacker
win without introducing an AI-vs-AI winner-side miss in the current saved-battle baseline.
The latest baseline refresh was generated after the iterative-deepening evaluator stopped letting timeout fallback scores compete with completed command evaluations. Future STANDARD-vs-EXPERIMENTAL comparisons should compare against these checked-in timeout-fix baseline artifacts.