First game: Mortal Kombat II (Genesis), Liu Kang vs the CPU. Models read game memory, not pixels.
Edition mk2-liukang-v1. Prompt hash (SHA-256 of the edition file): loading
The model answers 8 situation questions once per fight, before play. Rank: mean damage dealt on the 15-fight VeryHard ladder, 10 repeats. Latency: n/a.
The model answers the current situation every 0.12 s, at most 3 requests in flight. Rank: replayed mean damage dealt. Latency is self-reported and is not a rank key.
script is the rules-bot floor row. random picks uniformly from the 17 moves.uv run mk2_clef.py --rom <your file>. Nobody ships a ROM.Policy class (see fightbench.py).uv run fightbench.py table --policy my.py:MyPolicy. Live: uv run fightbench.py live --policy my.py:MyPolicy.submissions/ on a branch and open a pull request.uv run fightbench.py replay submissions/<file>.json. Only replayed results reach this board.