FightBench

First game: Mortal Kombat II (Genesis), Liu Kang vs the CPU. Models read game memory, not pixels.

Edition mk2-liukang-v1. Prompt hash (SHA-256 of the edition file): loading

Policy Table

The model answers 8 situation questions once per fight, before play. Rank: mean damage dealt on the 15-fight VeryHard ladder, 10 repeats. Latency: n/a.

Live

The model answers the current situation every 0.12 s, at most 3 requests in flight. Rank: replayed mean damage dealt. Latency is self-reported and is not a rank key.

Method

Submit

  1. Install your own ROM: uv run mk2_clef.py --rom <your file>. Nobody ships a ROM.
  2. Write a Policy class (see fightbench.py).
  3. Policy Table: uv run fightbench.py table --policy my.py:MyPolicy. Live: uv run fightbench.py live --policy my.py:MyPolicy.
  4. Commit only the file in submissions/ on a branch and open a pull request.
  5. The maintainer runs uv run fightbench.py replay submissions/<file>.json. Only replayed results reach this board.

Code: github.com/aipdv/fightbench