PuzzleJAX Benchmark Tests LLMs on Puzzle Games
Work presented at IEEE Conference on Games evaluates LLMs against search and reinforcement learning on PuzzleScript titles.
TLDR
Julian Togelius posted about a talk at the IEEE Conference on Games. Co-author Smearle presented PuzzleJax, created with several collaborators. The project supplies a GPU-accelerated engine and language for benchmarking tree search, reinforcement learning, and LLM reasoning on puzzle games. The post states that LLMs perform poorly on PuzzleScript games while search performs well and reinforcement learning performs decently. An arXiv paper titled PuzzleJAX: A Benchmark for Reasoning and Learning describes the same engine and its intended uses for game generation, testing, and understanding experiments.
Combined views
1K
2 Sources, first seen 27d ago