DeliveryGym unveiled for adaptive AI training in simulated Paris
DeliveryGym's creators report a 54.3% improvement in Qwen3-VL-4B's net income during a simulated shift, using earnings as rewards and failures to guide training.
TLDR
DeliveryGym's creators describe an Unreal Engine 5 reinforcement-learning environment where AI agents navigate Paris, deliver food and earn money. They say it identifies where agents struggle and generates harder tasks targeting those weaknesses, reporting a 54.3% improvement in Qwen3-VL-4B's net income in a shift. A contributor describes the engineering challenge as balancing photorealistic rendering with the pace of training runs. They say the system uses SPEAR for its Python interface to Unreal, alongside parallel simulator workers and asynchronous coordination of environment execution, inference and policy training.