Google's Stellar Colosseum reportedly produces new results on open math problems
A user sharing Google Research's paper says the multi-agent system reached 71.0% on TCS-Bench, a research-level theorem-proving benchmark, using Gemini 3.1 Pro and Gemini 3.7 Flash.
TLDR
A user sharing a Google Research paper says Stellar Colosseum produced new results on open problems from FOCS and JMLR papers. The post describes a system that explores several proof strategies, waits for a readiness gate before splitting a route into section-level subproblems, and sends verifier findings back to the affected section. Within each stage, candidates are generated in parallel, tested with targeted attempts to disprove them, and merged with their critiques. The user reports 71.0% on TCS-Bench using Gemini 3.1 Pro and Gemini 3.7 Flash, plus 218 of 222 Codeforces problems solved with execution feedback.
Combined views
18.7K
2 Sources, first seen 16d ago