GPT 5.6 Sol Tops ProgramBench With Two Perfect Rebuilds
Researcher John Yang reports GPT 5.6 Sol leads ProgramBench with two perfect program rebuilds.
John Yang posted that GPT 5.6 Sol (xhigh) ranks first on ProgramBench after perfectly rebuilding cmatrix and hex. The benchmark, a collaboration among Meta Superintelligence Labs, Stanford, and Harvard, requires models to reimplement terminal programs using only documentation and runtime access. Yang states the solutions are novel rather than regurgitated training data, and the model writes in Python far more often than the originals' mix of Rust, Go, and C. Ofir Press noted the top resolved score sits at 1 percent while the easier metric reaches 16.5 percent.
Combined views
22.9K
9 posts, first seen 21d ago
