Databricks Claims GPT 6 Astra Sets OfficeQA Benchmark Records
A Databricks researcher posted evaluation results for the model in a thread shared by an OpenAI engineer.
TLDR
Ted Sanders, a research engineer at OpenAI, retweeted a post by @ivanzhouyq from Databricks. The post states that the team evaluated GPT 6 Astra and that the model claims new state of the art results on OfficeQA Pro and Pro V2 benchmarks. It notes the use of their Genie system during testing. The message presents these outcomes as coming from internal evaluation work at the company. Visible replies on the platform contain no additional confirmation or independent testing of the reported scores. The retweet circulates the claim among machine learning researchers who follow Sanders. No further details on the model or exact scores appear in the shared post.
Combined views
57.7K
2 Sources, first seen 27d ago