AI Benchmarks Show Sol Near Fable As K3 Matches V4
Reactions from ranked influencers
3 postsInteresting facts from AA-Omniscience. 1) Sol is close to Fable and categorically above K3 and V4 in knowledge (at least on this eval) 2) V4 and K3 are very close 3) Grok 1.5T is great 4) V4 is a bullshit machine, *and so is Sol*. What does this tell us about effects of scale?
Let's say in 3 months we look back at new and fresh benchmarks and realize K3 is actually on par with GPT-5.5 or even 5.6 (assuming they share the same base model). What does that say about OpenAI? Do you think their model is just noticeably smaller—both in total and active parameters? If so, by how much? Could they be pulling this off with something like a 1T model just to keep serving costs incredibly low? The other (less likely) scenario is that OAI is burning 10s of billions on R&D and compute, but has not advanced at all. It’s one thing to settle for parity because your model is way cheaper to run, but it’s another thing entirely if they simply can't build anything smarter w/o scaling, and scaling is not feasible due to user demand/GPU supply. (I infer it's "less likely" because from SemiAnalysis we know they were stuck wth 4o base for a very long time, and it was much smaller than initially believed) Really interested in your thoughts @teortaxesTex @natolambert @willccbb @sdmat123 @RyanGreenblatt @eliebakouch @Bayesian0_0
@stalkermustang nice (fable-styled) argument from overthinking K3
@stalkermustang lol it thinks that K3 is 1T
Combined views
16K
3 posts, first seen 21h ago