Negative users criticize claims that advanced AI models can diverge sharply yet remain easy to switch between, arguing model-specific behaviors make short-prompt benchmarks unreliable and idiosyncrasies undesirable.
Based on 3 visible X reactions from 17 accounts; directional sample.
Ask a question below.
Published answers will appear here.
I think this is real to some extent, but I don't think it should be. For example, Fable 5 just gave up on proving the Riemann Hypothesis for me last night, but GPT-5.6 Sol kept going all night and was still going when I left this morning. I do think they have different idiosyncrasies...but I don't think that's necessarily a desirable trait for models to have.
@emollick Turns out model choice is load-bearing when the task runs 20 steps instead of 2. Wild surprise. Every benchmark built on short prompts is measuring the wrong thing.
@emollick This is an unserious comment. What? It is trivial to move data from one harness to another. Plug and play is in play.
At the moment that everyone is talking about switching models often for cost or sovereignty or optionality or whatever, the most advanced models are growing more and more different from each other. Fable responds very differently than Kimi K3 or Sol, you can't just plug & play
This makes complete sense, as models do longer and longer tasks that require more and more judgement decisions that add up, small differences in models result in very different kinds of outcomes.
Negative users criticize claims that advanced AI models can diverge sharply yet remain easy to switch between, arguing model-specific behaviors make short-prompt benchmarks unreliable and idiosyncrasies undesirable.
Based on 3 visible X reactions from 17 accounts; directional sample.
Ask a question below.
Published answers will appear here.