Guardrails removed spam, off-topic, unclear, or duplicate replies.
Ask a question below.
Published answers will appear here.
with cost numbers. funny that people were talking about muse spark and grok-4.20 as ~equivalent bc cheap good new models, but here muse spark is ~best and grok is terrible.
i fed in all of my texts with a friend (~600k tokens over the last year) and asked all the models to comment on our relationship (with a few specifics). fable then graded how good it thought each comment was.
there's some amount of variability in fable's grading, but not too much. the rankings stay mostly the same
sophia relationship understanding eval results, as graded by fable 5
i messed up the cost numbers for some of these, correct cost numbers can be found in the other image i posted
Guardrails removed spam, off-topic, unclear, or duplicate replies.
Ask a question below.
Published answers will appear here.