• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Engineer Proposes 1-17 Scale for LLM Response Ratings

    Florian Brand replies to Hamel Husain suggesting a 1-17 scale for GPT self-ratings.

    BT
    FB
    HH
    6 Sources, 29d ago, first seen 29d ago

    TLDR

    Research Engineer Florian Brand at Prime Intellect replied to @HamelHusain in a thread on LLM evaluations and benchmarking. Brand wrote that evaluators should instead use a scale and gave the example of asking gpt 4.1 mini how much it likes a response on a scale from 1-17. A second reply from the same account simply stated yes. The posts are visible among replies on X and address methods for obtaining model feedback during benchmarking work.

    Combined views

    12K

    6 Sources, first seen 29d ago

    Combined views

    12K

    6 Sources, first seen 29d ago

    36 likes
    36 likes
    6 comments
    1 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 comments
    1 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    @xlr8harder@xeophon did sol write it
    @HamelHusain@xeophon is that a reasoning trace you saw while running an eval? 🤣
    @xeophon@HamelHusain they should instead use a scale! "hey gpt 4.1 mini, how much do you like the response on a scale from 1-17?"
    @tunguzReasoning traces? You mean tweets?

    6 Sources

    @xlr8harder@xeophon did sol write it
    @HamelHusain@xeophon is that a reasoning trace you saw while running an eval? 🤣
    @xeophon@HamelHusain they should instead use a scale! "hey gpt 4.1 mini, how much do you like the response on a scale from 1-17?"
    @tunguzReasoning traces? You mean tweets?