Engineer Proposes 1-17 Scale for LLM Response Ratings
Florian Brand replies to Hamel Husain suggesting a 1-17 scale for GPT self-ratings.
Research Engineer Florian Brand at Prime Intellect replied to @HamelHusain in a thread on LLM evaluations and benchmarking. Brand wrote that evaluators should instead use a scale and gave the example of asking gpt 4.1 mini how much it likes a response on a scale from 1-17. A second reply from the same account simply stated yes. The posts are visible among replies on X and address methods for obtaining model feedback during benchmarking work.
Combined views
1K
4 posts, first seen 16h ago
Engineer Proposes 1-17 Scale for LLM Response Ratings
Florian Brand replies to Hamel Husain suggesting a 1-17 scale for GPT self-ratings.
Research Engineer Florian Brand at Prime Intellect replied to @HamelHusain in a thread on LLM evaluations and benchmarking. Brand wrote that evaluators should instead use a scale and gave the example of asking gpt 4.1 mini how much it likes a response on a scale from 1-17. A second reply from the same account simply stated yes. The posts are visible among replies on X and address methods for obtaining model feedback during benchmarking work.