• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    AutoBenchmark tests human input in AI-generated research benchmarks

    A thread introducing AutoBenchmark reports that detailed human guidance during the proposal stage worked better than a one-sentence idea or no human feedback in its tests.

    JW
    SW
    SC
    7 Sources, ,

    TLDR

    The AutoBenchmark thread describes a process in which agents propose benchmarks, other agents try to solve them, and an LLM reviews them. Its author reports that, in tests on two benchmarks, human feedback helped and detailed guidance worked better than a one-sentence idea. The thread also says human direction helped when an automated research loop stalled.

    Combined views

    35.1K

    7 Sources, first seen 11h ago

    Combined views

    35.1K

    7 Sources, first seen 11h ago

    415 likes
    11h ago
    first seen 11h ago
    415 likes
    4 comments
    392 saves
    100 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    4 comments
    392 saves
    100 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    7 Sources

    @jaseweston⏱️⚙️Introducing *AutoBenchmark* 📊🏁 - Creating benchmarks automatically - Benchmarking benchmark creation - Studying the role & impact of humans in the loop Blog post: https://facebookresearch.github.io/RAM/blogs/autobench Key takeaways: 1) We find human-agent collaboration gives big wins over agents alone - fine-grained feedback in ideation stage crucial - autobench can be used to measure this in the future with stronger agents 2) We show that it's possible to make *autoresearch benchmarks* for AI research using this recipe – full recursive improvement loop! 3) Important Ingredients: Autobenchmark creation works best with feedback from two sources: benchmark solvers + external verifiers (human+AI). 🧵1/5
    @wellecksRT @jaseweston: ⏱️⚙️Introducing *AutoBenchmark* 📊🏁 - Creating benchmarks automatically - Benchmarking benchmark creation - Studying the ro…
    @sarahcat21Human-in-the-loop semi-unsupervised environment/benchmark design. So so so cool.

    7 Sources

    @jaseweston⏱️⚙️Introducing *AutoBenchmark* 📊🏁 - Creating benchmarks automatically - Benchmarking benchmark creation - Studying the role & impact of humans in the loop Blog post: https://facebookresearch.github.io/RAM/blogs/autobench Key takeaways: 1) We find human-agent collaboration gives big wins over agents alone - fine-grained feedback in ideation stage crucial - autobench can be used to measure this in the future with stronger agents 2) We show that it's possible to make *autoresearch benchmarks* for AI research using this recipe – full recursive improvement loop! 3) Important Ingredients: Autobenchmark creation works best with feedback from two sources: benchmark solvers + external verifiers (human+AI). 🧵1/5
    @wellecksRT @jaseweston: ⏱️⚙️Introducing *AutoBenchmark* 📊🏁 - Creating benchmarks automatically - Benchmarking benchmark creation - Studying the ro…
    @sarahcat21Human-in-the-loop semi-unsupervised environment/benchmark design. So so so cool.