• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    LongHarness introduced as a benchmark for long-context harnesses

    The original post describes LongHarness as a challenging evaluation for systems such as RLMs and coding agents.

    CL
    XY
    2 Sources, 14h ago, first seen 14h ago

    TLDR

    A post introducing LongHarness shares a new paper on evaluating long-context harnesses, including RLMs and coding agents. It describes LongHarness as a challenging benchmark for that work.

    Combined views

    3.5K

    2 Sources, first seen 14h ago

    54 likes

    Combined views

    3.5K

    2 Sources, first seen 14h ago

    54 likes
    3 comments
    28 saves
    22 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    3 comments
    28 saves
    22 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @xiye_nlpNew 📜 on evaluating long-context harnesses like RLMs and coding agents Introducing LongHarness, a challenging benchmark evaluating both accuracy and efficiency. We find 1) clear accuracy gaps 2) >10× efficiency differences, even with same LM and similar accuracy 🧵
    @ChengleiSiRT @xiye_nlp: New 📜 on evaluating long-context harnesses like RLMs and coding agents Introducing LongHarness, a challenging benchmark eval…

    2 Sources

    @xiye_nlpNew 📜 on evaluating long-context harnesses like RLMs and coding agents Introducing LongHarness, a challenging benchmark evaluating both accuracy and efficiency. We find 1) clear accuracy gaps 2) >10× efficiency differences, even with same LM and similar accuracy 🧵
    @ChengleiSiRT @xiye_nlp: New 📜 on evaluating long-context harnesses like RLMs and coding agents Introducing LongHarness, a challenging benchmark eval…
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet