• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Rohan Paul on OpenAI Agent Task Durations

    Rohan Paul shares OpenAI data on agent performance across task lengths.

    RP
    1 Source, 24d ago, first seen 24d ago

    TLDR

    Rohan Paul posted that time horizon is a useful measure of agent capability. He stated that at OpenAI, success without human intervention falls from 86 percent on tasks under 15 minutes to around 16 percent on tasks in the 64- to 128-hour range. The post includes a stacked bar chart titled Longer t and notes that runs requiring intervention increase with duration. Paul is a Bengaluru-based machine learning engineer who writes an AI newsletter.

    Combined views

    6.7K

    1 Source, first seen 24d ago

    Combined views

    6.7K

    1 Source, first seen 24d ago

    26 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    26 likes
    7 comments
    10 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 comments
    10 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @rohanpaul_aiI think "time horizon" is becoming one of the more useful ways to talk about agent capability. At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common. A model can be extremely capable locally and still be unreliable over a long execution trajectory. Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene? maybe we are moving towards a benchmark, something like human minutes consumed per completed task.

    1 Source

    @rohanpaul_aiI think "time horizon" is becoming one of the more useful ways to talk about agent capability. At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common. A model can be extremely capable locally and still be unreliable over a long execution trajectory. Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene? maybe we are moving towards a benchmark, something like human minutes consumed per completed task.