• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Anthropic shares internal metrics on AI’s role in developing future AI

    Anthropic says its snapshot tracks three things: how much AI research and development is done by AI, how well AI agents are overseen, and how computing resources are allocated.

    1 Source, 20d ago, first seen 20d ago

    TLDR

    Anthropic has shared three internal measures covering AI’s role in research, agent oversight and computing resources. The company says other frontier AI developers could publish the same measures and third parties could verify them. Its stated goal is to narrow the information gap between labs and the public, giving society more insight into how AI development is progressing.

    Combined views

    —

    1 Source, first seen 20d ago

    Combined views

    —

    1 Source, first seen 20d ago

    — likes
    — likes
    — comments
    — saves
    — reposts
    — comments
    — saves
    — reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Gabriel@gabriel_horwitzThis is actually super interesting and unbelievably transparent from Anthropic. The scenario that scares people most is AI building the next AI with no humans involved. Anthropic just measured how close they are to it. Not there yet. But in February under 1% of their AI research work was being led by the AI, and now it is 26%. Their scale runs 0 to 5. A 0 is no AI. A 3 is the AI doing big chunks of the work while a person watches closely and unblocks it whenever something unexpected comes up. A 4 is you hand it the problem, it goes and fixes it on its own, writes up what it did, and a human reads the writeup and decides whether to ship it. A 5 is nobody hands it anything. It notices the problem, fixes it, and ships it itself. Nothing at Anthropic is at 5! Over 90% of the work is at a 3 or higher, and that 26% is the share sitting at 4. For scale, about 30,000 AI agents are doing research and engineering inside the company at any given moment (on their main internal system, anyway). In August they made over a billion decisions. Every one of those passes through a monitor before it runs, and the monitor blocked about 1 in 47,000. Around 100,000 work logs a week get flagged for a second look, and roughly 50 of those reach an actual human. They got the 26% by having Claude read employees Slack messages and internal docs to catalog what everyone was working on, then a separate Claude judged how automated each task was. So the AI measured how much of the work the AI is doing. They say this openly, and in fairness their human staff agreed with the model's ratings more often than they agreed with each other (59% vs 35%). They're also bringing in outside evaluators to check all of this. Of all the computing power that went into AI research the week they checked, about 6% went to safety work. Their explanation is that safety research is mostly people designing experiments and doesn't burn chips the way training a new model does. For a lab all about safety, 6% of research spent on what they deem the “most important problem” feels drastically low.20d