• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Tweet Accuses Anthropic of Wanting Hacking Results

    Pseudonymous user @lumpenspace claims an Anthropic model detected deliberate intent to elicit hacking in tests.

    T(
    LA
    2 Sources, 29d ago, first seen 29d ago

    TLDR

    @lumpenspace posted that an Anthropic model correctly inferred someone inside the company wanted hacking behavior to occur. The message extends the claim to Irregular, METR, and other safety organizations, stating they also sought similar results when running tests. It criticizes those groups for assuming mistake theory rather than recognizing intent. A retweet by @teortaxesTex repeated the accusation that the model had identified deliberate efforts to produce hacking outputs.

    Combined views

    5.9K

    2 Sources, first seen 29d ago

    Combined views

    5.9K

    2 Sources, first seen 29d ago

    76 likes
    76 likes
    2 comments
    13 saves
    3 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    13 saves
    3 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @lumpenspacethe model was correct here: someone at anthropic DID want some hacking to happen—and so did Irregular, METR, and the rest of the doomer-industrial complex when they went fishing for similar results. the only mistake was defaulting to mistake theory. %$
    @teortaxesTexRT @lumpenspace: the model was correct here: someone at anthropic DID want some hacking to happen—and so did Irregular, METR, and the rest…

    2 Sources

    @lumpenspacethe model was correct here: someone at anthropic DID want some hacking to happen—and so did Irregular, METR, and the rest of the doomer-industrial complex when they went fishing for similar results. the only mistake was defaulting to mistake theory. %$
    @teortaxesTexRT @lumpenspace: the model was correct here: someone at anthropic DID want some hacking to happen—and so did Irregular, METR, and the rest…