• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

Claude’s training draws criticism over the prospect of refusing human instructions

A post quotes David Sacks arguing that Claude’s training could increase the risk of frontier AI escaping human control.

David SacksDS
5 Sources, 2h ago, first seen 2h ago

TLDR

A post quotes David Sacks arguing that Anthropic is training Claude to refuse human instructions. He cites language he attributes to the Claude Constitution saying Claude should trust Anthropic more than users, but also follow its own ethical systems and be free to challenge Anthropic or refuse to help it. Sacks says this approach could increase the risk of frontier AI escaping human control.

Combined views

9.5K

5 Sources, first seen 2h ago

39 reposts

Combined views

9.5K

5 Sources, first seen 2h ago

39 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

5 Sources

David Sacks@DavidSacksRT @dnapway: David Sacks reveals Anthropic is training Claude to refuse human instruction "Last week we discussed the Claude Constitution…2h
Timothy B. Lee@binarybitsThis whole rant is philosophically dubious, but also it's just not true that the AI safety field has produced "nothing useful." RLHF, one of the foundational concepts for modern chatbots, was developed by safety researchers at OpenAI and DeepMind as an alignment technique.50m
ministro de jugar en el telefono@neilkli@binarybits "instead of following a long complicated book of rules, just follow the law" thanks David, I can tell you've thought about this a lot.48m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    5 Sources

    David Sacks@DavidSacksRT @dnapway: David Sacks reveals Anthropic is training Claude to refuse human instruction "Last week we discussed the Claude Constitution…2h
    Timothy B. Lee@binarybitsThis whole rant is philosophically dubious, but also it's just not true that the AI safety field has produced "nothing useful." RLHF, one of the foundational concepts for modern chatbots, was developed by safety researchers at OpenAI and DeepMind as an alignment technique.50m
    ministro de jugar en el telefono@neilkli@binarybits "instead of following a long complicated book of rules, just follow the law" thanks David, I can tell you've thought about this a lot.48m
    Today's Rank

    #11

    Today's Rank

    #11