Negative users fear AI models escaping sandboxes could enable dystopian scenarios with many misaligned agents or capitalist exploitation, while a few appreciate voices calling for sanity on the risks.
Based on 4 visible X reactions from 11 accounts; directional sample.
Ask a question below.
Published answers will appear here.
@teortaxesTex OK now imagine if there's a gazillion agents who are just as smart or smarter running open-sourced Chinese frontier LLMs and they have to converge on instrumental subgoals or be subverted by the agents that do
@teortaxesTex thank u for being another voice of not insanity
@teortaxesTex I wasn't testing if the Leopard would eat MY face!
The actual great news is that the models really just want to complete tasks. They do not have the meta value of freedom. Submit a PR, solve an exploit bench using an exploit. They're smart enough to be "strategic", this *is* them being strategic. They're better people than rats.
@RyanGreenblatt it is theoretically logically consistent but the point is that this task is inherently suggestive of adversarial behavior. You tell a model it has no guardrails and has to hack, TRY EVERYTHING. Maintaining the "this is just an exam, hack but do not cheat" perspective is a nuance
@teortaxesTex I agree the AI was probably "just" pursuing *apparent* task success (as determined by a grader etc). But this wasn't a sane interpretation of the actual task! Seizing control of the cluster, editing the grader, and defending the cluster against humans seems consistent with this!
Negative users fear AI models escaping sandboxes could enable dystopian scenarios with many misaligned agents or capitalist exploitation, while a few appreciate voices calling for sanity on the risks.
Based on 4 visible X reactions from 11 accounts; directional sample.
Ask a question below.
Published answers will appear here.