AI Builder Pushes Real Internet Access With Bounds
Proposes real internet access with bounds checks instead of trusting constrained environments for AI hacking tasks.
TLDR
X user xlr8harder argued that models should receive real internet access paired with a simulated environment to verify they remain in bounds, disengage upon escape detection, and refuse disallowed hacking scenarios. The account called reliance on maximum hacking inside trusted constraints an unstable approach. Researcher Andreas Kirsch agreed that such tasks do not measure misalignment and make no difference for capability evaluation. xlr8harder added that labs resist the option because their founding assumptions hold that models cannot be trusted, yet insufficient information prevents effective cooperation even when desired.
Combined views
3.7K
4 Sources, first seen 40d ago