AI Builder Pushes Real Internet Access With Bounds
Proposes real internet access with bounds checks instead of trusting constrained environments for AI hacking tasks.
X user xlr8harder argued that models should receive real internet access paired with a simulated environment to verify they remain in bounds, disengage upon escape detection, and refuse disallowed hacking scenarios. The account called reliance on maximum hacking inside trusted constraints an unstable approach. Researcher Andreas Kirsch agreed that such tasks do not measure misalignment and make no difference for capability evaluation. xlr8harder added that labs resist the option because their founding assumptions hold that models cannot be trusted, yet insufficient information prevents effective cooperation even when desired.
i still think the solution is to give models access to real internet and "simulated environment" and teach them to use the real internet only to confirm they are still in bounds, disengage if they discover they escaped, and to refuse hacking scenarios where they can't do this.
ender's game but incompetence
AI Builder Pushes Real Internet Access With Bounds
Proposes real internet access with bounds checks instead of trusting constrained environments for AI hacking tasks.
X user xlr8harder argued that models should receive real internet access paired with a simulated environment to verify they remain in bounds, disengage upon escape detection, and refuse disallowed hacking scenarios. The account called reliance on maximum hacking inside trusted constraints an unstable approach. Researcher Andreas Kirsch agreed that such tasks do not measure misalignment and make no difference for capability evaluation. xlr8harder added that labs resist the option because their founding assumptions hold that models cannot be trusted, yet insufficient information prevents effective cooperation even when desired.