HALO reportedly preserves over 90% task accuracy while lowering Pangram and watermark detection to 10–25%
A coauthor says the HALO method has a capable agent plan responses, then stitch together text sampled from a base language model. They say the resulting text is unwatermarked even if the agent is watermarked.
TLDR
A researcher introducing the HALO paper says the method separates an agent’s reasoning and planning from the generation of its final text. The team reports over 90% task accuracy, with Pangram and watermark detection down to 10–25%, across tasks including factual grounding and health Q&A. The researcher says the team does not endorse evasion or plan to release a tool to aid it.
Combined views
1.1K
5 Sources, first seen 2h ago
