Guardrails removed spam, off-topic, unclear, or duplicate replies.
Ask a question below.
Published answers will appear here.
A task is only useful if it produces a clean reward signal: gold patch → tests pass, no patch → tests fail. A surprising amount of open task data fails this. We ran gold-patch and no-op validation on every default dataset - multiple passes, 10x retries to separate flaky from broken.
RL agents share a live sandbox with the grading machinery, and anything readable is fair game for a reward hack. Our integrations withhold test patches, expected outputs, and grading scripts until scoring time - even where the original images ship them readable.
The cleaned datasets are re-uploaded with every exclusion and the generation scripts preserved - fully auditable, fully reproducible. All of it lives in our SWE RL collection on Hugging Face.
Guardrails removed spam, off-topic, unclear, or duplicate replies.
Ask a question below.
Published answers will appear here.