Report
Qwen3-14B reportedly goes from 22.2% to 54.8% on Spider 2.0-SQLite with harness training
A post summarizing HarnessSQL contrasts static-query training with deployed database agents that inspect schemas and revise queries.
TLDR
A post describing the HarnessSQL paper says Qwen3-14B's Spider 2.0-SQLite score rose from 22.2% to 54.8% when trained in the same execution harness used at deployment. It says HarnessSQL keeps only verified teacher trajectories for supervised fine-tuning, then applies execution-reward reinforcement learning. The post also reports Qwen3-8B rose from 15.5% to 45.2%, with both models transferring to BIRD-Interact and LiveSQLBench.
Combined views
1 Source, first seen ago
20 reposts
