Report
Alternating prompt updates and model retraining reportedly lifted a science agent’s accuracy from 42.2% to 73.3%
A post about ScienceBuddy says it turns user corrections into scored test tasks, then alternates prompt and skill changes with model retraining.
TLDR
A post sharing the ScienceBuddy research paper says alternating improvements to prompts, skills and the model raised a science agent’s accuracy from 42.2% to 73.3%. On biology tasks with a small 4B model, prompt and skill changes alone raised accuracy from 31.1% to 51.1%. Retraining alone raised the share of problems solved within four tries from 48.3% to 67.8%, according to the post.
Combined views
2.9K
1 Source, first seen ago
