democratizing high quality evals & environments for your specific use-cases is a huge win for Open Intelligence.
so we’re sharing skills & workflows that plug right into your coding agent to help everyone own this process using their Trace data using the open @harborframework
A design principle we adopt is: The best way to make good evals is to interview the user about what the agent should be good at!
Evals are training data for agents. The goal of evals is to transfer knowledge and specialization encoded in Tasks into Agent behavior. A lot of that knowledge lives in human minds, and interviewing them and iteratively testing evals is an effective way to drive that transfer.
Creating evals is hard. People are unfamiliar with the tooling and goals are often underspecified in a first pass. Agent optimization is iterative because you don’t know what agents will do until you run them. The same holds for evals which is why we find that iteratively refining them is usually much better than one-shot generation.
A large part of our research mission towards Continual Learning revolves around large-scale data mining and good eval creation. Every team should be able to use their data to continuously improve their agents
looking forward to getting community feedback on this flow, refining it, releasing more data, and helping more users own their intelligence :)
try it here https://github.com/langchain-ai/langchain-skills/tree/main/config/skills/eval-engineering