• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Claude’s developer account announces plugin evaluations for Claude Code

    The account says `claude plugin eval` lets users create test cases, run and score a plugin or skill against them, then rerun each case without the plugin to compare results.

    EL
    TH
    HH
    3 Sources, ,

    TLDR

    Claude’s developer account introduced claude plugin eval to help users assess what a plugin adds. A post sharing the announcement says feedback highlighted how hard it is to know whether skills still work with new model releases, and presents plugin evaluations as a way to help. It directs users to run claude plugin eval init in their plugin folder.

    Combined views

    295.4K

    3 Sources, first seen 19d ago

    Combined views

    295.4K

    3 Sources, first seen 19d ago

    2.3K likes
    19d ago
    first seen 19d ago
    2.3K likes
    118 comments
    2.1K saves
    137 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    118 comments
    2.1K saves
    137 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @trq212we heard feedback that it's hard to know if your skills are still working with new model releases plugin evals are here to help run `claude plugin eval init` in your plugin folder
    @HamelHusainIt's cool that there is more interest in eval tools! Some opportunities for improvement: 1. Right now, this workflow tries to quiz you up front about your recollection of your experience with a plugin. It would be better if it was more "in-situ", meaning you could give feedback on the plugins as you are using them. 2. It goes off and builds datasets and judges automatically based on what you tell it in up front as well as what's documented in the plugin. I would like to see it try to do error discovery with you first to allow you to annotate real traces / session history etc so you can figure out what's not working better. 3. It was clunky to do this against a skill that wasn't a plugin. I had to ask claude to set it up for me and it took 10 minutes to figure that out. 4. I had to stare at the generated HTML for a while before I could understand what everything meant. There is some onboarding experience that is missing but I'm not sure what that is yet. I'm sure this will get better over time and the things above seem doable. It's also positive to see affordances for evals directly in our agents!
    @omarsar0This is interesting. Any thoughts on open-sourcing it? It feels like something that could be more useful if it were hackable/customizable. Or is it just built to work with Claude Code and Claude models? I have been developing my own skills for this, but with new models, it's hard to stay up to date and be effective across my tasks. I think broader community knowledge/input could be useful for an evals plugin. Not sure; just thinking out loud, as this is something I do work on quite often in my own meta harness.

    3 Sources

    @trq212we heard feedback that it's hard to know if your skills are still working with new model releases plugin evals are here to help run `claude plugin eval init` in your plugin folder
    @HamelHusainIt's cool that there is more interest in eval tools! Some opportunities for improvement: 1. Right now, this workflow tries to quiz you up front about your recollection of your experience with a plugin. It would be better if it was more "in-situ", meaning you could give feedback on the plugins as you are using them. 2. It goes off and builds datasets and judges automatically based on what you tell it in up front as well as what's documented in the plugin. I would like to see it try to do error discovery with you first to allow you to annotate real traces / session history etc so you can figure out what's not working better. 3. It was clunky to do this against a skill that wasn't a plugin. I had to ask claude to set it up for me and it took 10 minutes to figure that out. 4. I had to stare at the generated HTML for a while before I could understand what everything meant. There is some onboarding experience that is missing but I'm not sure what that is yet. I'm sure this will get better over time and the things above seem doable. It's also positive to see affordances for evals directly in our agents!
    @omarsar0This is interesting. Any thoughts on open-sourcing it? It feels like something that could be more useful if it were hackable/customizable. Or is it just built to work with Claude Code and Claude models? I have been developing my own skills for this, but with new models, it's hard to stay up to date and be effective across my tasks. I think broader community knowledge/input could be useful for an evals plugin. Not sure; just thinking out loud, as this is something I do work on quite often in my own meta harness.