• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Every is adding work-based benchmarks to its AI reviews, a team member says

    An Every team member argues that higher benchmark scores say little about performance on real work. The team has built an internal platform for creating personal benchmarks around day-to-day tasks, they say.

    DS
    M🍥
    MT
    14 Sources, ,

    TLDR

    An Every team member says the team is making its AI model “vibe checks” more quantitative. For three years, those long-form reviews have relied on hands-on testing with real work, according to the announcement. Now, an internal platform lets team members create personal benchmarks based on their day-to-day tasks.

    Combined views

    646.3K

    14 Sources, first seen 19d ago

    Combined views

    646.3K

    14 Sources, first seen 19d ago

    5.4K likes
    19d ago
    first seen 19d ago
    5.4K likes
    163 comments
    3.4K saves
    206 reposts

    Sentiment

    Positive85.1%14.9%Negative

    Based on 74 sentiment-bearing replies from 67 accounts across 4 conversations.

    163 comments
    3.4K saves
    206 reposts

    Sentiment

    Positive85.1%14.9%Negative

    Based on 74 sentiment-bearing replies from 67 accounts across 4 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    14 Sources

    @danshipperbetter benchmark scores don't tell you much about how well a model does on your real work that's why for the last 3 years @every we've done vibe checks on new models: long-form reviews based on hands-on testing on each model for our real work now we're doubling down, and getting more quantitative @hammer_mt and @nityeshaga built an internal platform for all of us to make personal benchmarks based on our real day-to-day work our vibe checks on new models used to be vibes only now we're starting to get more quantitative with checks very psyched about this :)
    @hammer_mtI've been testing this and it's wild. It read everything I've ever written and checked it all for 21 different AI writing tells in under a second.
    @MatthewTeschkeInteresting read. It’s tempting to use LLMs for everything, but at scale you need cheaper options. TypeSafe is an interesting approach
    @kplikethebirdOne of the hardest parts of getting AI to write like you is getting it to make the same language decisions you would, word by word and sentence by sentence. The interesting question Jev raises is whether a writing model could use its probabilities while it generates. At each decision point, it could ask: - Which word would this writer be most likely to choose? - Which sentence shape fits their past decisions? - Which version sounds most like them? If Jev can help a writing model choose better language as it goes, that could be a huge level-up for AI-assisted writing.
    @JosephV04824396This is insanely good, its basically a perfect workflow model. Imagine how good it would be calling this thing in the middle of using skills that require LLM evaluation. Your workflow will be WAY more accurate and you can actually get to the point of scripting everything.
    @dotpem@hammer_mt @DSPyOSS and if anybody wants to look/try/use as input for vibes I already made a dspy fork that lets you swap out LLMs in your Signatures for TypeSafe with one decorator: https://github.com/typesafeainate/dspy-typesafeify
    @ryanvogelthis model is actually insane at email classification i tested it on 1500 of my own emails to see how well it works and I am blown away
    @mayferfirst Jev test: pretty good but 90% agreement rate with verified gemini flash workflow after some tuning it would prob get where i need but not a drop in replacement atm

    14 Sources

    @danshipperbetter benchmark scores don't tell you much about how well a model does on your real work that's why for the last 3 years @every we've done vibe checks on new models: long-form reviews based on hands-on testing on each model for our real work now we're doubling down, and getting more quantitative @hammer_mt and @nityeshaga built an internal platform for all of us to make personal benchmarks based on our real day-to-day work our vibe checks on new models used to be vibes only now we're starting to get more quantitative with checks very psyched about this :)
    @hammer_mtI've been testing this and it's wild. It read everything I've ever written and checked it all for 21 different AI writing tells in under a second.
    @MatthewTeschkeInteresting read. It’s tempting to use LLMs for everything, but at scale you need cheaper options. TypeSafe is an interesting approach
    @kplikethebirdOne of the hardest parts of getting AI to write like you is getting it to make the same language decisions you would, word by word and sentence by sentence. The interesting question Jev raises is whether a writing model could use its probabilities while it generates. At each decision point, it could ask: - Which word would this writer be most likely to choose? - Which sentence shape fits their past decisions? - Which version sounds most like them? If Jev can help a writing model choose better language as it goes, that could be a huge level-up for AI-assisted writing.
    @JosephV04824396This is insanely good, its basically a perfect workflow model. Imagine how good it would be calling this thing in the middle of using skills that require LLM evaluation. Your workflow will be WAY more accurate and you can actually get to the point of scripting everything.
    @dotpem@hammer_mt @DSPyOSS and if anybody wants to look/try/use as input for vibes I already made a dspy fork that lets you swap out LLMs in your Signatures for TypeSafe with one decorator: https://github.com/typesafeainate/dspy-typesafeify
    @ryanvogelthis model is actually insane at email classification i tested it on 1500 of my own emails to see how well it works and I am blown away
    @mayferfirst Jev test: pretty good but 90% agreement rate with verified gemini flash workflow after some tuning it would prob get where i need but not a drop in replacement atm