• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    Every seeks someone to lead an AI model benchmarking product

    Every’s head of evals says they’re working with 15 people internally to build their own personal benchmarks.

    Lenny RachitskyLR
    Dan ShipperDS
    Mike TaylorMT
    5 Sources, ,

    TLDR

    Every’s head of evals says CEO Dan Shipper asked for a way to quantify “vibe checks” of new frontier AI models. They’re working with 15 people internally to build personal benchmarks, and the new hire would own the product’s direction. The post describes an ideal candidate as someone who values both quantitative and qualitative judgment.

    Combined views

    10.7K

    5 Sources, first seen 3h ago

    Combined views

    10.7K

    5 Sources, first seen 3h ago

    71 likes
    3h ago
    first seen 3h ago
    71 likes
    10 comments
    24 saves
    17 reposts
    Featured Source
    10 comments
    24 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    Mike Taylor@hammer_mtCome work with us and own this product! This idea came out of @danshipper 's brief to me to find a way to quantify our vibe checks of new models from the frontier labs. I'm working with 15 people so far internally to build their own personal benchmark including @ArielleShipper who runs ops and was our guinea pig for improving onboarding and @kieranklaassen who created compound engineering and spends billions of tokens per day! You'll work with @nityeshaga who came over to evals with me from the consulting team, where he built workflows and skills for hedge funds, and before that was an early engineer with Kieran on Cora. You also have support from @JannikJung who helped rescue our vibe coded markdown editor product then automated our internal AI editor KatePass, a clone of Kate Lee our editor in chief. On design you have @TylerNishida who helped design much of our new @every agent product and had his projects highlighted by the Anthropic team in a new model release. You'd be working with @bigwilliestyle our head of platform, who just launched the Every agent and previously ran AI operations at Lyft. I'll be helping you push things forward with the crazy experiments I'm running as head of Evals, but you'll own the product direction as this turns into something real. Dan is personally very invested in this product and so you get to work a lot with a CEO that actually dogfoods the products that we build (and is spending thousands of dollars in tokens per week himself). If you have run your own company before (or plan to) this would be a great role for you. Ideally you're a quant that hasn't lost sight of the value of qual (or vice versa). Working at Every is an excuse to get paid to do the fun stuff with AI you would do anyway for free.3h
    Dan Shipper@danshipperRT @hammer_mt: Come work with us and own this product! This idea came out of @danshipper 's brief to me to find a way to quantify our vibe…3h
    yash poojary@poojary_yashhelp us checkmate AI slop with Checks and work with most ai pilled team on the planet @hammer_mt @nityeshaga @JannikJung2h
    Lenny Rachitsky@lennysanRad PM job alert2h

    5 Sources

    Mike Taylor@hammer_mtCome work with us and own this product! This idea came out of @danshipper 's brief to me to find a way to quantify our vibe checks of new models from the frontier labs. I'm working with 15 people so far internally to build their own personal benchmark including @ArielleShipper who runs ops and was our guinea pig for improving onboarding and @kieranklaassen who created compound engineering and spends billions of tokens per day! You'll work with @nityeshaga who came over to evals with me from the consulting team, where he built workflows and skills for hedge funds, and before that was an early engineer with Kieran on Cora. You also have support from @JannikJung who helped rescue our vibe coded markdown editor product then automated our internal AI editor KatePass, a clone of Kate Lee our editor in chief. On design you have @TylerNishida who helped design much of our new @every agent product and had his projects highlighted by the Anthropic team in a new model release. You'd be working with @bigwilliestyle our head of platform, who just launched the Every agent and previously ran AI operations at Lyft. I'll be helping you push things forward with the crazy experiments I'm running as head of Evals, but you'll own the product direction as this turns into something real. Dan is personally very invested in this product and so you get to work a lot with a CEO that actually dogfoods the products that we build (and is spending thousands of dollars in tokens per week himself). If you have run your own company before (or plan to) this would be a great role for you. Ideally you're a quant that hasn't lost sight of the value of qual (or vice versa). Working at Every is an excuse to get paid to do the fun stuff with AI you would do anyway for free.3h
    Dan Shipper@danshipperRT @hammer_mt: Come work with us and own this product! This idea came out of @danshipper 's brief to me to find a way to quantify our vibe…3h
    yash poojary@poojary_yashhelp us checkmate AI slop with Checks and work with most ai pilled team on the planet @hammer_mt @nityeshaga @JannikJung2h
    Lenny Rachitsky@lennysanRad PM job alert2h