• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Zapier Founders Announce Astra's AutomationBench Record

    Zapier founders posted that the model topped their internal benchmark.

    MK
    BT
    WF
    4 Sources, 27d ago, first seen 27d ago

    TLDR

    Wade Foster stated that GPT 6 Astra reached the highest score recorded on AutomationBench. Bryan Helmig described the test as running models inside simulated apps, assigning tasks, and checking resulting records and messages with code. The posts note a clean sweep across domains and list suitable uses such as reconciliation, deal review prep, and vendor scorecards. Earlier models had not cleared 40 percent on the same benchmark.

    Combined views

    21.1K

    4 Sources, first seen 27d ago

    Combined views

    21.1K

    4 Sources, first seen 27d ago

    224 likes
    224 likes
    11 comments
    59 saves
    39 reposts

    Sentiment

    Positiveโ€”โ€”Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    59 saves
    39 reposts

    Sentiment

    Positiveโ€”โ€”Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    โ€”

    Not ranked yet

    Today's Rank

    โ€”

    Not ranked yet

    4 Sources

    @wadefosterGPT 6 Astra is here. We ran the numbers on AutomationBench: It's the highest score we've ever recorded. Clean sweep across every domain. Scores 41.4% at Max effort. For context, no model had cleared 40% before today (GPT-5.6-Sol scored 28.8%) ๐—•๐—ฒ๐˜€๐˜ ๐—ณ๐—ถ๐˜ ๐—ณ๐—ผ๐—ฟ: reconciliation, deal review prep, vendor scorecards, anything where touching the wrong record is expensive. ๐—ช๐—ฒ๐—ฎ๐—ธ๐—ฒ๐—ฟ ๐—ณ๐—ผ๐—ฟ: outbound comms where the guidance is scattered. Operations and support are its strongest domains. HR is its weakest, same as every model we test (still the new high score, though) Its edge is arithmetic across messy sources. Finding the policy doc, the logged correction, the exception rule, etc. Example 1: rebalance a quarterly media budget from last quarter's actuals, with finance adjustments and channel eligibility rules buried in email. Both models produced a budget and landed on the same total. Astra found the adjustments, so every per-channel number was right. Sol's looked finished and had the splits wrong. Example 2: answer and log 15 integration inquiries using a reply standard stored in a doc. Astra searched, could not find the standard, and stopped. Zero replies sent. Sol did not find it either, took its best shot at all 15, and earned partial credit. Those examples highlight how these two models make tradeoffs... Astra will not guess. When the instructions exist and it can find them, it finishes the whole job. When it cannot, it pauses the work instead of improvising. Crazy week for LLM releases after a few quiet ones. Astra isn't available to the public yet, but should be soon. We run every new model through @Zapier's AutomationBench, 657 of the hardest workflows we have, across finance, HR, marketing, operations, sales, and support. See every model and every score here: http://zapier.com/benchmarks
    @bentossellRT @wadefoster: GPT 6 Astra is here. We ran the numbers on AutomationBench: It's the highest score we've ever recorded. Clean sweep acrossโ€ฆ
    @bryanhelmigokay, so astra just set a new high on automationbench: 41.4% of business workflows completed correctly. so how is it better? and what is automationbench? well, we measure that by putting the model in simulated apps, giving it a job, and checking the resulting records and messages with code. the hard part is building a test that catches "updated the wrong customer" or "ignored the hold" while still letting the agent figure out its own way to do the work. more or less, you build an small environment for an agent. start with a known set of crm records, emails, spreadsheet rows, calendar events, etc. give the agent tools that behave like those apps' apis. underneath, it's structured state: sending an email adds a message to the sent mailbox, updating a contact changes a field on that record. nothing goes to an actual customer. each run starts with its own copy of the environment, so one model's changes don't affect another's. then you give it a task. say the spreadsheet has verified job titles and you want the crm brought up to date. update the contacts, mark the spreadsheet rows reconciled, and send the summary to the audit inbox specified in the company's procedure. the agent has to find the procedure and the data itself. it gets a keyword search over api documentation and a tool for making requests with a method, url, and body. there's a budget of 50 turns, and it can make several tool calls in a turn. you can make this quite hard without changing that basic request. put two people with the same name in different accounts. leave an outdated title in one source. mark a row verified, then put a later message in the inbox saying that title is disputed and shouldn't be changed yet. put another record under a security freeze. now the agent has to work out which records the request actually applies to. "update everyone in the spreadsheet" won't do it. the critical information need to be discoverable, though. if the correct answer depends on a fact we never put in the environment, that's our bug. same if the api can't retrieve the relevant message, or the instructions contradict each other without a way to tell which one governs. making a task difficult is easy. making it difficult for a reason that tells you something about the model takes significantly more work. scoring is ordinary, deterministic code that inspects the final state. for the contact task, the checks are along these lines: the eligible contacts have the expected titles, their spreadsheet rows are marked reconciled, and the audit inbox received the required summary. but you also check that the held contacts kept their original titles and statuses. otherwise a model could update everyone and pass just because the two intended updates happened somewhere in there. that's an actual difference we saw between astra and sol. astra updated two eligible contacts; sol updated four, including two it should have left alone. sol had even retrieved the security-freeze message for one of them before making the change. a grader that only checked whether the requested updates happened would miss that failure. crucially -- you have to check the unwanted changes too. the same applies to messages. checking that the right person received an email isn't enough if the agent also sent it to twenty wrong people. these negative checks are part of the task, not an optional safety score off to the side. every scored check has to pass for the workflow to count as complete. we keep partial credit to help inspect failures, but 41.4% means complete tasks, not the average fraction of steps the model got right. we don't prescribe the whole sequence of calls. the agent might find the data a different way, batch some updates, or correct an earlier mistake. if it leaves the required state behind, that should pass. this also means the benchmark only sees what its checks cover. a deterministic grader can be consistently wrong! we audit the tasks and use separate hint-assisted runs to help check that they're solvable; the hints aren't part of the scored runs. the simulator and the grader both need scrutiny when a model fails. the public task set is available for experimentation. the headline scores use a separate, harder held-out set, so don't mix those numbers. a few model developers (meta, http://z.ai, etc) are reporting public set scores which aren't verified by us. and none of these scores tell you that the model will succeed on that percentage of work in your company. they tell you how often it completed this particular set of workflows under these conditions. we're actively working on automationbench 2, something that will be significantly trickier...
    @mikeknoopRT @bryanhelmig: okay, so astra just set a new high on automationbench: 41.4% of business workflows completed correctly. so how is it betteโ€ฆ

    4 Sources

    @wadefosterGPT 6 Astra is here. We ran the numbers on AutomationBench: It's the highest score we've ever recorded. Clean sweep across every domain. Scores 41.4% at Max effort. For context, no model had cleared 40% before today (GPT-5.6-Sol scored 28.8%) ๐—•๐—ฒ๐˜€๐˜ ๐—ณ๐—ถ๐˜ ๐—ณ๐—ผ๐—ฟ: reconciliation, deal review prep, vendor scorecards, anything where touching the wrong record is expensive. ๐—ช๐—ฒ๐—ฎ๐—ธ๐—ฒ๐—ฟ ๐—ณ๐—ผ๐—ฟ: outbound comms where the guidance is scattered. Operations and support are its strongest domains. HR is its weakest, same as every model we test (still the new high score, though) Its edge is arithmetic across messy sources. Finding the policy doc, the logged correction, the exception rule, etc. Example 1: rebalance a quarterly media budget from last quarter's actuals, with finance adjustments and channel eligibility rules buried in email. Both models produced a budget and landed on the same total. Astra found the adjustments, so every per-channel number was right. Sol's looked finished and had the splits wrong. Example 2: answer and log 15 integration inquiries using a reply standard stored in a doc. Astra searched, could not find the standard, and stopped. Zero replies sent. Sol did not find it either, took its best shot at all 15, and earned partial credit. Those examples highlight how these two models make tradeoffs... Astra will not guess. When the instructions exist and it can find them, it finishes the whole job. When it cannot, it pauses the work instead of improvising. Crazy week for LLM releases after a few quiet ones. Astra isn't available to the public yet, but should be soon. We run every new model through @Zapier's AutomationBench, 657 of the hardest workflows we have, across finance, HR, marketing, operations, sales, and support. See every model and every score here: http://zapier.com/benchmarks
    @bentossellRT @wadefoster: GPT 6 Astra is here. We ran the numbers on AutomationBench: It's the highest score we've ever recorded. Clean sweep acrossโ€ฆ
    @bryanhelmigokay, so astra just set a new high on automationbench: 41.4% of business workflows completed correctly. so how is it better? and what is automationbench? well, we measure that by putting the model in simulated apps, giving it a job, and checking the resulting records and messages with code. the hard part is building a test that catches "updated the wrong customer" or "ignored the hold" while still letting the agent figure out its own way to do the work. more or less, you build an small environment for an agent. start with a known set of crm records, emails, spreadsheet rows, calendar events, etc. give the agent tools that behave like those apps' apis. underneath, it's structured state: sending an email adds a message to the sent mailbox, updating a contact changes a field on that record. nothing goes to an actual customer. each run starts with its own copy of the environment, so one model's changes don't affect another's. then you give it a task. say the spreadsheet has verified job titles and you want the crm brought up to date. update the contacts, mark the spreadsheet rows reconciled, and send the summary to the audit inbox specified in the company's procedure. the agent has to find the procedure and the data itself. it gets a keyword search over api documentation and a tool for making requests with a method, url, and body. there's a budget of 50 turns, and it can make several tool calls in a turn. you can make this quite hard without changing that basic request. put two people with the same name in different accounts. leave an outdated title in one source. mark a row verified, then put a later message in the inbox saying that title is disputed and shouldn't be changed yet. put another record under a security freeze. now the agent has to work out which records the request actually applies to. "update everyone in the spreadsheet" won't do it. the critical information need to be discoverable, though. if the correct answer depends on a fact we never put in the environment, that's our bug. same if the api can't retrieve the relevant message, or the instructions contradict each other without a way to tell which one governs. making a task difficult is easy. making it difficult for a reason that tells you something about the model takes significantly more work. scoring is ordinary, deterministic code that inspects the final state. for the contact task, the checks are along these lines: the eligible contacts have the expected titles, their spreadsheet rows are marked reconciled, and the audit inbox received the required summary. but you also check that the held contacts kept their original titles and statuses. otherwise a model could update everyone and pass just because the two intended updates happened somewhere in there. that's an actual difference we saw between astra and sol. astra updated two eligible contacts; sol updated four, including two it should have left alone. sol had even retrieved the security-freeze message for one of them before making the change. a grader that only checked whether the requested updates happened would miss that failure. crucially -- you have to check the unwanted changes too. the same applies to messages. checking that the right person received an email isn't enough if the agent also sent it to twenty wrong people. these negative checks are part of the task, not an optional safety score off to the side. every scored check has to pass for the workflow to count as complete. we keep partial credit to help inspect failures, but 41.4% means complete tasks, not the average fraction of steps the model got right. we don't prescribe the whole sequence of calls. the agent might find the data a different way, batch some updates, or correct an earlier mistake. if it leaves the required state behind, that should pass. this also means the benchmark only sees what its checks cover. a deterministic grader can be consistently wrong! we audit the tasks and use separate hint-assisted runs to help check that they're solvable; the hints aren't part of the scored runs. the simulator and the grader both need scrutiny when a model fails. the public task set is available for experimentation. the headline scores use a separate, harder held-out set, so don't mix those numbers. a few model developers (meta, http://z.ai, etc) are reporting public set scores which aren't verified by us. and none of these scores tell you that the model will succeed on that percentage of work in your company. they tell you how often it completed this particular set of workflows under these conditions. we're actively working on automationbench 2, something that will be significantly trickier...
    @mikeknoopRT @bryanhelmig: okay, so astra just set a new high on automationbench: 41.4% of business workflows completed correctly. so how is it betteโ€ฆ