• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Andrew Curran Shares Artificial Analysis Coding Agent Index v1.4

    Reply includes a line chart plotting model scores against API costs.

    EM
    AC
    TI
    16 Sources, 27d ago, first seen 27d ago

    TLDR

    Andrew Curran an independent writer and commentator focused on AI posted a reply on X. The post contains an attachment showing a dark-themed line chart. The chart is titled Artificial Analysis Coding Agent Index v1.4 and features an OpenAI spiral logo in the top right. It displays Score measured in index points on the vertical axis against API cost measured in USD on the horizontal axis. The visible source line describes the chart layout and axes but provides no additional performance data or commentary from the author beyond the attachment itself.

    Combined views

    1.2M

    16 Sources, first seen 27d ago

    Combined views

    1.2M

    16 Sources, first seen 27d ago

    10.6K likes
    10.6K likes
    576 comments
    739 saves
    412 reposts
    576 comments
    739 saves
    412 reposts

    Sentiment

    Positive48.2%51.8%Negative

    Summary

    Sentiment

    Positive48.2%51.8%Negative

    Positive replies highlighted Astra's efficiency and cost-per-accuracy gains in GPT-6 benchmarks, while negative replies criticized high token use, lost reasoning traces, and questionable data.

    Based on 57 sentiment-bearing replies from 56 accounts across 4 conversations.

    Summary

    Positive replies highlighted Astra's efficiency and cost-per-accuracy gains in GPT-6 benchmarks, while negative replies criticized high token use, lost reasoning traces, and questionable data.

    Based on 57 sentiment-bearing replies from 56 accounts across 4 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    16 Sources

    @AndrewCurran_GPT-6 benchmarks:
    @xeophonwhy are they reporting some scores as api cost and some in tokens ahhhhh
    @andrew_n_carrLove you guys, but what is this graph? Some use tokens and others use api cost?
    @ArtificialAnlysGPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in output tokens at max effort compared to GPT-5.6 Sol.
    @EMostaqueNext year frontier models are going to one shot everything 100x faster & cheaper than they are No reasoning traces to monitor Instinctive intelligence Which is why what you embed in pretraining is even more important
    @stevenheidelalso "tokens" isn't a standard unit of measurement across providers, or even models from the same provider. Anthropic snuck in a 30% cost increase to Opus 4.7 just by making their tokenizer less efficient
    @zacharynado⚡⚡⚡
    @thsottiauxAwesome to see GPT-6 Astra is #1 on Terminal Bench 4.0 using the Codex harness. At 50% of the cost of #2.
    @reach_vbThis! Astra provides much higher performance at lower reasoning efforts compared to Sol (see attached) low/ medium is a comfy default!

    16 Sources

    @AndrewCurran_GPT-6 benchmarks:
    @xeophonwhy are they reporting some scores as api cost and some in tokens ahhhhh
    @andrew_n_carrLove you guys, but what is this graph? Some use tokens and others use api cost?
    @ArtificialAnlysGPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in output tokens at max effort compared to GPT-5.6 Sol.
    @EMostaqueNext year frontier models are going to one shot everything 100x faster & cheaper than they are No reasoning traces to monitor Instinctive intelligence Which is why what you embed in pretraining is even more important
    @stevenheidelalso "tokens" isn't a standard unit of measurement across providers, or even models from the same provider. Anthropic snuck in a 30% cost increase to Opus 4.7 just by making their tokenizer less efficient
    @zacharynado⚡⚡⚡
    @thsottiauxAwesome to see GPT-6 Astra is #1 on Terminal Bench 4.0 using the Codex harness. At 50% of the cost of #2.
    @reach_vbThis! Astra provides much higher performance at lower reasoning efforts compared to Sol (see attached) low/ medium is a comfy default!