• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Apodex Shares TRACES Platform for Heavy-Duty AI Solvers

    Rohan Paul posted about Apodex TRACES for submitting AI problems and solvers.

    EL
    RP
    DA
    23 Sources, 29d ago, first seen 29d ago

    TLDR

    Rohan Paul, a Bengaluru machine learning engineer, posted a reply highlighting Apodex and its TRACES service. The post links to a technical report at apodex.com/pdf/20260820 and to traces.apodex.com pages. Those pages describe TRACES as a way to turn real-world problems into executable environments for AI, with routes to verification. Solvers can be submitted to run in those environments and receive profiles covering outcome, trace, speed, and cost. Apodex positions itself as handling the hardest problems that lack existing answers through deep research and step-by-step checks.

    Combined views

    205.7K

    23 Sources, first seen 29d ago

    Combined views

    205.7K

    23 Sources, first seen 29d ago

    246 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    246 likes
    62 comments
    134 saves
    71 reposts
    62 comments
    134 saves
    71 reposts

    Sentiment

    Positive96%4%Negative

    Summary

    Sentiment

    Positive96%4%Negative

    Many accounts welcomed Apodex 1.1 and its TRACES benchmark because it scores AI agent reasoning on unconfirmed discoveries and shows strong performance from smaller models.

    Based on 25 sentiment-bearing replies from 25 accounts across 2 conversations.

    Summary

    Many accounts welcomed Apodex 1.1 and its TRACES benchmark because it scores AI agent reasoning on unconfirmed discoveries and shows strong performance from smaller models.

    Based on 25 sentiment-bearing replies from 25 accounts across 2 conversations.

    23 Sources

    @rohanpaul_aiApodex released Apodex 1.1, its proprietary model that reached 44 on the Artificial Analysis Intelligence Index and performs strongly on agentic tasks versus models in its tier. on GDPval-AA v2, which measures real-world agentic work, it reaches 1,348 Elo, ahead of several models with much higher general intelligence scores. So Apodex 1.1, its performance seems concentrated around professional and agentic tasks rather than being evenly distributed across the evaluation suite. I increasingly think this distinction matters for model selection. A model that wins broad reasoning benchmarks is not automatically the model you want sitting inside an agentic execution loop. For agents, the relevant question is: once you give the model a goal and tools, how often does it actually get the job done? Apodex 1.1 from @Apodex_AI looks unusually concentrated in that direction.
    @Apodex_AIhttps://x.com/i/article/2079399139453132800
    @omarsar0New benchmark from @Apodex_AI worth digging into. TRACES scores AI systems on discoveries where the answer isn't confirmed yet. This is a crucial capability to measure in agents used for real research problems. Here's the breakdown:
    @dair_aiVery important paper accelerating research around AI discovery.

    23 Sources

    @rohanpaul_aiApodex released Apodex 1.1, its proprietary model that reached 44 on the Artificial Analysis Intelligence Index and performs strongly on agentic tasks versus models in its tier. on GDPval-AA v2, which measures real-world agentic work, it reaches 1,348 Elo, ahead of several models with much higher general intelligence scores. So Apodex 1.1, its performance seems concentrated around professional and agentic tasks rather than being evenly distributed across the evaluation suite. I increasingly think this distinction matters for model selection. A model that wins broad reasoning benchmarks is not automatically the model you want sitting inside an agentic execution loop. For agents, the relevant question is: once you give the model a goal and tools, how often does it actually get the job done? Apodex 1.1 from @Apodex_AI looks unusually concentrated in that direction.
    @Apodex_AIhttps://x.com/i/article/2079399139453132800
    @omarsar0New benchmark from @Apodex_AI worth digging into. TRACES scores AI systems on discoveries where the answer isn't confirmed yet. This is a crucial capability to measure in agents used for real research problems. Here's the breakdown:
    @dair_aiVery important paper accelerating research around AI discovery.