• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    DeepSeek V4.1 Flash’s coding comparisons and ARC-AGI 2 prospects

    A commenter thinks DeepSeek V4.1 Flash could do well on ARC-AGI 2 puzzles, citing a reported 61% result for Flash-0731 while also flagging regressions on CritPt and other tests.

    T(
    3 Sources, 15d ago, first seen 15d ago

    TLDR

    A commenter endorses a quoted assessment placing DeepSeek V4.1 Flash between GPT 5.6 Terra and Sol in subjective agentic coding performance and Catan play, while costing much less than Terra. They want to see ARC-AGI 2 results and argue the model has potential on those puzzles. In a follow-up, they cite Flash-0731’s claimed 61% result, but temper their optimism by noting regressions on CritPt and other tests.

    Combined views

    43.5K

    3 Sources, first seen 15d ago

    Combined views

    43.5K

    3 Sources, first seen 15d ago

    499 likes
    499 likes
    30 comments
    53 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    30 comments
    53 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @teortaxesTexit really has the makings of a model that'd do great on ARC-2 puzzles. Flash-0731 already got to 61%, and it had a much more rudimentary non-native multimodality. Then again, they had a strange regression on CritPt and some others. @arcprize

    3 Sources

    @teortaxesTexit really has the makings of a model that'd do great on ARC-2 puzzles. Flash-0731 already got to 61%, and it had a much more rudimentary non-native multimodality. Then again, they had a strange regression on CritPt and some others. @arcprize