• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    LeanLean is a new benchmark for compressing Lean codebases

    Its creators say Opus 5.5 leads the leaderboard with a 64.3% score, versus 39.9% for GPT 6.1 Sol.

    Christian SzegedyCS
    Kári RögnvaldssonKR
    2 Sources, 9h ago, first seen 9h ago

    TLDR

    LeanLean’s creators introduced the benchmark as LLM-written proofs grow longer, citing what they describe as Claude’s 13-million-line Lean proof of Fermat’s Last Theorem. They report a 64.3% leaderboard score for Opus 5.5, compared with 39.9% for GPT 6.1 Sol.

    Combined views

    3.1K

    2 Sources, first seen 9h ago

    Combined views

    3.1K

    2 Sources, first seen 9h ago

    65 likes
    65 likes
    5 comments
    11 saves
    17 reposts
    Featured Source
    5 comments
    11 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Kári Rögnvaldsson@kariroggNew paper! 🥳 LLM-written proofs are getting long (Claude's proof of Fermat's Last Theorem is 13M lines of Lean). We introduce LeanLean, a benchmark for compressing Lean codebases. Opus 5.5 dominates the leaderboard with a score of 64.3%, while GPT 6.1 Sol reaches only 39.9%.9h
    Christian Szegedy@ChrSzegedyRT @karirogg: New paper! 🥳 LLM-written proofs are getting long (Claude's proof of Fermat's Last Theorem is 13M lines of Lean). We introduc…1h

    2 Sources

    Kári Rögnvaldsson@kariroggNew paper! 🥳 LLM-written proofs are getting long (Claude's proof of Fermat's Last Theorem is 13M lines of Lean). We introduce LeanLean, a benchmark for compressing Lean codebases. Opus 5.5 dominates the leaderboard with a score of 64.3%, while GPT 6.1 Sol reaches only 39.9%.9h
    Christian Szegedy@ChrSzegedyRT @karirogg: New paper! 🥳 LLM-written proofs are getting long (Claude's proof of Fermat's Last Theorem is 13M lines of Lean). We introduc…1h