Reactions from ranked influencers
7 postsIn the Frontend Code Arena, Hy3 is #2 for open-weight models! -#16 overall - Top 20 in Reference-Based Design, Gaming, Simulations and Content Creation Tools
In agent arena overall, Hy3 is #25 with -2.2% net improvement based on 8K+ live agentic sessions. -4.6% Confirmed Success - 2.9% Praise vs Complaint - 7.1% Steerability +2.6% Bash Recovery + 0.9% Tool Hallucionation
Hy3 by Tencent is #5 in Agent Arena for open-weight models (#25 overall)! It also ranks as the #2 open model in the Frontend Code Arena (#16 overall)! In Agent Arena: Hy3 lands at #25 overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #25), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% (#30). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: User-satisfaction proxies - Confirmed Success : an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Congrats to the @TencentHunyuan team on this release!
🚀Hy3 is here. 295B MoE. Best in its size class. Rivals trillion-scale flagships. Reliable and affordable for most agentic usecases. Apache 2.0. Friendly for commercial use. FREE API for 2 weeks → https://openrouter.ai/tencent/hy3:free 🤗 https://huggingface.co/tencent/Hy3 📖 https://hy.tencent.com/research/hy3
Another strong Chinese OSS model for Agent workflows today.
Hy3 by Tencent is #5 in Agent Arena for open-weight models (#25 overall)! It also ranks as the #2 open model in the Frontend Code Arena (#16 overall)! In Agent Arena: Hy3 lands at #25 overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #25), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% (#30). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Congrats to the @TencentHunyuan team on this release!
In agent arena overall, Hy3 is #25 with -2.2% net improvement based on 8K+ live agentic sessions. -4.6% Confirmed Success - 2.9% Praise vs Complaint - 7.1% Steerability +2.6% Bash Recovery + 0.9% Tool Hallucination
Hy3 by Tencent is #5 in Agent Arena for open-weight models (#25 overall)! It also ranks as the #2 open model in the Frontend Code Arena (#16 overall)! In Agent Arena: Hy3 lands at #25 overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #25), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% (#30). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Congrats to the @TencentHunyuan team on this release!
🚀Hy3 is here. 295B MoE. Best in its size class. Rivals trillion-scale flagships. Reliable and affordable for most agentic usecases. Apache 2.0. Friendly for commercial use. FREE API for 2 weeks → https://openrouter.ai/tencent/hy3:free 🤗 https://huggingface.co/tencent/Hy3 📖 https://hy.tencent.com/research/hy3
Combined views
21.4K
7 posts, first seen 4h ago