• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Qwen 3.8 flash and the case for laptop-based AI

    A user running the model on a Strix Halo laptop praises llama.cpp's progress and thinks local AI is close to handling a large share of tasks.

    vitalik.ethVI
    1 Source, 21d ago, first seen 21d ago

    TLDR

    A user calls Qwen 3.8 flash impressive on their Strix Halo laptop and says llama.cpp is rapidly improving at processing it. They think local models are nearing two practical uses: handling a large share of tasks locally, and coordinating queries to more powerful models for advanced work without leaking personal information.

    Combined views

    278.7K

    1 Source, first seen 21d ago

    Combined views

    278.7K

    1 Source, first seen 21d ago

    1.9K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1.9K likes
    237 comments
    522 saves
    175 reposts
    237 comments
    522 saves
    175 reposts

    1 Source

    vitalik.eth@VitalikButerinQwen 3.8 flash is truly impressive, and llama.cpp has been rapidly getting better and better at processing it columns are: pre-existing prompt, new prompt, generated, input tok/s, output tok/s This is on my laptop (strix halo). I think we're very close to the point where you can just use local models for a large share of tasks, and for anything more advanced, workflows like "use your local model to orchestrate queries to powerful models so your queries don't leak your personal information" actually become viable.21d

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    vitalik.eth@VitalikButerinQwen 3.8 flash is truly impressive, and llama.cpp has been rapidly getting better and better at processing it columns are: pre-existing prompt, new prompt, generated, input tok/s, output tok/s This is on my laptop (strix halo). I think we're very close to the point where you can just use local models for a large share of tasks, and for anything more advanced, workflows like "use your local model to orchestrate queries to powerful models so your queries don't leak your personal information" actually become viable.21d