• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI homelab user estimates their setup is 60–70% local

    A user says their homegrown routing plugin decides when to use local AI or switch to cloud models. One hardware choice they wouldn't recommend: the external GPU, which they regret buying.

    AC
    1 Source, 17d ago, first seen 17d ago

    TLDR

    One user describes an AI homelab that defaults to local models, estimating it is 60–70% local with a goal of reaching 100%. A homegrown plugin uses Arch-Router to decide when to switch to cloud models. They report Qwen 3.8 27B sometimes reaching 150-plus tokens per second on a 5090 external GPU. Two DGX Spark machines run Deepseek v4 Flash 0731 and handle background batch jobs and longer development builds. Their setup also includes a Mac mini for development and a Pi 5 for monitoring. The user says they started with just the Mac mini, stress that all this hardware isn't necessary, and regret buying the external GPU.

    Combined views

    30.5K

    1 Source, first seen 17d ago

    Combined views

    30.5K

    1 Source, first seen 17d ago

    275 likes
    275 likes
    77 comments
    289 saves
    16 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    77 comments
    289 saves
    16 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @andrewchencurrent homelab setup for local AI experimentation: - hermes box hosted on a Framework Desktop Mainboard AI Max+ 395 - 5090 eGPU running Qwen 3.8 27B for fast tok/s LLM use - sometimes 150+ tok/s - 2x DGX Spark: running Deepseek v4 Flash 0731 - better but slower model - pi 5 for monitoring - Mac mini as a dev box - use Herdr and ohmypi/codex/claude depending on the use case - housed in a 10" DeskPi mini rack (mostly) Hermes is defaulted to local AI but with a homegrown routing plugin hitting a small low TTFT model (Arch-Router) to decide whether to go local or upgrade to cloud/frontier. Trying to get to 100% local over time, but right now probably more like 60-70% The Sparks are for batch processing background runs (all the cron jobs, longer dev builds, etc) do you need all of this? Absolutely not lol. I started with the mac mini and couldn't help myself but to add over time! Also I regret getting the eGPU so I wouldn't recommend that to anyone

    1 Source

    @andrewchencurrent homelab setup for local AI experimentation: - hermes box hosted on a Framework Desktop Mainboard AI Max+ 395 - 5090 eGPU running Qwen 3.8 27B for fast tok/s LLM use - sometimes 150+ tok/s - 2x DGX Spark: running Deepseek v4 Flash 0731 - better but slower model - pi 5 for monitoring - Mac mini as a dev box - use Herdr and ohmypi/codex/claude depending on the use case - housed in a 10" DeskPi mini rack (mostly) Hermes is defaulted to local AI but with a homegrown routing plugin hitting a small low TTFT model (Arch-Router) to decide whether to go local or upgrade to cloud/frontier. Trying to get to 100% local over time, but right now probably more like 60-70% The Sparks are for batch processing background runs (all the cron jobs, longer dev builds, etc) do you need all of this? Absolutely not lol. I started with the mac mini and couldn't help myself but to add over time! Also I regret getting the eGPU so I wouldn't recommend that to anyone