AI homelab user estimates their setup is 60–70% local
A user says their homegrown routing plugin decides when to use local AI or switch to cloud models. One hardware choice they wouldn't recommend: the external GPU, which they regret buying.
TLDR
One user describes an AI homelab that defaults to local models, estimating it is 60–70% local with a goal of reaching 100%. A homegrown plugin uses Arch-Router to decide when to switch to cloud models. They report Qwen 3.8 27B sometimes reaching 150-plus tokens per second on a 5090 external GPU. Two DGX Spark machines run Deepseek v4 Flash 0731 and handle background batch jobs and longer development builds. Their setup also includes a Mac mini for development and a Pi 5 for monitoring. The user says they started with just the Mac mini, stress that all this hardware isn't necessary, and regret buying the external GPU.
Combined views
30.5K
1 Source, first seen 17d ago