Model loading reportedly accounts for 55–70% of cold-start latency for small quantized LLMs on serverless CPUs · Digg