xLLM promises flexible LLM training without giving up throughput
IFM says xLLM lets teams change tokenizers, data mixtures, model architectures and training stages without rebuilding the dataset or the system around them.
TLDR
IFM introduced xLLM for pre-training and fine-tuning dense and mixture-of-experts LLMs. It reports 6,295 tokens per second per GPU on K2-Horizon-MoVA-36B-A4B and 10,050 on Llama3-8B, both on H200s. IFM says xLLM also ships with K2 Horizon checkpoints, training logs and recipes.
Combined views
24.5K
6 Sources, first seen 7h ago
likes
