MOE Ultimate Guide to LocalLLM Workhorses and Assistants. LLMaxxing on Mini-Bucks < $500 - $1500, plus going higher into the DGX Spark. We follow up on the current trends of how to get capable and productive LLM's on a limited budget.
MTP World First! The Tom Pulls TurboQuant w/MTP (and it Works!) A world first! TurboQuant + MTP support from the same LLama.cpp! What a game changer!
Qwen3.5 Qwen3.5-122B-A10B-Q4_K_M.gguf - Run it at 13 Tokens/s with 262,000 Contexts on a Ryzen 9 3900 and a 4080ti. w/128GB RAM. Qwen3.5-122B-A10B-Q4_K_M.gguf - Run it at 13 Tokens/s with 262,000 Contexts on a Ryzen 9 3900 and a 4080ti. w/128GB RAM.