Simon Willison:寻找内存低于 60GB 的开放权重编程 MoE 模型
中文全文 · AI 翻译
能在不到 60GB 内存中运行、最适合编程的开放权重混合专家(MoE)大语言模型是什么?
我觉得,在我能使用的硬件上,要获得还算流畅的交互速度,可能必须用 MoE——我想要比每秒 12 个 token 更快的速度。
我喜欢用 Qwen3.5-35B-A3B 来做这件事,想知道还有没有更新一些、能力也更强一些的选择。
对照原文
What's the best open weight Mixture-of-Experts LLM for coding that fits in less than 60GB of RAM? I think MoE might be necessary to get reasonably interactive speeds on the hardware I have access to - I want something faster than 12 tokens/second
I like Qwen3.5-35B-A3B for this, wondering what else is out there that might be a bit more recent and pack a little more of a punch