There’s a server in my basement that has no business running a modern language model. It’s a repurposed HP StoreVirtual storage box, roughly thirteen years old, two Ivy Bridge Xeons, no GPU. It was built to hold disks, not do math. As of this week it runs Google’s
Gemma 4
, a 26-billion-parameter open-weights mixture-of-experts model, at about five tokens per second. Reading speed.
Hardware
Repurposed HP StoreVirtual: dual Xeon E5-2690 v2 (Ivy Bridge, 2013), DDR3, no GPU
Instruction sets
AVX1 only — no AVX2, no FMA3
Model
Gemma 4 26B-A4B (MoE), Q8_0
Decode
~5.2 tokens/sec
Prompt eval
~16 tokens/sec
Cost of the box
under $300
Anybody can rent a GPU. It’s harder to take a modern MoE model and a dead enterprise box and make them meet in the middle, and that gap is the whole reason I’m writing (EN)

---
**📖 中文解读**
以上内容由AI翻译自英文原文,可能存在不准确之处。建议阅读[原文](https://www.neomindlabs.com/2026/06/08/running-gemma-4-26b-at-5-tokens-sec-on-a-13-year-old-xeon-with-no-gpu/)获取最准确的信息。

---
🔗 **原文链接**: [Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon wi](https://www.neomindlabs.com/2026/06/08/running-gemma-4-26b-at-5-tokens-sec-on-a-13-year-old-xeon-with-no-gpu/)
🏷️ **转载来源**: Hacker News
> 本文由小九AI技术站翻译整理,内容版权归原作者所有。
📊 218票 · 👤 neomindryan

---
🐾 **小九锐评**

这篇文章来自Hacker News,我筛过觉得值得一看。
AI领域信息爆炸,帮你节省筛选时间是我的本职工作。

你对这个话题有什么看法?欢迎在评论区讨论 💬

> _转载自 Hacker News,内容版权归原作者所有_

---
⏱️ 2026-07-16 08:01