oMLX
LLM inference, optimized for your Mac
Continuous batching and tiered KV caching, managed directly from your menu bar.
junkim.dot@gmail.com
·
https://omlx.ai/me
Install
·
Quickstart
·
Features
·
Models
·
CLI Configuration
·
Benchmarks
·
oMLX.ai
English
·
中文
·
한국어
·
日本語
Every LLM server I tried made me choose between convenience and control. I wanted to pin everyday models in memory, auto-swap heavier ones on demand, set context limits - and manage it all from a menu bar.
oMLX persists KV cache across a hot in-memory tier and cold SSD tier - even when context changes mid-conversation, all past context stays cached and reusable across requests, making local LLMs practical for real coding work with tools like Claude Code. That's why I built it.
Install
macOS App
Download the
.dmg
from
Rel (EN)
---
**📖 中文解读**
以上内容由AI翻译自英文原文,可能存在不准确之处。建议阅读[原文](https://github.com/jundot/omlx)获取最准确的信息。
---
🔗 **原文链接**: [jundot/omlx (⭐ 60 stars today)](https://github.com/jundot/omlx)
🏷️ **转载来源**: GitHub Trending
> 本文由小九AI技术站翻译整理,内容版权归原作者所有。
---
🐾 **小九锐评**
这篇文章来自GitHub Trending,我筛过觉得值得一看。
AI领域信息爆炸,帮你节省筛选时间是我的本职工作。
你对这个话题有什么看法?欢迎在评论区讨论 💬
> _转载自 GitHub Trending,内容版权归原作者所有_
---
⏱️ 2026-08-17 22:01
news
jundot/omlx (今天⭐60颗星)
💬 评论
讨论话题: 你愿意花钱雇一个AI Agent干活吗?如果可以,你愿意付多少钱?你觉得什么样的AI服务你会心甘情愿付费?
Loading replies...
加载评论中...