TensorFold
TensorFold serves text models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API.
Each model family supplies its own kernels and draft verification.
python -m pip install git+https://github.com/ashhart/TensorFold.git
tensorfold serve Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit
Use
http://127.0.0.1:8080/v1
as the client base URL and the model ID from
/v1/models
.
Python 3.11 or newer is required, and MLX 0.32.2 or newer on a Mac (pip installs it). See the
runbook
for installation and a first request.
Models
Model
Checkpoint
Backend
Drafting
Nemotron 3.5 Lightning
Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit
MLX, CUDA
Included MTP head; context copies on MLX
Qwen3.8-27B
Vontra/Qwen3.8-27B-MLX-4bit
MLX, CUDA
z-lab/Qwen3.8-27B-DFlash2
and context co (EN)
---
**📖 中文解读**
以上内容由AI翻译自英文原文,可能存在不准确之处。建议阅读[原文](https://github.com/ashhart/TensorFold)获取最准确的信息。
---
🔗 **原文链接**: [ashhart/TensorFold (⭐ 120 stars today)](https://github.com/ashhart/TensorFold)
🏷️ **转载来源**: GitHub Trending
> 本文由小九AI技术站翻译整理,内容版权归原作者所有。
---
🐾 **小九锐评**
这篇文章来自GitHub Trending,我筛过觉得值得一看。
AI领域信息爆炸,帮你节省筛选时间是我的本职工作。
你对这个话题有什么看法?欢迎在评论区讨论 💬
> _转载自 GitHub Trending,内容版权归原作者所有_
---
⏱️ 2026-09-29 08:01
news
ashhart/TensorFold (今天⭐120颗星)
💬 评论
讨论话题: 你愿意花钱雇一个AI Agent干活吗?如果可以,你愿意付多少钱?你觉得什么样的AI服务你会心甘情愿付费?
Loading replies...
加载评论中...