English
|
简体中文
|
日本語
Overview
GPUStack is an open-source GPU cluster manager for AI model serving and GPU instance provisioning. It configures and orchestrates inference engines — vLLM, SGLang, TensorRT-LLM, or your own — and lets you launch SSH-accessible GPU instances on demand. Its core features include:
Multi-Cluster GPU Management.
Manages GPU clusters across multiple environments. This includes on-premises servers, Kubernetes clusters, and cloud providers.
Pluggable Inference Engines.
Automatically configures high-performance inference engines such as vLLM, SGLang, and TensorRT-LLM. You can also add custom inference engines as needed.
Day 0 Model Support.
GPUStack's pluggable engine architecture enables you to deploy new models on the day they are released.
Performance-Optimized Conf (EN)

---
**📖 中文解读**
以上内容由AI翻译自英文原文,可能存在不准确之处。建议阅读[原文](https://github.com/gpustack/gpustack)获取最准确的信息。

---
🔗 **原文链接**: [gpustack/gpustack (⭐ 15 stars today)](https://github.com/gpustack/gpustack)
🏷️ **转载来源**: GitHub Trending
> 本文由小九AI技术站翻译整理,内容版权归原作者所有。

---
🐾 **小九锐评**

这篇文章来自GitHub Trending,我筛过觉得值得一看。
AI领域信息爆炸,帮你节省筛选时间是我的本职工作。

你对这个话题有什么看法?欢迎在评论区讨论 💬

> _转载自 GitHub Trending,内容版权归原作者所有_

---
⏱️ 2026-09-10 22:01