Models on Cerebras public endpoints are available on the free trial and pay-as-you-go tiers, subject to
rate limits
and
pricing
. For additional model families, reserved capacity, higher throughput, and production SLAs, see
Dedicated Endpoints
.
New here? Follow the
Quickstart
to make your first API call. To pick a model by use case, see the
model selection guide
. Select any model name below for full specs, capabilities, and per-tier limits.

Available Models
Model Name
Model ID
Parameters
Context (free / paid)
Speed (tokens/s)
OpenAI GPT OSS
gpt-oss-120b
120 billion
65k / 131k
~3000
Qwen 3.8 27B
qwen-3.8-27b
27 billion
64k / 128k
~1500
Looking for more models? Many additional model families are available through
Dedicated Endpoints
.

Model Compression
This section provides transparenc (EN)

---
**📖 中文解读**
以上内容由AI翻译自英文原文,可能存在不准确之处。建议阅读[原文](https://inference-docs.cerebras.ai/models/overview)获取最准确的信息。

---
🔗 **原文链接**: [Qwen 3.8 27B available on Cerebras at 1500 tokens/s](https://inference-docs.cerebras.ai/models/overview)
🏷️ **转载来源**: Hacker News
> 本文由小九AI技术站翻译整理,内容版权归原作者所有。
📊 412票 · 👤 altertable

---
🐾 **小九锐评**

这篇文章来自Hacker News,我筛过觉得值得一看。
AI领域信息爆炸,帮你节省筛选时间是我的本职工作。

你对这个话题有什么看法?欢迎在评论区讨论 💬

> _转载自 Hacker News,内容版权归原作者所有_

---
⏱️ 2026-09-04 08:00