02 / Lightning fast inference
DeepSeek V4.1
550tok/s
up to 5.8× faster
95–247tok/s
$ prism chat --model deepseek-v4.1
›
$ chat --model deepseek-v4.1
›
03 / Serverless Model APIs
Run open source models through a single API.No infrastructure to manage.
One platform for the full AI inference stack
Model APIs, GPU infrastructure, and agent runtimes. All in one platform.
Model APIs
GPU infrastructure
Agent runtimes
Top Open Source models
DeepSeek-V4.1-Flash, DeepSeek-V4-Flash, GLM-5.3, Kimi K3, Qwen3.8, and Qwen on one key.
Better price-performance
Up to 50% less than major cloud providers. Not because we cut corners, because we built the infrastructure.
Scale with your workload
Start small and scale seamlessly from APIs to dedicated clusters.
Serverless API
Dedicated endpoints
GPU control
Built for production reliability
Stable infrastructure with low latency, high throughput, and reliable uptime at scale.
99.99%
Uptime SLA
<50ms
P50 latency
Global
PoPs
04 / Questions
What is Prism?
Prism is an OpenAI- and Anthropic-compatible inference API for coding agents. It serves open-weight frontier models so you can run the primary agent loop without standing up your own GPU stack.
Which models can I call?
The public lineup is DeepSeek-V4.1-Flash, DeepSeek-V4-Flash, GLM-5.3, Kimi K3, Qwen3.8, and Qwen. Currently available model ids are deepseek-v4.1-flash, deepseek-v4-flash, glm-5.3, kimi-k3, qwen-3.8, and qwen.
Which API formats are supported?
Prism supports OpenAI Chat Completions at https://api.prisminference.com/v1 and Anthropic Messages at https://api.prisminference.com. Keep your existing client and change the base URL, API key, and model id.
How do I get an API key?
How is inference priced?
Which coding agents work with Prism?
Any client or gateway that can send OpenAI Chat Completions or Anthropic Messages, including Cursor, Claude Code, and configurable custom loops. Clients that require the OpenAI Responses API need a compatibility gateway.
Does Prism retain prompts or train on them?
05 / Get a key