Blue-lit server rack in a data center
One OpenAI-compatible API for

The open-weight inference cloud.

Enterprise-grade, OpenAI-compatible inference for the open-weight model ecosystem. One API, transparent per-token pricing, no lock-in.

QwenLlamaDeepSeekMistralGPT-OSSGLMKimi / MoonshotGemmaPhiYiMistral SmallSmolLMQwenLlamaDeepSeekMistralGPT-OSSGLMKimi / MoonshotGemmaPhiYiMistral SmallSmolLM
1M
Max context window
99.9%
Target uptime
1
OpenAI-compatible endpoint
The catalog

The families that matter. One endpoint.

We are onboarding the major open-weight model families behind a single OpenAI-compatible API. New models and families ship as they release.

Model catalog
01

Qwen family

Multiple sizes, quantizations, and context windows. More families onboarding soon.

Onboarding
02

Llama family

Multiple sizes, quantizations, and context windows. More families onboarding soon.

Onboarding
03

DeepSeek family

Multiple sizes, quantizations, and context windows. More families onboarding soon.

Onboarding
04

Mistral family

Multiple sizes, quantizations, and context windows. More families onboarding soon.

Onboarding
05

GPT-OSS family

Multiple sizes, quantizations, and context windows. More families onboarding soon.

Onboarding
06

GLM family

Multiple sizes, quantizations, and context windows. More families onboarding soon.

Onboarding

+ 6 more families on the roadmap, with new models added as they ship.

Why Open Scale

Built for production inference

The details that separate a toy endpoint from infrastructure you can route real traffic through.

262K native context

Every open-weight model we serve exposes its full native context window, up to 262,144 tokens, so long documents and long-running agents don't get truncated.

Prefix caching built in

Automatic prefix caching at the inference layer means repeated system prompts and conversation history are billed at a fraction of the cost and return far faster.

OpenAI-compatible, drop-in

OpenAI-compatible by design. The Vercel AI SDK, LangChain, n8n, and other standard tooling work against a single endpoint, so the open ecosystem plugs in without rewrites.

Tools, reasoning & vision

First-class function calling, streamed tool calls, toggleable reasoning mode, structured outputs, and multimodal (text + image) input on the models that support it.

Global, fault-tolerant fleet

Inference is spread across a global fleet of high-performance GPUs with automatic failover. One region going down never takes your requests down with it.

Zero-data-retention path

We don't train on your prompts. A zero-data-retention, no-log serving path is available for workloads that require it, a property of the build and not a promise.

Infrastructure

A global fleet, engineered for tokens-per-dollar

From a single request to a fleet of GPUs, the same open-weight model served reliably wherever you are.

01

One OpenAI-compatible endpoint

A single base URL for the full model catalog. Auth, validation, and streaming are handled at the edge of the network.

02

Smart routing to the fleet

Requests are routed to the best-fit healthy GPU based on model, region, and load, with warm-session affinity for cache reuse.

03

Hardware-efficient inference

Models run on optimized vLLM deployments across our global fleet, tuned for tokens-per-second and time-to-first-token.

04

Streamed back, usage included

Responses stream immediately with accurate per-token usage and cached-token accounting, so you bill what you actually use.

Reach the open-weight models that matter.

One OpenAI-compatible endpoint, reachable through the platforms you already use. No lock-in.