Row of server racks in a data center hallway
About

The infrastructure layer for the open ecosystem.

Open Scale is an independent inference cloud. We exist to make open-weight models fast and reliable to run, behind a single OpenAI-compatible API.

Our mission

Frontier capability is moving into open weights at an unprecedented pace. But open models only matter if anyone can actually run them at scale. We build the serving layer: the GPUs, the routing, the caching, the billing. So that developers and companies can use these models the same way they use any other API.

Row of server racks in a data center hallway

What we're building

  • OpenAI-compatible API for open-weight models
  • Hardware-efficient inference on optimized vLLM
  • Transparent per-token pricing, no markup
  • Prefix caching for lower cost and faster returns
  • Zero-data-retention serving path
Principles

How we operate

Open by default

The open-weight ecosystem is where the most exciting models ship. We build infrastructure that lets anyone serve them, with no closed walled garden.

No markup, ever

Our pricing reflects the silicon. What you see on OpenRouter is what we charge.

Your data is yours

We don't train on prompts and we don't sell them. A zero-data-retention path is a feature we build, not a policy we hope you believe.

Built for routing

We optimize for the metrics routers actually measure: uptime, first-token latency, and throughput, so your traffic lands where it should.

Milestones

Where we're at

  1. 2025

    Founded

    Open Scale established in Wyoming, US.

  2. 2026

    Fleet online

    GPU fleet serving open-weight models with an OpenAI-compatible API.

Let's route some traffic through us.

Whether you're evaluating, integrating, or want to talk enterprise, we're easy to reach.