Measured open-model inference.
Priced by the yield, not the hype.

Inference Yield serves high-demand open-weight models on dedicated GPUs, with every serving configuration chosen by continuous benchmarking — quality-gated quantization, prefix caching with honest cached-token pricing, and published performance numbers measured on the live endpoint, never projected.

Current model

ModelQuantizationContextInput /MCached input /MOutput /M
Qwen3.8 27BFP8 (disclosed)262,144$0.40$0.08$2.70

OpenAI-compatible API · streaming · tool calling · structured outputs · usage accounting on every request (cached tokens itemized).

How we're different

Get started

curl https://inferenceyield.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen/qwen3.8-27b", "stream": true,
       "messages": [{"role": "user", "content": "Hello!"}]}'

Full examples in the docs. For keys, higher limits, or design-partner pricing: [email protected].

US datacenter NVIDIA H200 vLLM-based stack Zero data retention available