Inference on Open-Source Models vs. Proprietary APIs at Scale

Open-source models beat APIs on cost only above tens of millions monthly tokens.

Staff Writer · · 8 min read
Cover illustration for “Inference on Open-Source Models vs. Proprietary APIs at Scale”
Inference Serving · October 2, 2026 · 8 min read · 1,776 words

Two years ago the choice was effectively binary: OpenAI or expensive GPU rigs, but today there is a genuine marketplace with meaningfully different economics on each path. That binary has dissolved into a genuine marketplace, and the economics on each side now differ enough that the old shorthand no longer holds. Closed-source frontier models, GPT-5.2, Claude Opus 4.6, Gemini 3.1 Pro among them, still lead on benchmark ceiling and on how little setup an API call demands, while open-weight families such as Llama, Mistral, DeepSeek, and Qwen win on cost at scale, on customization, on data sovereignty, and on freedom from being tied to one vendor's roadmap. Many teams still reach for a simple rule: prototype on an API, move to self-hosting once volume justifies it. That rule is a heuristic, not an analysis, and it hides the variables that actually decide whether a path survives contact with a real production workload. Those variables live in the serving layer: KV cache memory pressure, the storage hierarchy underneath it, throughput under real concurrency, and the precision format a team chooses for its weights. A team that adopts open-source models without reckoning with what free weights demand from infrastructure can end up spending more than the API bill it was trying to avoid.

What each path costs once you move past per-token headline pricing

Per-token pricing and the promise of free weights both obscure what each path actually costs at scale. Open-weight models served through managed inference providers such as Fireworks AI, Together AI, and Groq cost a fraction of frontier APIs: Llama 4 Maverick on Fireworks AI is dramatically cheaper per token than Claude Opus 4.6 on AWS Bedrock. That gap makes the open-weight path look like an easy win until the cost of self-hosting enters the picture. Minimal internal deployments can run into six figures annually, and implementations built for enterprise scale can cross into seven figures. The point at which self-hosting starts to beat per-token API costs runs somewhere between tens of millions and low hundreds of millions of tokens a month, depending on the workload. Below that volume, the simplicity of an API wins on total cost; above it, self-hosted infrastructure pays for itself quickly. A third model complicates the picture further: flat-rate inference subscriptions, such as Featherless.ai's monthly tiers covering unlimited tokens, offer predictable costs regardless of volume, but they come with concurrency limits that matter once a workload needs real throughput. Given all of this, most production teams in 2026 don't pick one path exclusively, since closed-source frontier models lead on benchmark ceiling and API simplicity while open-weight families lead on cost at scale, customizability, data sovereignty, and freedom from vendor lock-in. They run closed models on user-facing surfaces where the capability ceiling is visible to end users, and they run open-weight models for batch inference, classification, embedding, and any task where the gap in capability is small relative to the gap in cost.

KV Cache: The Central Infrastructure Variable for Both Paths

Whichever path a team chooses, KV cache pressure is the serving-side bottleneck that governs memory capacity, batching efficiency, and cost per request. The KV cache holds the intermediate attention states a model needs so it doesn't recompute attention across the entire sequence every time it generates a new token, and at production scale that cache frequently takes up more memory than the model's own weights. The scale of the problem is concrete: a large model running at long context can require tens of gigabytes of cache for a single request, and even a modest batch of concurrent requests pushes total cache demand into the hundreds of gigabytes, well past what a single GPU's high-bandwidth memory can hold. Traditional inference setups waste most of that memory through fragmentation. vLLM's PagedAttention technique addresses the waste directly, cutting it to under 4% and unlocking meaningful throughput gains as a result.

Fragmentation is only one source of pressure. Long-context and agentic workloads compound the problem by design, since both demand the model hold far more context in memory per request. GLM-5.2, a large mixture-of-experts model with a long context window released under the MIT license, uses DeepSeek Sparse Attention alongside a newer technique called IndexShare, which reuses the same sparse-attention indexer across groups of four layers and cuts per-token compute at full context length by a meaningful factor. MiniMax-M3, another large MoE model with a long context window, uses its own MiniMax Sparse Attention to achieve substantially faster prefilling and decoding at long context compared to its prior generation. Both models point to the same shift: very long context windows have become close to a design standard across the open-source frontier, not an edge case engineering teams can treat as rare. Compression techniques are advancing fast enough to change the picture within a single model generation. Between DeepSeek-V3.2 and DeepSeek-V4-Pro, a technique called Compressed Sparse Attention cut KV cache size for long contexts by roughly an order of magnitude. Research activity in this area has accelerated correspondingly: TurboQuant applies online vector quantization to cache storage, KVzip handles query-agnostic eviction by reconstructing context rather than discarding it blindly, RocketKV pairs coarse-grained eviction with fine-grained sparse attention, and KVShare enables multi-tenant reuse of KV cache across separate users and sessions.

The proprietary API path doesn't eliminate any of this pressure. It simply hides it from the customer, since the provider absorbs the cost of managing cache at scale and folds it into the per-token price it charges. Self-hosting exposes that same pressure directly to the team running the model, which is precisely what makes it something an engineering team can tune, measure, and optimize rather than pay for blindly.

The storage hierarchy that makes self-hosted inference viable at long context

Self-hosted inference at long context only works if KV cache blocks can move across a hierarchy spanning GPU HBM, CPU DRAM, and NVMe, and building that hierarchy correctly is a genuine engineering problem that shapes whether the open-source path delivers on its cost promise. The scale of the constraint is specific: on a single H100 serving a large model at long context, one user's KV cache can consume tens of gigabytes, and a handful of concurrent users can demand more capacity than the H100's entire HBM provides. At that point, moving cold KV blocks off the GPU stops being optional.

What emerges is a three-tier architecture, each tier defined by a different bandwidth and a different job. Hot KV blocks, the ones involved in active generation, stay in GPU HBM, running at roughly 3.35 TB/s. Warm blocks, completed but likely to be needed again in a multi-turn conversation, move to CPU DRAM, reached over PCIe 5.0 at roughly 63 GB/s. Cold blocks, historical context and prefix caches unlikely to be touched again soon, sit on NVMe SSD, reached over PCIe 4.0 at roughly 7 GB/s.

The industry's response to this problem is itself evidence of how unresolved it remains. NVIDIA announced its Inference Context Memory Storage Platform at CES 2026 to standardize NVMe offload of inference context, built around Rubin GPU cluster-level cache capacity and the coming BlueField-4 DPU, with hardware-accelerated cache placement handled through the DOCA Memos framework, the Dynamo KV cache offload engine, and NVIDIA's own Inference Transfer Library, NIXL. Partners are building toward the same problem from different angles: one vendor is developing a 1RU token memory product built on CXL memory, PCIe Gen 5, NVMe, and GPUDirect with RDMA, designed to let KV cache get reused across sessions, models, and nodes with very low latency through RDMA over NVMe-oF. Another vendor's memory grid product, first shown at GTC 2025 and validated with a hardware partner on NVIDIA Grace CPUs and BlueField-3 DPUs, pools and persists KV cache outside GPU memory entirely, treating it as a dedicated memory extension layer rather than a cache that lives only on the GPU. Research is also exploring CXL as a tier in its own right: a proposal called CXL-SpecKV uses disaggregated FPGA-based speculative KV cache that moves data from CXL memory into GPU L2 cache on demand.

None of this is incidental tuning work that a team can leave until after launch. NVMe throughput, DRAM bandwidth, and the eviction policy that decides when a block moves between tiers all have to be treated as first-class design decisions from the start, because getting them wrong under live load is far more expensive than designing for them up front. This is exactly the complexity that the proprietary API path abstracts away, and it's a legitimate reason teams choose APIs in early scale, because they defer this entire category of infrastructure work to someone else.

How the three dominant self-hosted serving frameworks handle KV pressure differently

Once a team commits to self-hosting, the serving framework it picks decides how KV pressure actually gets managed in practice, and vLLM, TensorRT-LLM, and SGLang each take a different architectural position on that question. Those differing bets produce measurably different performance depending on the shape of the workload a team is running.

In Spheron's H100 testing, TensorRT-LLM outperformed vLLM at every concurrency level tested, from a modest edge at a single request to a noticeably wider lead at 50 concurrent requests. That performance comes at a cost: TensorRT-LLM demands one to two weeks of setup time and ties a team into a single vendor's toolchain. SGLang takes a different architectural approach entirely, built around RadixAttention, and its advantage depends heavily on what the workload looks like. On workloads with heavy prefix sharing, long shared system prompts combined with shared retrieval-augmented context, SGLang produces throughput gains over vLLM that run up to several times higher, with output tokens generated more than twice as fast when requests share context. On unique-prompt traffic, where little or nothing is shared between requests, that advantage narrows to within normal benchmark variance. SGLang also supports Mooncake and NIXL as transfer backends for disaggregated serving, and published results show meaningfully higher decoding throughput on NVIDIA GB200 NVL72 clusters using that configuration.

None of these rankings hold universally. A September 2026 benchmark run on a smaller Llama model, using vLLM 0.23.0 and SGLang 0.5.13 as of June 20, 2026, found the opposite ordering at full saturation: vLLM reached the highest throughput and the lowest cost per token of the frameworks tested. Framework performance depends on model size and concurrency regime as much as it depends on architecture, and a ranking that holds at one scale can invert at another. Choosing a serving framework, in other words, is not a one-time decision made in the abstract. It has to be tested against the actual shape of the workload a team expects to run, at the actual scale that workload will reach.

Sources

  1. LLM API Pricing Comparison 2026: The Complete Guide to Inference Costs - Featherless
  2. Open Source vs Closed LLMs: Technical Comparison 2026
  3. Open Source LLM Cost: Hidden Expenses in 2026
  4. KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
  5. An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
  6. Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

More in Inference Serving