Executive Summary
The AI inference market is in the midst of a silent profitability crisis that will reshape the competitive landscape within 12-24 months. Detailed server economics reveal that operating commodity 120B parameter open-source models at current market prices is literally unprofitable—not marginally unprofitable, but losing money at scale regardless of whether hardware is rented or owned.
A single 8-card H100 server running GPT-OSS-120B at Together AI's published pricing ($0.15/input token, $0.60/output token) generates $0.67/hour in revenue against $24.51/hour in fully-loaded costs, producing a loss of -$23.84/hour. Even small models (20B) lose $0.66-1.44/hour at $2/hour rented GPU pricing.
This economics collapse reverberates through the entire open-source model infrastructure. Most startups operating inference endpoints for open-source models are hemorrhaging capital. This would be unsustainable for even 24 months before requiring either (1) massive price cuts driving competitors into bankruptcy, (2) consolidation into larger entities that can cross-subsidize with proprietary models, or (3) wholesale pivot to custom-optimized models with substantially higher margins.
China's Kimi K2, by contrast, achieves positive unit economics through a combination of aggressive MoE optimization, low-precision inference (BF16 + INT8), and consumer-direct pricing that reflects true infrastructure costs. At ¥0.58/input token and ¥2.29/output token, Kimi K2 generates approximately $3.20/hour revenue, producing roughly $1.20/hour profit.
This divergence between unsustainable Western open-source hosting and profitable proprietary models will drive three critical market transitions: (1) aggressive commoditization of open-source model quality (smaller models becoming substantially better), (2) migration of profitable inference workloads to proprietary models, and (3) infrastructure consolidation around companies with sufficient scale to achieve positive unit economics on commodity models.
Subscriber Content
Continue reading with a subscription.
You are reading a free preview. The full analysis, unit economics tables, and investment-relevant conclusions are available to Research and Full Quant Intelligence subscribers.
This research was produced by InAI Capital Advisor as part of our ongoing coverage of the global AI investment landscape. The analysis represents proprietary research conducted through expert network consultations and primary technical evaluation.