Executive briefing: Cost & Platform Landscape
The headline cost of model inference is a small and shrinking fraction of what enterprises actually spend to run generative AI in production. The production evidence reviewed here — modelled from public pricing and vendor disclosures rather than audited enterprise ledgers — indicates inference accounting for 20–40% of run-rate cost for mature deployments. Retrieval, evaluation, observability, governance, and human review consume the remainder — and are largely invisible at planning time.
2 min readCost & Platform Landscape