The Applied Layer

Topic

Cost & Platform Landscape

Pillar 4 of The Applied Layer research programme.

14 Jul 2026BriefingMember

Executive briefing: Cost & Platform Landscape

The headline cost of model inference is a small and shrinking fraction of what enterprises actually spend to run generative AI in production. The production evidence reviewed here — modelled from public pricing and vendor disclosures rather than audited enterprise ledgers — indicates inference accounting for 20–40% of run-rate cost for mature deployments. Retrieval, evaluation, observability, governance, and human review consume the remainder — and are largely invisible at planning time.