The practical takeaway from a new VentureBeat Pulse survey of 107 enterprises (all 100+ employees): companies are committing to AI infrastructure much faster than they can account for what they already run. Only about 21% operate AI in production at scale, yet 83% of GPU operators report utilization at 50% or less, and just 44% rigorously track what their compute actually costs. If you're scaling AI, the first fix isn't more hardware — it's instrumentation.
Today's stack is boringly familiar: hyperscalers (Google Cloud leads at 48%, alongside Microsoft, AWS, Oracle) plus model APIs from OpenAI, Anthropic, and Google. The specialized GPU "neoclouds" that dominate infrastructure headlines — CoreWeave, Lambda, Crusoe, Nebius — barely register. But intent points elsewhere: the top planned evaluation area over the next year is AI-specialized clouds (45%), a layer almost none of these firms use now. Roughly a third want to assess non-Nvidia accelerators, and 64% plan to switch or add a provider within twelve months, 38% within a quarter. That's unusually high churn for something this foundational.
How buyers actually decide is instructive. Integration with the existing stack (41%) and total cost of ownership (35%) dominate selection, while headline cost-per-million-tokens ranks dead last at just 8%. That's rational — but it exposes a contradiction. Enterprises claim to buy on TCO, yet most can't measure it. You can't optimize for total cost of ownership when the majority of your accelerator fleet sits idle and fewer than half of you track spend rigorously.

Why this matters: idle GPUs are expensive GPUs. With nearly half of operators running at 25% utilization or below, the efficiency headroom in existing fleets is enormous and mostly unmeasured. Before signing the next infrastructure contract, the higher-return move is to stand up real cost-and-utilization telemetry — per-workload GPU usage, cost attribution, and return tracking — so re-platforming decisions rest on evidence rather than vendor pitches.
There's also a bottleneck on the horizon most teams aren't watching. As inference scales, the binding constraint shifts from raw GPU compute to memory bandwidth and KV-cache capacity. Responses here scatter (Dell 31%, Nvidia 16%, the rest fragmented), and about 18% don't recognize the constraint at all. If your roadmap includes heavy inference, factor memory economics into architecture now — it will reshape both cost and hardware choices.
One caveat on the data: this is a single Q2 2026 wave of self-selected respondents skewed toward the mid-market and earlier-stage adopters, so treat it as directional rather than precise. Still, the direction is clear and worth acting on. The compute gap isn't a capacity problem more hardware solves — it's a visibility problem. Close the measurement gap before you buy the next layer, or you'll scale the same blindness you already have.
