Spot rental prices for GPU clusters have doubled from under $2 per hour to approximately $4 per hour over a seven-month period (as of August 4, 2026), according to Gavin Baker, CIO of Atreides Management. This surge reflects intense competition for AI compute capacity as hyperscalers and enterprises scale model training and inference workloads.
Consumer GPU pricing has similarly spiked due to GDDR6/GDDR7 memory constraints. NVIDIA's RTX 5090 is selling for a median $4,699.99 (up from $4,299.99 in June), representing a 135% markup over its $1,999 MSRP, while mid-range RTX 5060 Ti 16GB units climbed 39% from June at $804.99. The root cause: AI data centers now consume an estimated 70% of global memory output, leaving consumer GPUs starved for GDDR supply.
For architects: Doubled spot rental rates and 135%+ consumer GPU markups signal that compute capacity is now the binding constraint across every AI tier—from consumer fine-tuning to enterprise model training. Baker forecasts hyperscaler operating cash-flow growth will accelerate from 31% in Q1 2026 to 50% in Q2, driven by repricing of existing infrastructure contracts. Teams planning new model infrastructure should lock capacity now rather than wait for supply normalization; price relief is unlikely before late 2026 at earliest.