AI cloud compute pricing is moving higher again, adding fresh pressure to organizations that rely on rented GPUs for model training, inference, simulation, and other accelerated workloads.
On Sept. 17, Investing.com reported, citing Bloomberg News and an emailed company statement, that Nebius will raise prices for its on-demand GPU services starting Oct. 1. The reported change applies to Nvidia H100, H200, B200, and B300 resources, along with certain CPU and memory resources.
Nebius’s public pricing page, viewed before the reported October change takes effect, listed on-demand rates of $3.85 per GPU-hour for HGX H100, $4.50 for HGX H200, $7.15 for HGX B200, and $7.85 for HGX B300.
What changed at AWS
AWS’s current EC2 Capacity Blocks for ML pricing page shows how expensive guaranteed high-end AI capacity has become. The page says Capacity Block reservation prices are updated regularly based on supply and demand trends, with the next update scheduled for October 2026.
Current AWS listed rates include $112.32 per hour for a p6-b300.48xlarge instance in several listed regions, or $14.04 per B300 accelerator. For B200 capacity, AWS lists p6-b200.48xlarge at $98.84 per hour, or $12.355 per accelerator, in regions including US East (Ohio), US East (N. Virginia), and US West (Oregon).
Hopper-generation GPU capacity remains costly as well. AWS lists p5e.48xlarge, which uses eight H200 accelerators, at $47.76 per instance-hour, or $5.97 per accelerator. AWS lists p5en.48xlarge in several U.S. regions at $54.92 per instance-hour, or $6.865 per H200 accelerator. These are reservation rates; AWS says operating system charges are billed separately when instances run.
The p5e.48xlarge example shows the scale of the increase. Network World reported in January that the H200-based p5e.48xlarge in US East (Ohio) had increased from $34.608 per instance-hour to $39.799. AWS’s current page lists that same instance type in US East (Ohio) at $47.76 per instance-hour, roughly 38% above the pre-January level reported by Network World.
Why guaranteed GPU capacity is getting more expensive
The price pressure is not happening in isolation. Gartner forecast on Sept. 16 that worldwide AI spending will reach $2.7 trillion in 2026, a 49.5% year-over-year increase. Gartner also projected AI infrastructure spending at about $1.48 trillion in 2026, making infrastructure the largest single category in its AI spending table.
That demand is hitting the parts of cloud computing that depend on scarce physical capacity: GPUs, memory, power, networking, and data center space. AWS acknowledged the supply problem in a June 2025 blog post announcing price reductions for some Nvidia GPU-accelerated EC2 instances, saying growth in GPU demand had outpaced industry-wide supply and made GPUs a scarce resource.
The latest pricing moves point to a split market. General-purpose cloud compute still has many optimization paths, including reserved commitments, architecture changes, and workload scheduling. High-end AI accelerators are different. For buyers that need guaranteed access to clusters of H100, H200, B200, or B300 GPUs, the cloud provider is not only selling compute time. It is selling certainty in a constrained market.
Not every AI bill will move the same way
The increases do not mean every AI product, chatbot, API call, or automation workflow will immediately cost more. Many AI services are sold through token-based pricing, bundled software subscriptions, or enterprise contracts that hide the underlying GPU-hour cost from the end user.
There are also differences between on-demand, preemptible, reserved, and committed-use models. Nebius lists lower preemptible pricing than on-demand rates. AWS Capacity Blocks are designed for reserved accelerator capacity rather than casual experimentation. Contracted customers may also have pricing terms that differ from public rate cards.
Still, public price moves matter because they reveal where supply is tight. If providers can raise list prices on scarce GPU capacity without immediately losing demand, it signals that buyers are still competing for access to the hardware needed to run large AI workloads.
What it means for smaller businesses and AI teams
For small and midsize organizations, the lesson is less about panic and more about cost visibility. AI pilots can look affordable at low volume, then become harder to justify when workloads run continuously, require guaranteed GPU capacity, or scale from a few internal users to customer-facing production systems.
Cost planning now needs to include model choice, workload timing, inference volume, data movement, storage, idle capacity, and whether a workload truly needs premium GPUs. Smaller models, batching, caching, and managed APIs can sometimes reduce the need for direct GPU rentals, while committed capacity can make sense for steady, predictable workloads.
The pricing pressure also reinforces a broader point Tech Help Canada has made in its coverage of using AI tools for SEO without publishing low-value content: AI spending should be tied to a clear business outcome. More automation, more tokens, or more GPU-hours do not automatically create more value.
October is the next pricing checkpoint
Two timing details now stand out. The reported Nebius increases are set to begin Oct. 1, while AWS says current EC2 Capacity Block prices are scheduled to be updated in October 2026.
That makes October a key month for cloud AI buyers watching GPU capacity costs. If rates keep climbing, AI teams may face a tougher tradeoff between speed, certainty, and budget control. If capacity eases or more efficient infrastructure changes the equation, some pricing pressure may soften. For now, the signal from public GPU cloud pricing is clear: guaranteed AI compute remains expensive, and demand is still strong enough for providers to test higher rates.

Tech Help Canada Staff researches, writes, and reviews practical content for business owners and professionals. Our coverage spans business, marketing, SEO, technology, and the tools and systems people use to grow and operate online. We focus on clear, useful information backed by research, hands-on experience, and editorial review. Learn more about our team and editorial standards. Need help with something? Contact Us







