DeepSeek launched V4 Pro this week and raised its API prices with it, by between 50% and 1,100% depending on the model, the token type and the time of day. The clearest example is V4 Pro output, which cost a flat $0.87 per million tokens for nearly three months and now runs $3.96 per million at peak and $1.98 off peak. The new rates take effect from 16 August. The company said the change lets it allocate resources more reasonably.
The structure is the genuinely new part. DeepSeek has split its pricing into peak and off-peak windows, with peak defined as 01:00 to 04:00 and 06:00 to 10:00 UTC, and everything else discounted. Cache-hit input, cache-miss input and output are each priced separately in both windows, which gives V4 Pro six different rates rather than three.
This is a reversal by the company most responsible for the direction prices had been moving. DeepSeek’s cheap models triggered the round of global price cuts that followed its R1 release, and the pitch was always that frontier-adjacent capability could be served for a fraction of what American labs charged. Charging four times more at busy hours is a different argument, and it is an argument about capacity rather than capability.
The context is what makes it legible. On 13 August OpenAI previewed an Ultrafast tier running GPT-5.6 Sol on Cerebras hardware at up to 750 output tokens a second, roughly 14 times its standard speed, as a limited preview for selected customers with no price attached. One lab is metering demand by the hour and another is rationing speed by invitation. Both are behaving like operators of a scarce utility rather than sellers of software, and priced electricity is what that eventually looks like.



