> ## Content Index
> Fetch the complete content index at: https://globalfeed.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Ollama cloud adds off-peak token rates: DeepSeek-V4-Flash and V4-Pro half price outside 12:00 to 18:00 UTC on weekdays, all weekend
- URL: https://globalfeed.ai/en/ollama-cloud-adds-off-peak-token-rates-deepseek-v4-flash-and-v4-pro-half-price-outside-12-00-to-18-00-utc-on-weekdays-all-weekend/
- Published: 2026-09-06T01:59:38.000Z
- Updated: 2026-09-06T01:59:38.000Z
- Description: Ollama has applied the logic of an electricity tariff to AI tokens on its cloud: DeepSeek-V4-Flash and DeepSeek-V4-Pro are half price outside 12:00 to 18:00 UTC on weekdays and all day on weekends. The company says the same discount will reach more models soon. The move follows its 31 August switch
- Author: GlobalFeed Editor
- Tags: Ollama, DeepSeek, cloud pricing, local AI, x-ollama, x-deepseek_ai, dil-en, elle, video, x-gitti

**Ollama** has introduced off-peak token rates on its cloud service. According to [the 6 September announcement on the company's official X account](https://x.com/ollama/status/2096374001119744190?ref=globalfeed.ai), **DeepSeek-V4-Flash** and **DeepSeek-V4-Pro** are now half price outside 12:00 to 18:00 UTC on weekdays and all day on weekends. In Türkiye, peak hours fall between 15:00 and 21:00\. The logic is a night-time electricity tariff: cheaper when demand is low.

## The prices

Ollama's pricing page puts the peak rate at Monday to Friday, 12:00 to 18:00 UTC, with the off-peak rate at 50 percent of peak. Off-peak, per million tokens:

- **DeepSeek-V4-Flash:** input $0.22, cached input $0.007, output $0.66.
- **DeepSeek-V4-Pro:** input $0.66, cached input $0.022, output $1.98.

At peak, each figure doubles. Only these two models carry the split so far; Ollama says off-peak pricing "will be available soon for more models".

## Background: the move to per-token pricing

It follows Ollama's 31 August pricing change: cloud billing moved from GPU time to per-token prices, with a monthly usage allowance in each plan. Pro is $20 a month ($60 of usage), Max $100 a month ($300 of usage), Team $500 a month ($1,000 of shared usage, unlimited users). The free plan has a small allowance for starter models; unused allowance does not roll over.

Ollama says requests run on dedicated compute in the US and Europe, with Singapore for a limited number of Qwen models, while its pricing page says hosting is primarily in the US. Ollama also cites zero data retention (ZDR) for the DeepSeek models, saying it does not log prompts or train on customer data.

## For the corporate reader

An organization that moves batch jobs, or overnight summarization and classification runs, into off-peak hours can halve its bill on these two models. The promise of more models soon is still a promise, with no date. Whether data is processed in the US or Europe matters for GDPR and KVKK: Ollama names two regions but does not say whether customers can choose.

## Sources

- **Primary source:** Ollama's X post, 6 September: [https://x.com/ollama/status/2096374001119744190](https://x.com/ollama/status/2096374001119744190?ref=globalfeed.ai)
- **Pricing page:** [https://ollama.com/pricing](https://ollama.com/pricing?ref=globalfeed.ai)
- **Ollama blog, 31 August:** [https://ollama.com/blog/transparent-pricing](https://ollama.com/blog/transparent-pricing?ref=globalfeed.ai)