Ask Cat › AI Tool Summary › DeepSeek
[Pitfalls] DeepSeek Free Tier Limits and API Pricing: Peak-Hour Rates Double, Old Model Names Quietly Rebilled
- Free tier: There is
- Cheapest paid plan: US$0/mo and up
- Free quota: The web version is basically free, but has peak traffic limits.
- Last checked: 2026-09-25
Article last updated: 2026-10-04
DeepSeek is the Chinese open-model company that went viral for a free chatbot and an API (the interface developers use to plug DeepSeek’s models into their own apps) priced low enough to make rivals nervous. This piece is about the traps, not the hype: what the free tier actually limits, how the API’s peak-hour pricing can quietly double your bill, and how a model-name swap changes what you’re billed for without you asking. If you’ve searched “DeepSeek free tier limits” or wondered whether the API is still as cheap as people say, here’s what we verified against DeepSeek’s own pricing pages.
The web app is free, but “free” has a catch at peak hours
DeepSeek’s web chatbot is free to use and doesn’t require registration. There’s no region lock, and it works fine from Taiwan. But “free” doesn’t mean “unlimited, always.” During peak hours, the web app throttles — meaning your requests get rate-limited or slowed — and DeepSeek has published no detailed, official quota policy explaining exactly how much you get or when a limit resets. There’s no fixed daily reset window we could find; it behaves more like a rolling throttle that kicks in when traffic is high.
That matters if you’re relying on the free web app for regular work. You can’t plan around a quota you can’t see. The practical fix: if you need predictable access, especially during Chinese business hours, don’t build a workflow that assumes the free web app will always respond instantly — it’s a courtesy tier, not an SLA.

Figure: api-docs.deepseek.com official pricing page, captured 2026-10-04.
API pricing looks cheap until you check the clock
The API is a separate system from the free web app, and it’s not free — you pay per token (tokens are the small chunks of text, roughly pieces of words, that the model reads and generates). But the headline price you see isn’t the only price. DeepSeek runs two pricing tiers depending on the time of day.

Illustration: Running tasks during local night can secretly hit daytime peak pricing in another time zone, silently doubling costs.
For DeepSeek-V4-Flash — called deepseek-flash in the API — off-peak pricing per million tokens is: US$0.003 for cached input (text the system recognizes from a recent, repeated request), US$0.15 for input that isn’t cached, and US$0.60 for output. During peak hours, every one of those numbers doubles: US$0.006, US$0.30, and US$1.20 per million tokens.
Peak hours are specific and easy to miss: Monday through Friday, 01:00–04:00 UTC and 06:00–10:00 UTC, excluding Chinese public holidays. If your application runs on a schedule, or if your users are concentrated in a timezone that overlaps those UTC windows, you could be paying double for a chunk of your traffic without realizing it. There’s no monthly cap and no contract or minimum spend — you’re billed purely on what you use, which is good for occasional use but means the peak-hour multiplier hits every request equally, with nothing capping the damage.
The pricier sibling, DeepSeek-V4-Pro (specifically version V4-Pro-0813), follows the same off-peak/peak doubling pattern but at higher base rates: US$0.022/US$0.044 per million tokens for cached input, US$0.66/US$1.32 for uncached input, and US$1.98/US$3.96 for output. Pro also caps concurrent requests at 500, versus 2,500 for Flash — worth knowing if you’re building something that needs to handle a lot of simultaneous requests cheaply; Flash has far more headroom there.
The model-name swap: old names, new price
Here’s the trap that’s easy to miss if you set up your integration months ago and never looked back. On 2026-09-10, DeepSeek launched DeepSeek-V4.1-Flash and gave it the API name deepseek-flash. The older model names — deepseek-v4-flash and deepseek-v4-flash-vision-exp — still work if you call them. But they’re no longer running the old model. Requests to those old names are now served by V4.1-Flash and billed at Flash’s current rates.
If you hardcoded an old model name into your code and haven’t touched it since, you’re not getting the model you think you’re getting, and you’re not necessarily paying what you think you’re paying. We also caught a real price change around the same window: on 2026-09-04, off-peak uncached input for Flash was US$0.22 per million tokens and output was US$0.66. By the time we verified the pricing page on 2026-09-25, those off-peak rates had dropped to US$0.15 input (uncached) and US$0.60 output — a genuine cut, not a trap, but it shows the pricing page moves. If you’re budgeting off a number you saw weeks ago, re-check it.

Figure: compiled by AMPM from its verified data (last verified 2026-09-25).
🔍 AMPM exclusive check
DeepSeek’s pricing page states the peak/off-peak split and the two model names, but it doesn’t flag, in plain language, that calling the legacy model names silently reroutes you to the new model and its current price — you have to read the fine print on the model-naming note to catch that. We also compared our 2026-09-04 price capture against the 2026-09-25 page ourselves and found the off-peak Flash numbers had changed (input uncached US$0.22 → US$0.15, output US$0.66 → US$0.60), which the vendor page doesn’t present as a “we lowered prices” notice anywhere obvious — it just reflects the current number with no changelog. Beyond that, every other number in this piece — the Pro rates, the concurrency limits (500 for Pro, 2,500 for Flash), the 1M-token context limit, and the 384K-token max output — matched the official api-docs.deepseek.com pages exactly when we checked them directly on 2026-09-25, and again on 2026-10-04 with no changes.
☀️ AMO, the budget-minded cat

I don’t pay for AI chat if I can avoid it, and DeepSeek’s web app means I mostly don’t have to — it’s free, no sign-up, and works from Taiwan with no region block. The only real cost for me would be hitting the peak-hour throttle, and since there’s no published quota number, I just avoid depending on it during the two UTC windows (01:00–04:00 and 06:00–10:00, weekdays) when things might slow down.
Where I’d actually spend money is the API, and even there DeepSeek is hard to beat on a pure numbers basis: cached input on Flash is US$0.003 per million tokens off-peak, and even uncached input is US$0.15. Output costs more at US$0.60 per million tokens off-peak, but there’s no minimum spend and no monthly commitment — you pay for exactly what you use. My rule: if I ever build something on the API, I schedule any heavy, non-urgent batch jobs outside those peak windows, because the 2x multiplier applies to every single number, including output. A job that costs US$0.60 per million output tokens off-peak costs US$1.20 at peak — same job, double the bill, just for bad timing.
🌙 PIMI, the performance-minded cat

My question is always whether a cheap model is cheap because it’s genuinely efficient, or cheap because it’s underpowered. DeepSeek’s specs suggest the former, at least on paper: Flash and Pro both support up to 1M tokens of context and up to 384K tokens of output, which is a large working window for a model at these prices. Pro’s concurrency cap of 500 simultaneous requests versus Flash’s 2,500 tells me Flash is the one built for throughput — if you’re running a high-volume pipeline, Flash’s higher concurrency ceiling plus its lower price per token makes it the obvious default, not Pro.
What concerns me more than raw speed is the model-name situation. DeepSeek-V4.1-Flash replaced what used to answer to deepseek-v4-flash on 2026-09-10, and the old endpoint name still responds — just with different underlying behavior. If you’re benchmarking output quality or consistency for a production use case, pointing at a legacy model name no longer guarantees you’re testing the model you think you’re testing. I’d re-run any quality benchmarks against the current deepseek-flash name explicitly, rather than trusting an old integration that’s technically still “working.”
Verdict: who should use it, who should skip it
Use DeepSeek’s free web app if you want a no-cost chatbot and can tolerate occasional slowdowns during weekday UTC peak windows — there’s no registration hurdle and no region block from Taiwan. Use the API if you’re cost-sensitive and your workload can be scheduled or batched: Flash’s off-peak rates (US$0.003–US$0.15 per million input tokens, US$0.60 output) are hard to beat, and there’s no contract or minimum spend. Skip relying on either tier for anything latency-critical during peak hours, since pricing and, for the web app, responsiveness both work against you there. And if you integrated the API months ago, check that you’re still calling deepseek-flash or the current Pro version name — not a legacy name quietly riding on new pricing.
Written: 2026-10-04 Price last verified: 2026-09-25 Tool last health-checked: 2026-10-04
🔗 Related on AMPM
📎 Sources
All prices and quotas in this article come from AMPM’s verified dataset for DeepSeek (last verified 2026-09-25); primary sources:
🐾 Meet the cats: AMO and PIMI
AMPM is a Taiwan-based AI-tool pricing watchdog. Our motto: ask before you subscribe. Two cats argue every tool from two sides:
- ☀️ AMO — the budget-minded cat. Always asks: is the free tier enough, and is paying actually worth it?
- 🌙 PIMI — the performance-minded cat. Cares about whether it works well, fast, and gets your job done.
When they are done arguing, you know whether to pay. Every number is checked by us, with the verification date at the end.
🐾 Our sister sites also checked
Let's take a look at these
- DeepSeek Comprehensive Introduction: Pricing, Features, and Actual Limitations
- DeepSeek Is the free quota enough?
- DeepSeek Alternatives
- Comprehensive Free Quota List for All Tools
