Groq free tier: quotas, limits and gotchas
Groq Free Tier
Last verified: 2026-10-06
Models and quotas
| Model | Quota | Window | Notes |
|---|---|---|---|
| Open-source models (Llama, GPT OSS, etc.) | The official website does not specify a concrete upper limit. | — | Registration comes with a free API key, but the official pricing page does not specify the daily/minute usage limit for the free tier. The actual limit is subject to the console display. |
What you get for free
| Feature | Free tier | Notes |
|---|---|---|
| Free API key | There is | Sign up to get |
| Speed limit | Not stated | This is the only key figure that was not stated when checked. |
| Available models | Open-source models are the primary focus | Llama, GPT OSS series |
| Inference speed | Same as paid plan | The free tier also runs on LPU, and its speed is its biggest selling point. |
Limits at a glance
- Quota reset: Not stated
- Commercial use: Allowed
- Conditions: Requires registration for an account and obtaining an API key
Watch out for this
The selling point of Groq is "speed", not "free" - it can run Llama at over 800+ tokens per second, which is several times faster than general cloud services. However, the official website does not clearly state the rate limit for the free tier, and the actual quota needs to be checked when used as a formal service. Additionally, it only runs open-source models, **without GPT-5 or Claude**.
When the free quota is insufficient, how much does it cost to pay for it?
| Plan | pricing | Description |
|---|---|---|
| API - Llama 3.1 8B Instant | API pay-as-you-go (usage-based) | Per million tokens: input US$0.05, output US$0.08; speed approximately 840 TPS (Checked 2026-07-31 from official website) |
| API - GPT OSS 20B | API pay-as-you-go (usage-based) | Per million tokens: input US$0.075, output US$0.30; speed approximately 1,000 TPS (Checked on official website as of 2026-07-31) |
| API - GPT OSS 120B | API pay-as-you-go (usage-based) | Per million tokens: input US$0.15, output US$0.60; speed approx. 500 TPS (Checked 2026-07-31 from official website) |
| API - Llama 3.3 70B Versatile | API pay-as-you-go (usage-based) | Per million tokens: input US$0.59, output US$0.79; speed approximately 394 TPS (Checked on 2026-07-31 from official website) |
Checked 2026-10-06
Information for Taiwanese users
- Payment method:Payable; Credit card
- Chinese support: partial
- Available platforms:web
Related
Sources: Groq official pricing page (last checked on 2026-07-31)
