The selling point of Groq is "speed", not "free" - it can run Llama at over 800+ tokens per second, which is several times faster than general cloud services. However, the official website does not clearly state the rate limit for the free tier, and the actual quota needs to be checked when used as a formal service. Additionally, it only runs open-source models, **without GPT-5 or Claude**.
View Change History
2026-08-05~2026-09-26Routine checks ran 24 times; no price or plan changes
2026-09-25~2026-10-07Automated checks could not be matched line by line with the official site 13 times (page redesign, region-based pricing or usage-based billing); not judged to be a price change; the verified date was not updated, pending manual review
Continuously re-checked, every fact dated
Has a free plan
What Is This
Custom LPU chips deliver the fastest inference for open-source models, billed per token with no monthly fee.
Free Version Limitations
The Free Plan needs no card and each model has its own per-minute and per-day limits: Llama 3.1 8B Instant is 30 requests per minute and 14,400 per day, 6,000 tokens per minute and 500,000 tokens per day; Llama 3.3 70B Versatile is 30 per minute and 1,000 per day, 12,000 tokens per minute and 100,000 tokens per day; GPT OSS 120B and 20B are 30 per minute and 1,000 per day, 8,000 tokens per minute and 200,000 tokens per day; Groq Compound is 30 per minute and 250 per day, 70,000 tokens per minute; Whisper Large V3 is 20 per minute and 2,000 per day, 7,200 audio seconds per hour and 28,800 audio seconds per day. Going over the limit returns HTTP 429, and cached tokens do not count towards the limits.
Available Models And Quantity
Model
Free quota
Calculate Time Window
Description
Open-source models (Llama, GPT OSS, etc.)
The official website does not specify a concrete upper limit.
—
Registration comes with a free API key, but the official pricing page does not specify the daily/minute usage limit for the free tier. The actual limit is subject to the console display.
Free Usage Limit
Features
Free quota
Description
Free API key
There is
Sign up to get
Speed limit
Not stated
This is the only key figure that was not stated when checked.
Available models
Open-source models are the primary focus
Llama, GPT OSS series
Inference speed
Same as paid plan
The free tier also runs on LPU, and its speed is its biggest selling point.
Quota Resets At
Not stated
Terms Of Use
Requires registration for an account and obtaining an API key
Best for: Applications requiring extremely low latency, large capacity, and cost savings (Chatbots, real-time translation)
API pay-as-you-go (usage-based)
Per million tokens: input US$0.05, output US$0.08; speed approximately 840 TPS
Which Model
Llama 3.1 8B
Usage Quota
Pricing based on tokens, pay-as-you-go
Max Reading Time
Based on model specifications
This Plan Includes
US$0.05 per million input tokens
US$0.08 per million output tokens
Approximately 840 TPS generation speed
Batch API can save an additional 50%
Not Included
GPT-5/Claude etc. closed models
Monthly flat-rate with unlimited usage
What Sets Us Apart
Similarly, running Llama, Groq's speed advantage is most obvious. However, it **does not provide GPT or Claude**, and those who need these models have to go through aggregators like OpenRouter.
An inference-acceleration platform: it doesn't train its own models, but runs open models (Llama, GPT-OSS, etc.) on its custom LPU chips at some of the fastest speeds available. Billed per token, no monthly fee. zh_support is marked partial: the platform UI is in English, and Chinese capability depends on the model selected. ⚠️ As of 2026-08-09, groq.com/pricing now 302-redirects to the homepage; official pricing and rate limits have moved to the developer docs at console.groq.com/docs/models and /docs/rate-limits. Automated price-checkers pulling groq.com/pricing will fail to find pricing — this is expected behavior and should not be read as the data being invalid.
Looking for Groq deals, discounts or promo codes? We check the official site automatically every day, review changes by hand, and date-stamp every entry.