Total — unique visitors
Browse by category Chatbots Image Generation Video Generation Audio & Voice Coding Writing Productivity Research AI Agents Free Tier Table
🕐 Last checked: 2026-10-08 09:37

Groq pricing, free tier and how to use it

AMPM AI Ops × Groq: verified pricing and free tier

Groq The fastest model inference platform— Groq

🌙 PIMI in one lineFree to use

Free-tier friendliness
★★★★★ 4/5
  • ✓ Has a free version
  • ✕ Quota reset is slow: Not stated
  • ✓ Free version allows commercial use
  • ✓ No watermark on free plan
  • ✓ The free quota has a publicly stated specific number.

This score is calculated from the fields checked on the site and is not a subjective review. Price confirmed. 2026-10-06

AMPM100 Confirmed · Chatbots Seat #7 (74.2 pts) · Selected 2026-08-24 · on the list 45 days→
Common Pitfalls

The selling point of Groq is "speed", not "free" - it can run Llama at over 800+ tokens per second, which is several times faster than general cloud services. However, the official website does not clearly state the rate limit for the free tier, and the actual quota needs to be checked when used as a formal service. Additionally, it only runs open-source models, **without GPT-5 or Claude**.

View Change History
2026-08-05~2026-09-26 Routine checks ran 24 times; no price or plan changes
2026-09-25~2026-10-07 Automated checks could not be matched line by line with the official site 13 times (page redesign, region-based pricing or usage-based billing); not judged to be a price change; the verified date was not updated, pending manual review

Continuously re-checked, every fact dated

Has a free plan

What Is This

Custom LPU chips deliver the fastest inference for open-source models, billed per token with no monthly fee.

Free Version Limitations

The Free Plan needs no card and each model has its own per-minute and per-day limits: Llama 3.1 8B Instant is 30 requests per minute and 14,400 per day, 6,000 tokens per minute and 500,000 tokens per day; Llama 3.3 70B Versatile is 30 per minute and 1,000 per day, 12,000 tokens per minute and 100,000 tokens per day; GPT OSS 120B and 20B are 30 per minute and 1,000 per day, 8,000 tokens per minute and 200,000 tokens per day; Groq Compound is 30 per minute and 250 per day, 70,000 tokens per minute; Whisper Large V3 is 20 per minute and 2,000 per day, 7,200 audio seconds per hour and 28,800 audio seconds per day. Going over the limit returns HTTP 429, and cached tokens do not count towards the limits.

Available Models And Quantity

ModelFree quotaCalculate Time WindowDescription
Open-source models (Llama, GPT OSS, etc.)The official website does not specify a concrete upper limit.—Registration comes with a free API key, but the official pricing page does not specify the daily/minute usage limit for the free tier. The actual limit is subject to the console display.

Free Usage Limit

FeaturesFree quotaDescription
Free API keyThere isSign up to get
Speed limitNot statedThis is the only key figure that was not stated when checked.
Available modelsOpen-source models are the primary focusLlama, GPT OSS series
Inference speedSame as paid planThe free tier also runs on LPU, and its speed is its biggest selling point.
Quota Resets AtNot stated
Terms Of UseRequires registration for an account and obtaining an API key
Free Version For Commercial UseYes

Sources: Groq official pricing page (last checked on 2026-07-31)

🔍 AMO checked Official site checked 2026-10-06 sources ▾

✉️ Price looks wrong? Tell us

Pricing Plan Differences

API - Llama 3.1 8B Instant
Best for: Applications requiring extremely low latency, large capacity, and cost savings (Chatbots, real-time translation)
API pay-as-you-go (usage-based)

Per million tokens: input US$0.05, output US$0.08; speed approximately 840 TPS

Which ModelLlama 3.1 8B
Usage QuotaPricing based on tokens, pay-as-you-go
Max Reading TimeBased on model specifications
This Plan Includes
  • US$0.05 per million input tokens
  • US$0.08 per million output tokens
  • Approximately 840 TPS generation speed
  • Batch API can save an additional 50%
Not Included
  • GPT-5/Claude etc. closed models
  • Monthly flat-rate with unlimited usage
What Sets Us Apart

Similarly, running Llama, Groq's speed advantage is most obvious. However, it **does not provide GPT or Claude**, and those who need these models have to go through aggregators like OpenRouter.

Subscribe On Official Website
API - GPT OSS 20B
Best for: Developers billed by token usage (small model, faster)
API pay-as-you-go (usage-based)

Per million tokens: input US$0.075, output US$0.30; speed approximately 1,000 TPS

This Plan Includes
  • Per million tokens: input US$0.075, output US$0.30
  • Speed approx. 1,000 TPS
Subscribe On Official Website
API - GPT OSS 120B
Best for: Developers billed by token usage (large model)
API pay-as-you-go (usage-based)

Per million tokens: input US$0.15, output US$0.60; speed approx. 500 TPS

This Plan Includes
  • Per million tokens: input US$0.15, output US$0.60
  • Speed approx. 500 TPS
Subscribe On Official Website
API - Llama 3.3 70B Versatile
Best for: Applications that require stronger reasoning capabilities without sacrificing too much speed
API pay-as-you-go (usage-based)

Per million tokens: input US$0.59, output US$0.79; speed approximately 394 TPS

Which ModelLlama 3.3 70B
Usage QuotaPricing based on tokens
Max Reading TimeBased on model specifications
This Plan Includes
  • US$0.59 per million input tokens
  • US$0.79 per million output tokens
  • Approximately 394 TPS generation speed
Not Included
  • Closed model
What Sets Us Apart

About 10 times more expensive than 8B, but with significantly stronger reasoning capabilities. Still faster than most cloud services.

Subscribe On Official Website

Plan details verified from: Groq official pricing page (last checked on 2026-07-31)

Similar Options

Other tools in the same category as Groq — free tiers and pricing vary, so you can compare them side by side.

ChatGPT logo ChatGPT AI chat assistant

The world's most widely used AI assistant, handling conversation, search, image generation, and voice all in one place

FreeFree plan NT$0 per month for everyone
Claude logo Claude AI chat/programming

Long-form writing and document quality are its strengths, with the paid version including the Claude Code engineering tool

FreeThe quota is calculated based on a rolling 5-hour usage wind…
Gemini logo Gemini AI chat assistant

Google has the deepest ecosystem integration of AI, with a generous free version and frequent discounts for student plans.

FreeNT$0/month, free to anyone with a Google account

Using Information

DeveloperGroq
CategoryChatbots Coding
PlatformWeb
Chinese SupportPartial support
Official LinkVisit Groq's official site →

Notes

An inference-acceleration platform: it doesn't train its own models, but runs open models (Llama, GPT-OSS, etc.) on its custom LPU chips at some of the fastest speeds available. Billed per token, no monthly fee. zh_support is marked partial: the platform UI is in English, and Chinese capability depends on the model selected. ⚠️ As of 2026-08-09, groq.com/pricing now 302-redirects to the homepage; official pricing and rate limits have moved to the developer docs at console.groq.com/docs/models and /docs/rate-limits. Automated price-checkers pulling groq.com/pricing will fail to find pricing — this is expected behavior and should not be read as the data being invalid.

Related Comparison Alternatives

Related Guide: [Special] Groq's Free Tier Caps the 70B and GPT OSS 120B Models at 1,000 Requests a Day

Special | Groq: free tier, limits and pricing checked by AMPM. Last verified 2026-10-03.

Read the full guide →

Groq Promotions & discounts

Looking for Groq deals, discounts or promo codes? We check the official site automatically every day, review changes by hand, and date-stamp every entry.

See the latest deals for all AI tools →

What Others Say

Share Your Experience Or Recommend Tools

Affiliate Links Notice