2026-08-07 21:24:582026-08-07 21:24 check failed: neither the automated system nor the AI detected a specific price. Source=https://groq.com/pricing
2026-08-06 12:13:04Checked 2026-08-06 12:13, verification failed: neither the automated checker nor the AI detected a specific price. Source = https://groq.com/pricing
看更多查核紀錄(2)
2026-08-05 22:36:282026-08-05 22:36 Check failed: AI did not find specific pricing information. Source=https://groq.com
2026-07-31Initial registration. (Official website groq.com/pricing direct read) Free = provides free API key but does not indicate rate limit; Paid = pure token billing without monthly fee, Llama 3.1 8B US$0.05/0.08, GPT OSS 20B US$0.075/0.30, GPT OSS 120B US$0.15/0.60, Llama 3.3 70B US$0.59/0.79 (per million tokens input/output); Batch API saves 50%, cached input half price.
Continuously re-checked, every fact dated
Has a free plan
What Is This
Self-developed LPU chip, open-source model runs the fastest inference platform
Free Version Limitations
The Free Plan needs no card and each model has its own per-minute and per-day limits: Llama 3.1 8B Instant is 30 requests per minute and 14,400 per day, 6,000 tokens per minute and 500,000 tokens per day; Llama 3.3 70B Versatile is 30 per minute and 1,000 per day, 12,000 tokens per minute and 100,000 tokens per day; GPT OSS 120B and 20B are 30 per minute and 1,000 per day, 8,000 tokens per minute and 200,000 tokens per day; Groq Compound is 30 per minute and 250 per day, 70,000 tokens per minute; Whisper Large V3 is 20 per minute and 2,000 per day, 7,200 audio seconds per hour and 28,800 audio seconds per day. Going over the limit returns HTTP 429, and cached tokens do not count towards the limits. (Read directly from the official docs at console.groq.com/docs/rate-limits on 2026-08-09)
Available Models And Quantity
Model
Free quota
Calculate Time Window
Description
開源模型(Llama、GPT OSS 等)
The official website does not specify a concrete upper limit.
—
Registration comes with a free API key, but the official pricing page does not specify the daily/minute usage limit for the free tier. The actual limit is subject to the console display.
Free Usage Limit
Features
Free quota
Description
Free API key
There is
Sign up to get
Speed limit
Not stated
This is the only key figure that was not stated when checked.
Available models
Open-source models are the primary focus
Llama, GPT OSS series
Inference speed
Same as paid plan
The free tier also runs on LPU, and its speed is its biggest selling point.
Quota Resets At
Not stated
Terms Of Use
Requires registration for an account and obtaining an API key
Free Version For Commercial Use
Yes
Common Pitfalls
The selling point of Groq is "speed", not "free" - it can run Llama at over 800+ tokens per second, which is several times faster than general cloud services. However, the official website does not clearly state the rate limit for the free tier, and the actual quota needs to be checked when used as a formal service. Additionally, it only runs open-source models, **without GPT-5 or Claude**.
Per million tokens: input US$0.05, output US$0.08; speed approximately 840 TPS (Checked 2026-07-31 from official website)
Which Model
Llama 3.1 8B
Usage Quota
Pricing based on tokens, pay-as-you-go
Max Reading Time
Based on model specifications
This Plan Includes
US$0.05 per million input tokens
US$0.08 per million output tokens
Approximately 840 TPS generation speed
Batch API can save an additional 50%
Not Included
GPT-5/Claude etc. closed models
Monthly flat-rate with unlimited usage
What Sets Us Apart
Similarly, running Llama, Groq's speed advantage is most obvious. However, it **does not provide GPT or Claude**, and those who need these models have to go through aggregators like OpenRouter.
The world's most widely used AI assistant, handling conversation, search, image generation, and voice all in one place
FreeNT$0/month, open to everyone. The official site lists the free-tier limits one by one: limited access to GPT-5.5 Instant, limited messages and uploads, limited and slower image generation, limited deep research, limited memory and context, limited Codex, limited ChatGPT Work desktop app. The only place the official site gives concrete numbers is the plan comparison table: free-tier GPT Instant total context window 27K (Go/Plus 54K, Pro 128K); free-tier single-input cap about 12 pages of text (Go/Plus about 40 pages, Pro about 250 pages); the context window for reasoning models is marked depends on the situation for the free tier (Go/Plus 256K, Pro 400K). Response time on the free tier is marked as limited by system resources and service conditions; only paid plans are fast. Chat history is unlimited on the free tier. Whether content is used for model training can be opted out of. The official site has never published a concrete messages-per-day figure, and this site does not fill in an estimate. (Read directly from chatgpt.com/zh-Hant/pricing over a Taiwan connection on 2026-08-09)
Long-form writing and document quality are its strengths, with the paid version including the Claude Code engineering tool
FreeThe quota is calculated based on a rolling 5-hour usage window (not simply resetting at a fixed number every day), and general conversations can be used for around 10-20 times, excluding Claude Code.
Google has the deepest ecosystem integration of AI, with a generous free version and frequent discounts for student plans.
FreeNT$0/month, free to anyone with a Google account. Model available: Gemini 3.6 Flash; the official site states plainly that access to 3.1 Pro may vary, meaning free-tier access to Pro fluctuates and is not guaranteed. Features included on the free tier: image generation and editing, Deep Research, Gemini Live, Canvas, Gems, Gemini Notebook (research and writing), the Google Flow creative studio, plus limited usage of Nano Banana Pro. Cloud storage 15 GB (shared across Gmail/Drive/Photos). How the quota is counted and when it resets, in the official footnote's own words: the Gemini app measures usage limits by compute, which depends on prompt complexity, the features you use and conversation length; usage resets every 5 hours, up to a weekly usage cap. In other words the free tier is not a fixed number of messages per day but a compute allowance that resets every 5 hours plus an overall weekly cap; long conversations, Deep Research and image generation burn through it faster than plain chat. When the quota runs out you can buy AI credits. The official site does not publish the specific compute figure for the free tier. (Read directly from gemini.google/subscriptions over a Taiwan connection on 2026-08-09)
An inference-acceleration platform: it doesn't train its own models, but runs open models (Llama, GPT-OSS, etc.) on its custom LPU chips at some of the fastest speeds available. Billed per token, no monthly fee. zh_support is marked partial: the platform UI is in English, and Chinese capability depends on the model selected. ⚠️ As of 2026-08-09, groq.com/pricing now 302-redirects to the homepage; official pricing and rate limits have moved to the developer docs at console.groq.com/docs/models and /docs/rate-limits. Automated price-checkers pulling groq.com/pricing will fail to find pricing — this is expected behavior and should not be read as the data being invalid.