Groq
Inference cloud that runs open language, speech-to-text and text-to-speech models on its own hardware (LPU) through an OpenAI-compatible API, billed per token.
Category: Development & Code · Agent & RAG frameworks
Features
Groq is an inference cloud with its own LPU (Language Processing Unit) architecture for very fast response times on open language, speech-to-text, and text-to-speech models. Free API access with lower rate limits simplifies onboarding.
Popular open-source models including Llama, Gemma, and Whisper are directly available. The pay-as-you-go model offers low cost per token and speech minute. Groq Playground enables direct browser-based testing without setup effort.
- Inference cloud for open language, speech-to-text, and text-to-speech models
- Free API access with lower rate limits
- Low price per token and per speech minute in pay-as-you-go model
- LPU (Language Processing Unit) architecture for very fast response times
- Support for popular open-source models (Llama, Gemma, Whisper and others)
- Groq Playground for direct browser-based testing
Prices and plans
| Tier | Price | Period | What you get |
|---|---|---|---|
| Free Plan | free | - | free API access with lower rate limits examples of the published free limits: openai/gpt-oss-120b and openai/gpt-oss-20b each 30 requests per minute, 1,000 requests per day, 8,000 tokens per minute, 200,000 tokens per day groq/compound and groq/compound-mini each 30 requests per minute, 250 per day, 70,000 tokens per minute whisper-large-v3 and whisper-large-v3-turbo each 20 requests per minute, 2,000 per day, 7,200 audio seconds per hour, 28,800 per day no batch or flex processing |
| Developer Plan | free | Monthly settlement afterwards | provider wording: 'When you upgrade, there's no immediate charge - you'll be billed for tokens at month-end or when you reach progressive billing thresholds' valid payment method required: credit card, US bank account or SEPA debit higher rate limits, chat support, flex service tier, batch processing, spend limits progressive billing at cumulative usage of $1, $10, $100, $500 and $1,000; above $1,000 lifetime usage only monthly billing invoices below $0.50 are not issued downgrade to Free possible at any time |
| Modellpreis openai/gpt-oss-120b | 0,15 $ Input / 0,60 $ Output | Per 1 million tokens | context window 131,072 tokens, max 65,536 completion tokens provider states about 500 tokens per second Developer rate limit 250,000 TPM, 1,000 RPM |
| Modellpreis openai/gpt-oss-20b | 0,075 $ Input / 0,30 $ Output | Per 1 million tokens | context window 131,072 tokens, max 65,536 completion tokens provider states about 1,000 tokens per second Developer rate limit 250,000 TPM, 1,000 RPM |
| Modellpreis openai/gpt-oss-safeguard-20b (Preview) | 0,075 $ Input / 0,30 $ Output | Per 1 million tokens | context window 131,072 tokens Developer rate limit 150,000 TPM, 1,000 RPM preview model, not intended for production |
| Modellpreis qwen/qwen3.6-27b (Preview) | 0,60 $ Input / 3,00 $ Output | Per 1 million tokens | context window 131,072 tokens, max 16,384 completion tokens, files up to 20 MB |
| Modellpreis qwen/qwen3.8-27b (Preview) | 0,80 $ Input / 4,00 $ Output | Per 1 million tokens | context window 131,042 tokens, max 16,384 completion tokens, files up to 20 MB |
| Modellpreis whisper-large-v3 (Spracherkennung) | 0,111 $ | Per hour of audio | files up to 100 MB Developer rate limit 200,000 audio seconds per hour, 300 RPM |
| Modellpreis whisper-large-v3-turbo (Spracherkennung) | 0,04 $ | Per hour of audio | Developer rate limit 400,000 audio seconds per hour, 400 RPM |
| Modellpreis canopylabs/orpheus-v1-english (Sprachausgabe, Preview) | 22,00 $ | 1 million Characters | context window 4,000 tokens Developer rate limit 50,000 TPM, 250 RPM |
| Modellpreis canopylabs/orpheus-arabic-saudi (Sprachausgabe, Preview) | 40,00 $ | 1 million Characters | context window 4,000 tokens Developer rate limit 50,000 TPM, 250 RPM |
| Modellpreis meta-llama/llama-prompt-guard-2-22m und -86m (Preview) | 0,03 $ Input und Output (22M) bzw. 0,04 $ Input und Output (86M) | Per 1 million tokens | context window 512 tokens Developer rate limit 30,000 TPM, 100 RPM |
| Enterprise-Modelle (llama-3.1-8b-instant, llama-3.3-70b-versatile, minimaxai/minimax-m2.7) | contact sales | - | the model table explicitly shows 'Contact Sales' for price and rate limit for these models |
| Performance Tier (Enterprise) | contact sales | - | provider wording: 'Performance is delivered as provisioned throughput: you purchase input and output capacity bundles and pay for that provisioned capacity rather than per-token usage. Reach out to inquire about pricing.' 99.9 % availability SLA and 99 % latency guarantee per the agreement enterprise plans only, context length under 8,192 tokens available models: openai/gpt-oss-120b, openai/gpt-oss-20b, llama-3.3-70b-versatile |
Limits
Accessible to European users. Per the billing FAQ, Groq explicitly accepts SEPA debit alongside credit cards (Visa, MasterCard, American Express, Discover) and US bank accounts - SEPA is the European payment route, so European customers are catered for. The provider runs its own data centres in the EU (EU-1 Vantaa, Finland) and the UK (EU-2 London). There is a free tier with published rate limits; the Developer tier requires a valid payment method. Groq is a US company; the privacy policy states: 'Groq is located in the United States, and maintains processing operations in various global jurisdictions.' Documentation page 'Your Data', verbatim: 'By default, Groq does not retain customer data for inference requests.' Exceptions, verbatim: data is retained only 'If you use features that require data retention to function (e.g., batch jobs, fine-tuning and LoRAs)' or 'If needed to protect platform reliability (e.g., to troubleshoot system failures or investigate abuse)'; such logs are 'retained for up to 30 days, unless legally required to retain longer' and can be switched off. On storage location, verbatim: 'All customer data is retained in Google Cloud Platform (GCP) buckets located in the United States. Groq maintains strict access controls and security standards as detailed in the Groq Trust Center. Where applicable, Customers can rely on standard contractual clauses (SCCs) for transfers between third countries and the U.S.' Additionally: 'All customers may enable Zero Data Retention (ZDR) in Data Controls settings.' The privacy policy names no use of customer inputs for model training.
Similar tools
FAQ
Is Groq available for free?
Yes. A permanently free plan is published, not a time-limited trial.
Where is Groq's vendor based?
Groq, Inc., based in USA. For data processing and privacy terms the vendor's own contract governs.
Is Groq still active?
Active. Groq describes itself as the 'premier neocloud for fast inference' and states it operates 13 data centres on four continents, including EU-1 in Vantaa (Finland) and EU-2 in London (UK). The developer offering is GroqCloud at console.groq.com; the platform page additionally names the building blocks GroqMetal (infrastructure), GroqCore (inference) and GroqAssured (enterprise control). The former groq.com/pricing page now redirects to the homepage; prices live in the console documentation.
How many plans are documented for Groq?
The catalog documents 14 plans with price, billing period, and included features.
Which use case is Groq categorized for?
Groq is listed in the catalog under “Development & Code.”
Where do the details about Groq come from?
Prices, limits, and product details are documented against the listed official vendor page.
How we compare
We compare free access, paid entry, plan limits, features and the listed official sources. Unsourced figures stay marked as not published.
Only confirmed entries appear in the catalog. Prices, plans and limits come from listed official vendor pages. Missing figures stay marked as not published. Result lists are alphabetical by default and can be switched to Free access first; neither is a quality ranking. The vendor page always wins.
Edited and checked by the AI Compare team
Sources
- https://groq.com/
- https://groq.com/platform
- https://groq.com/legal
- https://groq.com/privacy-policy
- https://console.groq.com/docs/models
- https://console.groq.com/docs/billing-faqs
- https://console.groq.com/docs/rate-limits
- https://console.groq.com/docs/performance-tier
- https://console.groq.com/docs/your-data
Visit site (opens in a new tab)
Spotted an outdated figure? Vendor pages change without notice. Report an error