AI Compare AI Compare

Check prices, features and limits.

Pricing market
Automatic
Language
EN

Groq

Inference cloud that runs open language, speech-to-text and text-to-speech models on its own hardware (LPU) through an OpenAI-compatible API, billed per token.

Free accessFree plan
Paid fromDeveloper Planprovider wording: 'When you upgrade, there's no immediate charge - you'll be billed for tokens at month-end or when you reach progressive billing …

Category: Development & Code · Agent & RAG frameworks

Features

Groq is an inference cloud with its own LPU (Language Processing Unit) architecture for very fast response times on open language, speech-to-text, and text-to-speech models. Free API access with lower rate limits simplifies onboarding.

Popular open-source models including Llama, Gemma, and Whisper are directly available. The pay-as-you-go model offers low cost per token and speech minute. Groq Playground enables direct browser-based testing without setup effort.

Prices and plans

TierPricePeriodWhat you get
Free Planfree-free API access with lower rate limits
examples of the published free limits: openai/gpt-oss-120b and openai/gpt-oss-20b each 30 requests per minute, 1,000 requests per day, 8,000 tokens per minute, 200,000 tokens per day
groq/compound and groq/compound-mini each 30 requests per minute, 250 per day, 70,000 tokens per minute
whisper-large-v3 and whisper-large-v3-turbo each 20 requests per minute, 2,000 per day, 7,200 audio seconds per hour, 28,800 per day
no batch or flex processing
Developer PlanfreeMonthly settlement afterwardsprovider wording: 'When you upgrade, there's no immediate charge - you'll be billed for tokens at month-end or when you reach progressive billing thresholds'
valid payment method required: credit card, US bank account or SEPA debit
higher rate limits, chat support, flex service tier, batch processing, spend limits
progressive billing at cumulative usage of $1, $10, $100, $500 and $1,000; above $1,000 lifetime usage only monthly billing
invoices below $0.50 are not issued
downgrade to Free possible at any time
Modellpreis openai/gpt-oss-120b0,15 $ Input / 0,60 $ OutputPer 1 million tokenscontext window 131,072 tokens, max 65,536 completion tokens
provider states about 500 tokens per second
Developer rate limit 250,000 TPM, 1,000 RPM
Modellpreis openai/gpt-oss-20b0,075 $ Input / 0,30 $ OutputPer 1 million tokenscontext window 131,072 tokens, max 65,536 completion tokens
provider states about 1,000 tokens per second
Developer rate limit 250,000 TPM, 1,000 RPM
Modellpreis openai/gpt-oss-safeguard-20b (Preview)0,075 $ Input / 0,30 $ OutputPer 1 million tokenscontext window 131,072 tokens
Developer rate limit 150,000 TPM, 1,000 RPM
preview model, not intended for production
Modellpreis qwen/qwen3.6-27b (Preview)0,60 $ Input / 3,00 $ OutputPer 1 million tokenscontext window 131,072 tokens, max 16,384 completion tokens, files up to 20 MB
Modellpreis qwen/qwen3.8-27b (Preview)0,80 $ Input / 4,00 $ OutputPer 1 million tokenscontext window 131,042 tokens, max 16,384 completion tokens, files up to 20 MB
Modellpreis whisper-large-v3 (Spracherkennung)0,111 $Per hour of audiofiles up to 100 MB
Developer rate limit 200,000 audio seconds per hour, 300 RPM
Modellpreis whisper-large-v3-turbo (Spracherkennung)0,04 $Per hour of audioDeveloper rate limit 400,000 audio seconds per hour, 400 RPM
Modellpreis canopylabs/orpheus-v1-english (Sprachausgabe, Preview)22,00 $1 million Characterscontext window 4,000 tokens
Developer rate limit 50,000 TPM, 250 RPM
Modellpreis canopylabs/orpheus-arabic-saudi (Sprachausgabe, Preview)40,00 $1 million Characterscontext window 4,000 tokens
Developer rate limit 50,000 TPM, 250 RPM
Modellpreis meta-llama/llama-prompt-guard-2-22m und -86m (Preview)0,03 $ Input und Output (22M) bzw. 0,04 $ Input und Output (86M)Per 1 million tokenscontext window 512 tokens
Developer rate limit 30,000 TPM, 100 RPM
Enterprise-Modelle (llama-3.1-8b-instant, llama-3.3-70b-versatile, minimaxai/minimax-m2.7)contact sales-the model table explicitly shows 'Contact Sales' for price and rate limit for these models
Performance Tier (Enterprise)contact sales-provider wording: 'Performance is delivered as provisioned throughput: you purchase input and output capacity bundles and pay for that provisioned capacity rather than per-token usage. Reach out to inquire about pricing.'
99.9 % availability SLA and 99 % latency guarantee per the agreement
enterprise plans only, context length under 8,192 tokens
available models: openai/gpt-oss-120b, openai/gpt-oss-20b, llama-3.3-70b-versatile

Limits

Accessible to European users. Per the billing FAQ, Groq explicitly accepts SEPA debit alongside credit cards (Visa, MasterCard, American Express, Discover) and US bank accounts - SEPA is the European payment route, so European customers are catered for. The provider runs its own data centres in the EU (EU-1 Vantaa, Finland) and the UK (EU-2 London). There is a free tier with published rate limits; the Developer tier requires a valid payment method. Groq is a US company; the privacy policy states: 'Groq is located in the United States, and maintains processing operations in various global jurisdictions.' Documentation page 'Your Data', verbatim: 'By default, Groq does not retain customer data for inference requests.' Exceptions, verbatim: data is retained only 'If you use features that require data retention to function (e.g., batch jobs, fine-tuning and LoRAs)' or 'If needed to protect platform reliability (e.g., to troubleshoot system failures or investigate abuse)'; such logs are 'retained for up to 30 days, unless legally required to retain longer' and can be switched off. On storage location, verbatim: 'All customer data is retained in Google Cloud Platform (GCP) buckets located in the United States. Groq maintains strict access controls and security standards as detailed in the Groq Trust Center. Where applicable, Customers can rely on standard contractual clauses (SCCs) for transfers between third countries and the U.S.' Additionally: 'All customers may enable Zero Data Retention (ZDR) in Data Controls settings.' The privacy policy names no use of customer inputs for model training.

Similar tools

FAQ

Is Groq available for free?

Yes. A permanently free plan is published, not a time-limited trial.

Where is Groq's vendor based?

Groq, Inc., based in USA. For data processing and privacy terms the vendor's own contract governs.

Is Groq still active?

Active. Groq describes itself as the 'premier neocloud for fast inference' and states it operates 13 data centres on four continents, including EU-1 in Vantaa (Finland) and EU-2 in London (UK). The developer offering is GroqCloud at console.groq.com; the platform page additionally names the building blocks GroqMetal (infrastructure), GroqCore (inference) and GroqAssured (enterprise control). The former groq.com/pricing page now redirects to the homepage; prices live in the console documentation.

How many plans are documented for Groq?

The catalog documents 14 plans with price, billing period, and included features.

Which use case is Groq categorized for?

Groq is listed in the catalog under “Development & Code.”

Where do the details about Groq come from?

Prices, limits, and product details are documented against the listed official vendor page.

How we compare

We compare free access, paid entry, plan limits, features and the listed official sources. Unsourced figures stay marked as not published.

Only confirmed entries appear in the catalog. Prices, plans and limits come from listed official vendor pages. Missing figures stay marked as not published. Result lists are alphabetical by default and can be switched to Free access first; neither is a quality ranking. The vendor page always wins.

Edited and checked by the AI Compare team

Sources

Visit site (opens in a new tab)

Spotted an outdated figure? Vendor pages change without notice. Report an error