AI Compare AI Compare

Check prices, features and limits.

Pricing market
Automatic
Language
EN

Together AI

Cloud platform where developers run open AI models via API, fine-tune them and rent dedicated GPUs.

Free accessNo free plan published
Paid fromServerless Inference (Pay-as-you-go)Examples per 1M tokens (input/output): gpt-oss-20B USD 0.05 / 0.20; LFM2.5-8B-A1B USD 0.03 / 0.12; gpt-oss-120B USD 0.15 / 0.60; Llama 3.3 70B USD …

Category: Development & Code · Agent & RAG frameworks

Features

Together AI is a cloud for open models: inference via API (application programming interface), fine-tuning, and rented GPUs (graphics processing units). Serverless bills by tokens; Batch costs half of real-time inference. Cached inputs are cheaper.

Dedicated Inference isolates GPUs. Clusters and Provisioned Throughput (PTU) reserve capacity. Together publishes prices in US dollars; the pricing page names no EU region. The full token and GPU list is in the pricing overview.

Prices and plans

TierPricePeriodWhat you get
Serverless Inference (Pay-as-you-go)$0.03After consumptionExamples per 1M tokens (input/output): gpt-oss-20B USD 0.05 / 0.20; LFM2.5-8B-A1B USD 0.03 / 0.12; gpt-oss-120B USD 0.15 / 0.60; Llama 3.3 70B USD 1.04 / 1.04; Qwen3.7-Max USD 1.25 / 3.75; Kimi K3 USD 3.00 / 15.00
Cached input is cheaper, e.g. Kimi K3 at USD 0.30 instead of USD 3.00
One model is listed at USD 0.00: Ternary Bonsai 27B
Batch API at half the real-time inference price
Further modalities: image, video, audio, transcription, embeddings, rerank, moderation
Dedicated Inference$5.49 / $8.99Per GPU and hour, on-demandSingle-tenant GPU instances with guaranteed performance
Support for custom models, autoscaling
H200, B300, GB200 NVL72 and GB300 NVL72 on request only; reserved capacity on request only
GPU Clusters$3.99 / $5.99 / $8.19Per GPU and hour, on-demandReserved H100: USD 3.69 (7-30 days), 3.45 (31-90 days), 3.19 (91-180 days)
Reserved H200: USD 4.99 / 4.15 / 3.99
Reserved B200: USD 7.99 / 7.79 / 6.79
181+ days and GB200/GB300/B300 on request
Provisioned Throughput (PTU)$0.05Per PTU and per minuteReserved capacity in throughput units; tokens per minute delivered depends on the model
MiniMax M3: 166,667 input TPM, 833,333 cached TPM, 41,667 output TPM per PTU
Kimi K3: 16,667 / 166,667 / 3,333 TPM per PTU
GLM-5.2: 35,714 / 192,308 / 11,364 TPM per PTU
Fine-Tuning, Sandbox und Storage$0.48After consumptionSupervised fine-tuning per 1M tokens: up to 16B USD 0.48 (LoRA) / 0.54 (full); 17B-69B 1.50 / 1.65; 70-100B 2.90 / 3.20
Direct preference optimization per 1M tokens: up to 16B USD 1.20 / 1.35; 17B-69B 3.75 / 4.12; 70-100B 7.25 / 8.00
Minimum charge per fine-tuning job: USD 4.00
Code Sandbox: USD 0.0446 per vCPU-hour, USD 0.0149 per GiB RAM-hour
Code Interpreter: USD 0.03 per 60-minute session
Shared filesystem: USD 0.16 per GiB per month

Limits

Together AI is a US vendor and publishes prices in US dollars. The pricing page names no EU data region. According to the privacy policy, the vendor does not use customer data to train its own models without explicit opt-in and consent. Data may be processed outside the user's country. In Zero Data Retention mode, the vendor says inputs and outputs are not stored and are not used for model training, product improvement, or other secondary purposes.

Similar tools

FAQ

Is Together AI available for free?

No. No free access is published on the vendor's own pages.

Where is Together AI's vendor based?

Together Computer, Inc., based in USA. For data processing and privacy terms the vendor's own contract governs.

Is Together AI still active?

Active. Cloud platform for inference, fine-tuning and GPU clusters; usage-based billing only, no subscription tiers.

How many plans are documented for Together AI?

The catalog documents 5 plans with price, billing period, and included features.

Which use case is Together AI categorized for?

Together AI is listed in the catalog under “Development & Code.”

Where do the details about Together AI come from?

Prices, limits, and product details are documented against the listed official vendor page.

How we compare

We compare free access, paid entry, plan limits, features and the listed official sources. Unsourced figures stay marked as not published.

Only confirmed entries appear in the catalog. Prices, plans and limits come from listed official vendor pages. Missing figures stay marked as not published. Result lists are alphabetical by default and can be switched to Free access first; neither is a quality ranking. The vendor page always wins.

Edited and checked by the AI Compare team

Sources

Visit site (opens in a new tab)

Spotted an outdated figure? Vendor pages change without notice. Report an error