Together AI
Cloud platform where developers run open AI models via API, fine-tune them and rent dedicated GPUs.
Category: Development & Code · Agent & RAG frameworks
Features
Together AI is a cloud for open models: inference via API (application programming interface), fine-tuning, and rented GPUs (graphics processing units). Serverless bills by tokens; Batch costs half of real-time inference. Cached inputs are cheaper.
Dedicated Inference isolates GPUs. Clusters and Provisioned Throughput (PTU) reserve capacity. Together publishes prices in US dollars; the pricing page names no EU region. The full token and GPU list is in the pricing overview.
- Run, fine-tune, and host open models on rented GPUs
- Serverless inference plus a Batch API at half the real-time rate
- Dedicated single-tenant GPUs with autoscaling for custom models
- Reserved GPU clusters (H100, H200, and others) and provisioned throughput
- Image, video, audio, transcription, embeddings, rerank, and moderation besides text
- Fine-tuning (LoRA or full), code sandbox, and shared filesystem
Prices and plans
| Tier | Price | Period | What you get |
|---|---|---|---|
| Serverless Inference (Pay-as-you-go) | $0.03 | After consumption | Examples per 1M tokens (input/output): gpt-oss-20B USD 0.05 / 0.20; LFM2.5-8B-A1B USD 0.03 / 0.12; gpt-oss-120B USD 0.15 / 0.60; Llama 3.3 70B USD 1.04 / 1.04; Qwen3.7-Max USD 1.25 / 3.75; Kimi K3 USD 3.00 / 15.00 Cached input is cheaper, e.g. Kimi K3 at USD 0.30 instead of USD 3.00 One model is listed at USD 0.00: Ternary Bonsai 27B Batch API at half the real-time inference price Further modalities: image, video, audio, transcription, embeddings, rerank, moderation |
| Dedicated Inference | $5.49 / $8.99 | Per GPU and hour, on-demand | Single-tenant GPU instances with guaranteed performance Support for custom models, autoscaling H200, B300, GB200 NVL72 and GB300 NVL72 on request only; reserved capacity on request only |
| GPU Clusters | $3.99 / $5.99 / $8.19 | Per GPU and hour, on-demand | Reserved H100: USD 3.69 (7-30 days), 3.45 (31-90 days), 3.19 (91-180 days) Reserved H200: USD 4.99 / 4.15 / 3.99 Reserved B200: USD 7.99 / 7.79 / 6.79 181+ days and GB200/GB300/B300 on request |
| Provisioned Throughput (PTU) | $0.05 | Per PTU and per minute | Reserved capacity in throughput units; tokens per minute delivered depends on the model MiniMax M3: 166,667 input TPM, 833,333 cached TPM, 41,667 output TPM per PTU Kimi K3: 16,667 / 166,667 / 3,333 TPM per PTU GLM-5.2: 35,714 / 192,308 / 11,364 TPM per PTU |
| Fine-Tuning, Sandbox und Storage | $0.48 | After consumption | Supervised fine-tuning per 1M tokens: up to 16B USD 0.48 (LoRA) / 0.54 (full); 17B-69B 1.50 / 1.65; 70-100B 2.90 / 3.20 Direct preference optimization per 1M tokens: up to 16B USD 1.20 / 1.35; 17B-69B 3.75 / 4.12; 70-100B 7.25 / 8.00 Minimum charge per fine-tuning job: USD 4.00 Code Sandbox: USD 0.0446 per vCPU-hour, USD 0.0149 per GiB RAM-hour Code Interpreter: USD 0.03 per 60-minute session Shared filesystem: USD 0.16 per GiB per month |
Limits
Together AI is a US vendor and publishes prices in US dollars. The pricing page names no EU data region. According to the privacy policy, the vendor does not use customer data to train its own models without explicit opt-in and consent. Data may be processed outside the user's country. In Zero Data Retention mode, the vendor says inputs and outputs are not stored and are not used for model training, product improvement, or other secondary purposes.
Similar tools
FAQ
Is Together AI available for free?
No. No free access is published on the vendor's own pages.
Where is Together AI's vendor based?
Together Computer, Inc., based in USA. For data processing and privacy terms the vendor's own contract governs.
Is Together AI still active?
Active. Cloud platform for inference, fine-tuning and GPU clusters; usage-based billing only, no subscription tiers.
How many plans are documented for Together AI?
The catalog documents 5 plans with price, billing period, and included features.
Which use case is Together AI categorized for?
Together AI is listed in the catalog under “Development & Code.”
Where do the details about Together AI come from?
Prices, limits, and product details are documented against the listed official vendor page.
How we compare
We compare free access, paid entry, plan limits, features and the listed official sources. Unsourced figures stay marked as not published.
Only confirmed entries appear in the catalog. Prices, plans and limits come from listed official vendor pages. Missing figures stay marked as not published. Result lists are alphabetical by default and can be switched to Free access first; neither is a quality ranking. The vendor page always wins.
Edited and checked by the AI Compare team
Sources
- https://www.together.ai/pricing
- https://docs.together.ai/docs/billing-credits
- https://www.together.ai/privacy
Visit site (opens in a new tab)
Spotted an outdated figure? Vendor pages change without notice. Report an error