NVIDIA Nemotron
Open NVIDIA model family with weights, training data, and recipes for specialized AI agents, plus a prototype API through NVIDIA NIM.
Category: Development & Code · Agent & RAG frameworks
Features
NVIDIA Nemotron is not a chat app with a monthly plan. It is an open model family for agents, with a prototype API and self-hosting.
- Open weights, training data, and recipes; Hugging Face download named on the topic page
- Nemotron 3: hybrid Mamba-Transformer MoE, context up to 1 million tokens according to NVIDIA
- Nemotron 3.5 Lightning: open 30B MoE with 3B active parameters for always-on agents
- Also named: Ultra 550B A55B, Super 120B A12B, Nano 30B A3B, and Nano Omni 30B A3B, plus Retriever, Parse, Speech, and Safety
- Prototype API with no credit card; base URL https://integrate.api.nvidia.com/v1, OpenAI Chat Completions
- Run via vLLM, SGLang, Ollama, llama.cpp, and NVIDIA NIM
Prices and plans
| Tier | Price | Period | What you get |
|---|---|---|---|
| Prototype API | free | free trial access, no token price | FAQ, verbatim: 'All models offer a free trial tier with no credit card required.' Base URL https://integrate.api.nvidia.com/v1, authentication with an NVIDIA API key The same endpoints implement the OpenAI Chat Completions API Example model on the Ultra page: nvidia/nemotron-3-ultra-550b-a55b |
| Production | contact sales | partner endpoint or self-hosted NIM | Ultra page: prototype on the free API endpoint, scale through partner endpoints or self-hosted deployments The topic page does not publish one NVIDIA token price for the model family Individual partner prices are those providers' prices, not an NVIDIA list price |
Limits
The models can be self-hosted. The hosted prototype runs on build.nvidia.com. The topic page states up to 1 million tokens of context for the Nemotron 3 family; that is a family figure, not a separately checked context length for every model. No single NVIDIA token list price for production is published there. Production uses the named partner endpoints or self-hosted NIM. The topic page does not say whether prototype-API prompts are used for training.
Alternatives and similar tools to NVIDIA Nemotron
FAQ
Is NVIDIA Nemotron available for free?
Not permanently. What is published is a time-limited free trial.
Where is the vendor of NVIDIA Nemotron based?
NVIDIA Corporation is based in the USA. Regardless of that location, data processing and privacy terms are governed by the vendor's own contract.
How many plans are documented for NVIDIA Nemotron?
The catalog documents 2 plans with price, billing period, and included features.
Is NVIDIA Nemotron still active?
NVIDIA Nemotron is still active: Open model family with a prototype API on build.nvidia.com and downloadable weights. Not a consumer chat with a monthly plan.
How many official sources are on file for NVIDIA Nemotron?
4 official sources are on file for NVIDIA Nemotron, and prices and limits are based on them.
How we compare
We compare free access, paid entry, plan limits, features and the listed official sources. Unsourced figures stay marked as not published.
Only confirmed entries appear in the catalog. Prices, plans and limits come from listed official vendor pages. Missing figures stay marked as not published. Result lists are alphabetical by default and can be switched to Free access first; neither is a quality ranking. The vendor page always wins.
Edited and checked by the AI Compare team
Sources
- https://developer.nvidia.com/topics/ai/nemotron
- https://build.nvidia.com/llms.txt
- https://build.nvidia.com/nvidia/nemotron-3-ultra-550b-a55b
- https://www.nvidia.com/en-us/about-nvidia/privacy-policy/
Visit site (opens in a new tab)
Spotted an outdated figure? Vendor pages change without notice. Report an error