Specs & substance
- Developed by
- NVIDIA
- Model family
- Nemotron
- Use case
- Flagship
- Modality
- Text
- Context window
- 128K tokens
- Architecture
- MoE NVFP4
- Version
- 3 Ultra
- License
- NVIDIA Open Model License
- Pricing
- $0.50 in · $2.50 out · $0.10 cache read
- Released
- Jun 2026
- Endpoint
- parasail-nemotron-3-ultra-550b-nvfp4
NVIDIA’s flagship open model optimized for enterprise reasoning workloads.
Nemotron 3 Ultra 550B is part of the Nemotron family by NVIDIA, and sits in the flagship category. It supports a 128K tokens context window and is built on a MoE NVFP4 architecture.
Key strengths: reasoning, enterprise. Parasail serves Nemotron 3 Ultra 550B on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.
Drop-in via the OpenAI SDK
Point any OpenAI-compatible client at Parasail and change the model name. That's it.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.parasail.io/v1"
)
response = client.chat.completions.create(
model="parasail-nemotron-3-ultra-550b-nvfp4",
messages=[
{"role": "user", "content": "Hello, what can you do?"}
],
stream=True,
max_tokens=1000
)
for chunk in response:
if chunk.choices and chunk.choices[0].delta.content is not None:
print(chunk.choices[0].delta.content, end="", flush=True) curl https://api.parasail.io/v1/chat/completions \
-H "Authorization: Bearer $PARASAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "parasail-nemotron-3-ultra-550b-nvfp4",
"messages": [{"role": "user", "content": "Hello, what can you do?"}],
"stream": true,
"max_tokens": 1000
}' Explore the library
Kimi K3
Moonshot's 2.8T-parameter open-weight multimodal agentic model — frontier reasoning, native vision, and a 1M-token context window.
GLM-5.3 Flash
Z.AI’s first natively multimodal GLM-5 model — 320B total / 18B active parameters at a tenth of GLM-5.2’s price.
GLM-5.3
Z.AI’s strongest GLM-5 agentic coder — same base as 5.2, all the gains from post-training, open-source SOTA on coding and cyber benchmarks.
Llama 4 Maverick
Meta’s multimodal Llama model with native image understanding.
Start building today
Instantly run any open model — popular or specialized.