Skip to content

status.openkey.ai is live — independent, worldwide uptime proof See it now ->

openkey

NVIDIA

11 models built · 7 served

NVIDIA is best known for GPUs and CUDA, but it also builds and ships open models under the Nemotron name. On OpenKey it has 11 models, and 7 of them are free. The lineup runs from small, cheap workhorses (Nemotron Nano 9B, 12B) up through a 120B Super tier to a 550B-parameter Ultra flagship with a 1M-token context window. There's also a dedicated content-safety model and multimodal variants that handle image, audio, and video input. NVIDIA's strategy looks less like a single flagship product and more like a full toolkit: pick the size and modality that fits the job, and most of them cost nothing to try.

Lineup

Start with the free tier — nvidia/nemotron-3-nano-30b-a3b:free, nvidia/nemotron-3-super-120b-a12b:free, and nvidia/nemotron-3-ultra-550b-a55b:free are all zero-cost and cover small, mid, and large capacity. If you need paid throughput with predictable pricing, nvidia/nemotron-3-nano-30b-a3b runs $0.05/$0.20 per 1M input/output tokens, nvidia/nemotron-3-super-120b-a12b is $0.085/$0.40, and the flagship nvidia/nemotron-3-ultra-550b-a55b is $0.50/$2.20 with a 1M-token context window. For vision or audio/video input, use nvidia/nemotron-nano-12b-v2-vl:free or nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free. For moderation pipelines, nvidia/nemotron-3.5-content-safety:free is the purpose-built option.

Prices per 1M tokens, flat 3% fee included.

Models NVIDIA serves

Beyond its own models, NVIDIA hosts 7 models in the OpenKey catalog. Prices are NVIDIA’s own list price per 1M input tokens, before our fee.

From the OpenRouter catalog, Jul 28, 2026. All providers

Privacy policy ↗Terms of service ↗

Questions

What's the price range for NVIDIA models on OpenKey?
Paid Nemotron models range from $0.05/$0.20 per 1M input/output tokens (nemotron-3-nano-30b-a3b) up to $0.50/$2.20 (nemotron-3-ultra-550b-a55b). On OpenKey that's provider price × 1.03: the Ultra tier works out to about $0.515 input / $2.266 output per 1M tokens after the 3% fee. Seven of NVIDIA's 11 models are free, so you can test most of the lineup at no cost first.
Which NVIDIA model should I use as the default?
For general text tasks with a large context window, nvidia/nemotron-3-super-120b-a12b is the middle ground at $0.085/$0.40 per 1M tokens with a 1M-token context — cheaper than the Ultra flagship but far more capable than the Nano tier. If cost isn't a concern and you need maximum context and scale, nemotron-3-ultra-550b-a55b ($0.50/$2.20) is the largest model in the lineup.
Does NVIDIA offer free models via OpenKey?
Yes — 7 of NVIDIA's 11 models are free, including nano, super, and ultra-scale variants (nemotron-3-nano-30b-a3b:free, nemotron-3-super-120b-a12b:free, nemotron-3-ultra-550b-a55b:free) plus multimodal and content-safety models. That covers text, vision, audio/video, and moderation use cases without a paid tier.