FactStream

Release of NVIDIA Nemotron 3 Ultra AI Model

Occurred 11 August 2026

NVIDIA released Nemotron 3 Ultra, a 550-billion-parameter open-weight mixture-of-experts (MoE) model designed for autonomous AI agents. While it leads other American open-source models, benchmarks indicate it remains behind Chinese competitors like Kimi K2.6.

Aggregated from 7 sources · Updated 25 Aug 2026, 03:07 UTC (UTC)

Share

Coverage Balance

7 sources
Center 100%

Blindspot Alert: Center Gap

This event is primarily covered by one side of the political spectrum. Niche or counter-narrative facts may be underrepresented.

Facts (10)

  • Established

    NVIDIA officially released Nemotron 3 Ultra, its most capable open-weight AI model, on June 4, 2026.

  • Established

    Nemotron 3 Ultra features 550 billion total parameters and utilizes a mixture-of-experts (MoE) architecture with 55 billion active parameters per token.

  • Established

    Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, outperforming other American open-weight models but trailing China's Kimi K2.6, which scores 53.9.

  • EstablishedNiche Signal

    Early enterprise adopters of Nemotron 3 Ultra include Perplexity, Palantir, ServiceNow, Glean, and CrowdStrike.

  • Established

    On August 11, 2026, NVIDIA released Nemotron 3.5 Lightning, a smaller 30B parameter model optimized for low-latency agent execution.

  • Established

    NVIDIA released the model weights, post-training recipes, and training corpora to Hugging Face and its own build.nvidia.com catalog.

  • Emerging

    The model supports a context window of up to 1 million tokens and utilizes a hybrid Mamba-Transformer architecture.

  • Emerging

    NVIDIA is collaborating with the Nemotron Coalition, including members like Nous Research and NAVER Cloud, to develop the foundation for a future Nemotron 4 family.

  • EmergingNiche Signal

    NVIDIA's cloud-compute budget for Nemotron development is reportedly capped at $7 billion through fiscal 2028.

  • AttributionNiche Signal

    NVIDIA claims the model delivers up to 5x faster inference and reduces costs for agentic tasks by up to 30% compared to previous iterations.

We'll extract atomic facts from your source and add them below for cross-checking. Single-source facts surface as Emerging or Unverified.