AI

·

5 MIN

Nemotron 3: NVIDIA's open models built for agents

NVIDIA's Nemotron 3 Ultra: 550B open-weights, agent-focused, strong on the index, free day one. A real on-prem option.
Nemotron 3: NVIDIA's open models built for agents — thumbnail for the iiterate Signals article
LOCATION
Germany
SERIES
Sovereign & Offline AI
AUTHOR
Aashwin Shrivastava
PUBLISHED

A frontier-scale, open-weight model built for long-running agents rather than chat, that runs fast on NVIDIA hardware many clients already own. That is what makes Nemotron 3 a real on-prem option. NVIDIA completed the family at Computex with Nemotron 3 Ultra, a 550-billion-parameter Mixture-of-Experts model (55B active), released 4 June under the Linux Foundation's Open MDW license.

01.

Strong on the index, interesting under the hood

Per Artificial Analysis, Ultra scores 47.7 on its Intelligence Index, the strongest US open-weights model, ahead of Gemma 4 31B (39.2) and the smaller Nemotron 3 Super (36.0).

The technical report details the interesting part: a hybrid Mamba-Attention MoE with LatentMoE, NVFP4 training, and multi-token-prediction layers for native speculative decoding, up to 5x higher throughput than the previous generation, with a context window up to 1M tokens.

02.

Built for agents, not chat

NVIDIA was explicit at launch: Nemotron 3 is for "agentic problems where an AI is trying to solve difficult tasks autonomously", long-running tool use rather than conversation. The family is sized to common NVIDIA footprints: Nano (30B) for consumer hardware, Super (120B), and Ultra at frontier scale.

The reaction split predictably. Developers liked the economics: one early test had Ultra completing a physics-heavy task for roughly 5 cents versus around 57 cents on a comparable closed model, about 10x cheaper, and it landed free on OpenCode and OpenRouter on day one. The skeptics, like Fluid Coding's harness review, pushed the right question: a strong benchmark table is not the same as staying coherent across a real coding-agent loop with tools, broken logs, and 10-turn memory. The r/LocalLLaMA head-to-heads are already underway.

03.

Why we are paying attention

A frontier-scale, open-weight, agent-focused model that runs fast on NVIDIA hardware many of our clients already own is a meaningful option for on-prem AI. Open weights plus high throughput plus a million-token context is the combination that makes long-running autonomous agents deployable inside a client's own environment: no per-token vendor bill, no data leaving the building.

As always, the proof is in the harness, not the leaderboard. But Nemotron 3 moves self-hosted agentic AI from possible to practical.

Wayne Dyer

"If you change the way you look at things, the things you look at change."

Deutschland
Mittelbachstraße 66, 53518 Adenau

Tel. +49 (0) 176 74709826

© iiterate Technologies GmbH
All rights reserved
Social Links

Wayne Dyer

"If you change the way you look at things, the things you look at change."

Deutschland
Mittelbachstraße 66, 53518 Adenau

Tel. +49 (0) 176 74709826

© iiterate Technologies GmbH
All rights reserved
Social Links

Wayne Dyer

"If you change the way you look at things, the things you look at change."

Deutschland
Mittelbachstraße 66, 53518 Adenau

Tel. +49 (0) 176 74709826

© iiterate Technologies GmbH
All rights reserved
Social Links

Wayne Dyer

"If you change the way you look at things, the things you look at change."
Deutschland
Mittelbachstraße 66, 53518 Adenau

Tel. +49 (0) 176 74709826

© iiterate Technologies GmbH
All rights reserved
Social Links