NVIDIA unveils at SIGGRAPH: integrating Agentic and Physical AI into a single technology stack

At SIGGRAPH 2026, NVIDIA packaged open models, real-time simulation, robotics, and graphics rendering into a comprehensive physical AI narrative, with Nemotron 3 and Cosmos 3 serving as its two main pillars.
NVIDIA Unveils at SIGGRAPH: Merging Agentic and Physical AI into One Tech Stack
At the SIGGRAPH 2026 event on July 20, NVIDIA didn’t talk about GPU specs anymore—instead, it focused entirely on two things: Agentic AI and Physical AI. Research leads Neil Ashton, Edward Liu, and Ming-Yu Liu took turns presenting across domains spanning open models, real-time simulation, rendering pipelines, and robot training—all giving one clear impression: Jensen Huang wants to redefine graphics as the simulation foundation for physical AI.
This wasn’t a one-off launch, but more of a strategic recalibration. For the past two years, NVIDIA’s AI narrative has had a hidden pain point: while GPUs sold like crazy, at the model layer, the closed-source camp included OpenAI, Anthropic, and Google, and the open-source camp had Meta, DeepSeek, and Qwen—while Nemotron barely made a splash. This SIGGRAPH combo marks the first time NVIDIA clearly articulated the chain linking open models, physical simulation, and robotics.

Nemotron 3 Ultra & Omni: Open Models with Real Momentum
Let’s start with the Agentic side. This time NVIDIA unveiled Nemotron 3 Ultra and Nemotron 3 Omni together, fine-tuned using LangChain’s Deep Agents Harness. Officially, they aim to “deliver excellent open-model agent performance at a lower cost”—in other words: you can run long-lived agents without paying for Claude Opus.
Key signals to note:
- Deep Agents Harness: Essentially LangChain’s scheduling framework for long-context, multi-step planning scenarios. Nemotron 3 Ultra is currently its recommended host model. For teams building long-task agents—say, “run overnight and finish a pull request”—this combo is well worth a try.
- Omni version: Multimodal, supporting all-input types—visual, audio, and text. NVIDIA demoed how it fits into designer workflows—give it a sketch, and it can infer layout logic and generate a 3D scene.
- Enterprise endorsements: Baseten, Cognition, DeepInfra, Together AI, and Cursor were all name-dropped as ecosystem adopters. Cursor’s inclusion is particularly telling—it suggests Nemotron’s coding capabilities are competitive enough for mainstream IDEs.
Worth noting: the Nemotron 3 series is fully open—weights, data recipes, and post-training scripts will all be released. Unlike Llama 4’s “open but restricted” approach, NVIDIA’s openness here even surpasses Meta’s. For developers, OpenAI Hub now supports Nemotron 3 Ultra and Omni, letting you compare them with GPT, Claude, Gemini, and DeepSeek using the same key—no need to spin up your own inference cluster.
Cosmos 3: The Real Barrier in Physical AI Isn’t the Model—it’s the Simulation
If Nemotron is the standard card, Cosmos 3 is the real highlight.
NVIDIA’s positioning of Cosmos 3 is direct: “reason, generate, and act to advance physical AI.” These three verbs reflect the capabilities of a world model—understanding physical scenes, generating plausible futures, and controlling real-world agents. That sets it apart from world models focused solely on video generation (like Sora or Genie).
Key technical advancements include:
- Real-time simulation: The demo showcased Cosmos 3 achieving near real-time, physics-level simulation, including fluids, soft bodies, and rigid-body collisions. This is a massive leap for robot training—previously, dense contact simulations in Isaac Sim could burn GPUs overnight; now it’s cut to minutes.
- Connected to the graphics pipeline: Cosmos 3 directly integrates with OpenUSD and the RTX rendering stack, meaning scenes built in Omniverse can flow seamlessly into world models for data augmentation. The value isn’t in single-point performance but in the closed loop—design, simulation, training, and deployment all running on the same USD asset.
- Actionable output: Cosmos 3 outputs not just pixels but action sequences that robot controllers can consume directly. This is the fundamental differentiator between NVIDIA’s approach and vision-only models like Google’s Genie line.

A practical example: Toyota expanded its collaboration with NVIDIA, integrating Physical AI into next-gen L2++ vehicles and factory robots. In this partnership, Cosmos 3’s job is generating corner-case data for autonomous driving—rare scenarios like sudden pedestrian crossings can be synthesized hundreds of times per hour. This synthetic data capability offers an alternative to the massive real-world data pipelines of Waymo and Tesla.
The Graphics Legacy, Reinvented
Let’s not forget—SIGGRAPH is a graphics conference at heart. NVIDIA didn’t ignore its roots—DLSS 4.5 and a suite of rendering updates also launched.
What’s interesting is the narrative shift: DLSS was once “AI for game image upscaling.” Now it’s framed as “neural rendering as the front-end of physical AI.” Technologies like RTX Neural Shaders and Neural Radiance Cache are reframed as the “visual cortex” powering physical AI perception.
This isn’t a semantic trick. For example, in robot training, the most expensive component isn’t neural nets—it’s rendering. The simulator must produce visuals so realistic that the sim-to-real gap stays minimal. Traditional methods rely on rasterization (fast but fake) or path tracing (real but slow). DLSS 4.5’s neural path guiding achieves “path-tracing quality at rasterization speed,” cutting Cosmos 3’s training data cost by an order of magnitude.
That explains why NVIDIA talks about agents and robots at SIGGRAPH—graphics is no longer the endpoint, but a middle layer within physical AI.
Vera Rubin DSX and the Infrastructure Layer Behind the Scenes
Another under-the-radar update: Vera Rubin DSX AI Factory Reference Design. Japan’s Data Center Japan will be the first to deploy a national-scale AI infrastructure built on DSX, exceeding 140 MW and open to all developers.
In the context of physical AI, this news serves as a supporting piece—to train world models like Cosmos 3 or large models like Nemotron 3 Ultra, you need DSX-level compute capacity. NVIDIA’s tactic is clear: models are open, compute is monetized—exactly like AWS once made S3 free while charging for EC2.
Straightforward Takeaways
After this SIGGRAPH, several takeaways stand out:
- Nemotron 3 is worth trying but won’t beat Claude: In terms of cost-to-performance for open models, the Ultra version is competitive for long-run agent tasks. But pure reasoning depth and coding ability still lag behind top closed models. Think “good enough and affordable,” not “best.”
- Cosmos 3 is the real technical breakthrough: NVIDIA’s fusion of graphics and robotics stacks gives it two extra moats over pure AI companies. This will reshape the cost structure of robot training in the next two years.
- Physical AI is NVIDIA’s second growth curve: The LLM inference boom will likely peak by 2027; physical AI (robotics, autonomous driving, industrial simulation) is the next trillion-dollar market. This SIGGRAPH signals NVIDIA’s formal entry.
- Developer perspective: If you’re building agents, Nemotron 3 + LangChain Deep Agents deserves evaluation; if you’re working on robotics or embodied intelligence, Cosmos 3 + Isaac Sim is almost a must-have stack.
SIGGRAPH transformed from a graphics conference to an AI conference in less than three years. With this keynote, NVIDIA decisively sealed that shift—the next chapter of graphics is Physical AI.
References
- NVIDIA Official Blog: SIGGRAPH 2026 Coverage — NVIDIA’s official post on Agentic and Physical AI at SIGGRAPH 2026
- NVIDIA China Official Site — Product pages and keynote access for Nemotron 3 Ultra, Cosmos 3, and Vera Rubin DSX


