Microsoft at NVIDIA GTC: AI Agents, Physical AI, and More
AI / Machine Learning | 6 min read
At NVIDIA GTC 2026 — held in San Jose, California from March 16–19 — Microsoft used the world's largest AI developer conference to make its most significant enterprise AI announcements to date. Spanning expanded Microsoft Foundry capabilities, next-generation Azure AI infrastructure, and a deepening collaboration on Physical AI, the announcements signal Microsoft's intent to compete at every layer of the AI stack — from GPU-accelerated data centres all the way to intelligent systems operating on factory floors and beyond.
"Whether powering always-on agents, scaling next-generation AI infrastructure or deploying intelligent systems in factories, energy facilities and sovereign environments, Microsoft and NVIDIA are helping customers move faster from insight to action."
— Yina Arenas, Corporate Vice President, Microsoft Foundry
Microsoft Foundry: The Operating System for Enterprise AI Agents
At the core of Microsoft's GTC announcements is Microsoft Foundry — positioned as the operating system for building, deploying, and operating AI at enterprise scale. Built on Azure, Foundry pulls together models, tools, data, and observability into a single environment designed specifically for production-grade agents rather than experimentation.
The Foundry Agent Service and Observability in Foundry Control Plane are now generally available — enabling organisations to build and operate AI agents that can reason, plan, and act across tools, data, and workflows at production scale. The updated Microsoft Foundry portal at ai.azure.com also reached general availability, consolidating agent building, model access, and governance into a single interface. Built on OpenAI's Responses API, the service ships with production-ready SDKs in Python, JavaScript, Java, and .NET — and supports both single- and multi-agent configurations.
Foundry Control Plane delivers continuous evaluation of agent performance — including out-of-the-box evaluators for coherence, relevance, and safety, as well as custom LLM-as-a-judge evaluators for business-specific standards and end-to-end execution tracing for production debugging. NVIDIA Nemotron models are now also available through Microsoft Foundry, joining the platform's wide selection of frontier, reasoning, and open-weight models — with a partnership with Fireworks AI enabling customers to fine-tune open-weight models into low-latency assets deployable to the edge. Real-world deployments are already underway: Corvus Energy is using Foundry to replace manual inspection workflows with agent-driven operational intelligence across its global fleet.
Azure AI Infrastructure: Vera Rubin and Hundreds of Thousands of Grace Blackwell GPUs
To power the inference-heavy workloads that agentic AI demands, Microsoft has been aggressively scaling its hardware infrastructure. The company has already deployed hundreds of thousands of liquid-cooled NVIDIA Grace Blackwell GPUs across its global data centres in under a year — within facilities optimised for power, cooling, networking, and rapid generational upgrades.
Azure was also announced as the first hyperscale cloud provider to power on the new NVIDIA Vera Rubin NVL72 systems in its labs — NVIDIA's next-generation architecture featuring seven new chips, five rack-scale designs, and a complete AI supercomputer platform delivering 10x performance-per-watt gains over prior generations. Vera Rubin NVL72 systems will roll out globally across Azure over the coming months, extending accelerated AI capabilities to a new tier of inference performance.
Microsoft's infrastructure innovation also extends to sovereign and regulated environments, with initial support for the NVIDIA Vera Rubin platform on Azure Local — giving organisations control of both where AI runs and how it evolves, while maintaining Azure-consistent operations, governance, and security through Azure Arc and Foundry Local.
Physical AI: From Digital Twins to Real-World Operations
Beyond cloud infrastructure, Microsoft and NVIDIA are sharpening their collaboration on Physical AI — intelligent systems that sense, simulate, and act in the real world. At GTC, this work centres on the NVIDIA Physical AI Data Factory Blueprint, with Microsoft Foundry serving as the platform for hosting and operating Physical AI systems at Azure's cloud scale.
Microsoft is introducing a public Azure Physical AI Toolchain GitHub repository — plugging directly into NVIDIA's Physical AI Data Factory and core Azure services — giving developers a repeatable way to build, train, and operate robotics and Physical AI workflows that link physical assets, high-fidelity simulation, and cloud training environments into enterprise-grade pipelines. The integration between Microsoft Fabric and NVIDIA Omniverse libraries is also deepening, bringing together live operational data, physically accurate digital twins, and simulation in a single unified loop — targeted at manufacturing, energy, and other industries moving from monitoring dashboards to AI-driven action across physical systems.
Voice AI and Security Integrations
Rounding out the GTC announcements, Microsoft is introducing Voice Live API integration with Foundry Agent Service in public preview — enabling developers to build voice-first, multimodal, real-time agentic experiences. The refreshed Foundry portal also adds deeper integrations with Palo Alto Networks' Prisma AIRS and Zenity, delivering richer diagnostics for builders and runtime security for compliance teams — ensuring governance across the entire agent lifecycle from development through production.
Key Takeaways
- • Microsoft Foundry Agent Service and Observability in Foundry Control Plane are now generally available — enabling production-grade AI agents that reason, plan, and act at enterprise scale.
- • NVIDIA Nemotron models are now available through Microsoft Foundry, with Fireworks AI enabling fine-tuning of open-weight models into low-latency, edge-deployable assets.
- • Azure is the first hyperscale cloud to power on NVIDIA Vera Rubin NVL72 systems — delivering 10x performance-per-watt gains — and has deployed hundreds of thousands of liquid-cooled Grace Blackwell GPUs in under a year.
- • Microsoft Fabric and NVIDIA Omniverse integration creates a unified loop connecting live operational data, digital twins, and simulation — targeting manufacturing, energy, and physical operations.
- • Voice Live API integration, Palo Alto Networks Prisma AIRS, and Zenity security integrations complete a full-stack agent platform spanning development, deployment, voice, and runtime governance.
