NVIDIA
40 recent stories
Also today

NVIDIA Releases Nemotron 3.5 Lightning: 30B MoE with Only 3B Active Parameters
NVIDIA AI has released Nemotron 3.5 Lightning, a Mixture-of-Experts model with 30 billion total parameters but only 3 billion active at inference time, paired with the NeMo Switchyard model router for intelligent request routing. The low active-parameter count means inference costs are dramatically reduced compared to dense models of equivalent capacity, making it viable for production deployments where latency and cost matter. NVIDIA also released the NeMo Switchyard router alongside it, which lets developers automatically route requests to the most appropriate model in a fleet — a key primitive for multi-model agentic systems. For developers building with NVIDIA's ecosystem, this is a direct path to running capable reasoning at dense-model quality with MoE-level efficiency. The combination of a strong open MoE and a production-ready router makes this a meaningful infrastructure upgrade for teams running self-hosted inference.
NVIDIA
NVIDIA Outlines New 800V DC Power Architecture for AI Factory Scale
NVIDIA has published details on a new 800-volt DC power architecture designed specifically for AI factory deployments, addressing the growing power density demands of large-scale GPU clusters. The shift from traditional AC distribution to high-voltage DC reduces conversion losses and enables more efficient power delivery to densely packed compute racks running AI workloads. For infrastructure engineers and data center architects planning AI factory builds, this represents a meaningful design departure that affects facility planning, cabling, and UPS systems. NVIDIA is positioning this architecture as a prerequisite for operating next-generation GPU clusters at full efficiency. Developers at companies planning to build or expand private AI compute infrastructure should engage their facilities teams with this specification change early in the planning cycle.
NVIDIA
NVIDIA and Local AI Community Advance Open Source Models and Intelligent Agents with Nemotron
NVIDIA has announced collaborative efforts with the local AI community to accelerate open-source model development and intelligent agent deployment using Nemotron as the foundation. The initiative focuses on enabling developers to run capable, open-weight models locally alongside agentic frameworks, reducing dependence on cloud API calls for inference. For developers prioritizing data privacy, offline capability, or cost control, this expands the practical options for deploying performant agents without API overhead. NVIDIA's Nemotron lineup is being positioned as the open-source alternative to proprietary frontier models for agent-centric workloads. This aligns with a broader industry trend of community-driven model refinement and local inference optimization, particularly relevant for edge and enterprise deployments.
NVIDIA

NVIDIA Releases Nemotron 3.5 Lightning and NeMo Switchyard for Faster Agentic AI
NVIDIA has launched Nemotron 3.5 Lightning, a new model optimized for speed and efficiency in agentic workloads, alongside NeMo Switchyard, a framework designed to route and orchestrate AI agents across RTX and DGX hardware. The release targets developers building multi-agent systems who need low-latency inference without sacrificing task-execution quality. NeMo Switchyard specifically addresses a key pain point in agentic architectures: intelligent task routing between models and compute resources. Together, these tools lower the barrier to deploying production-grade agentic pipelines on NVIDIA hardware. Developers working on autonomous agents or complex orchestration layers should evaluate both for integration into their existing NeMo-based stacks.
NVIDIA

NVIDIA Magpie TTS Enables Low-Latency Multilingual Voice Agents with Open Weights
NVIDIA has released Magpie TTS, an open-weight multilingual text-to-speech system optimized for building low-latency voice agents with full local deployment control. The model supports multiple languages and is designed to give developers complete ownership over inference infrastructure, avoiding cloud TTS dependency and associated latency and cost penalties. Hugging Face's blog post walks through deployment patterns and integration strategies for building production voice agent pipelines. For developers building conversational AI, customer service bots, or voice-first applications, Magpie TTS offers a viable open alternative to hosted TTS APIs with controllable latency profiles. The open-weight approach also enables fine-tuning for domain-specific pronunciation, accent, or vocabulary needs.
NVIDIA

NVIDIA Releases NemotronLabs VoiceChat 11B: Open Full-Duplex Speech Model with ~450ms Turn-Taking and Live Tool Calling
NVIDIA has released VoiceChat 11B under its NemotronLabs initiative, an open full-duplex speech-to-speech model capable of natural conversational turn-taking with approximately 450ms latency. Unlike traditional pipeline-based voice systems, this model handles real-time interruptions and overlapping speech natively, making it far more suitable for natural dialogue applications. A standout feature is live tool calling during voice conversations, enabling the model to invoke external APIs or functions mid-conversation without breaking the speech flow. For developers building voice agents, customer service bots, or any real-time spoken AI interface, this represents a meaningful open-source alternative to proprietary voice APIs. The model's openness means it can be self-hosted, fine-tuned, and integrated into custom stacks without vendor lock-in.
NVIDIA

Firebird Launches CIS Region's Largest AI Factory in Armenia Powered by NVIDIA Blackwell and Rubin
Firebird has inaugurated what is described as the largest AI factory in the CIS region, located in Armenia, built on NVIDIA's latest Blackwell and Rubin GPU architectures alongside the DGX SuperPOD (DSX) platform. This represents a significant expansion of sovereign AI infrastructure into a region that has historically had limited access to frontier compute. The deployment signals growing demand for localized AI compute outside the US, EU, and East Asia — a trend with implications for data residency, latency-sensitive inference workloads, and regional model development. For developers building or deploying in the CIS region, this creates new options for on-premise or regionally hosted inference and training capacity. NVIDIA's continued role as the infrastructure backbone for new AI factories globally reinforces its position at the center of the AI compute supply chain.
NVIDIA

NVIDIA's Omniverse Open World Models Push the Frontier of Physical AI
NVIDIA has published a detailed look at open world models within its Omniverse platform, focusing on how these models advance physical AI — systems that must understand and operate within complex, unstructured real-world environments. The post details how open world modeling enables robots and autonomous agents to generalize beyond scripted scenarios to handle novel situations, a key unsolved problem in physical AI. For developers working on robotics, simulation, or embodied AI, Omniverse's open world models represent a significant infrastructure investment by NVIDIA to make physical AI training more tractable. The integration with NVIDIA's existing simulation stack means teams can potentially leverage these tools without building custom world-modeling pipelines from scratch. This is directly relevant to anyone working on autonomous systems that need to operate outside controlled environments.
NVIDIA

NVIDIA and Partners Announce U.S.-Based AI Manufacturing Push
NVIDIA has announced a major initiative with manufacturing and supply chain partners to build AI infrastructure domestically in the United States, framing it as a strategic commitment to American-made AI hardware and data center capacity. The announcement covers chip production, systems integration, and broader AI supply chain components that NVIDIA and its partners plan to localize. For developers and enterprises planning large-scale AI infrastructure investments, domestic production could reduce supply chain risk and potentially affect lead times for high-demand hardware like Blackwell GPUs. This is also strategically significant in the context of ongoing export controls and geopolitical pressure on semiconductor supply chains. The initiative positions NVIDIA to benefit from both domestic policy tailwinds and enterprise demand for supply chain resilience.
NVIDIA

AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency in AI Systems
A coalition of AI leaders, with NVIDIA among prominent contributors, has proposed the SAFE (Secure AI Framework for Evaluation) guidelines to establish a standardized approach to cybersecurity transparency across AI systems and deployments. The proposal targets the growing gap between the pace of AI deployment and the maturity of security disclosure practices, aiming to give enterprises and developers clearer expectations about what vendors should disclose regarding vulnerabilities and mitigations. For developers building production AI systems, these guidelines could soon define baseline compliance expectations, particularly in regulated industries. The initiative ties into broader efforts like the Open Secure AI Alliance and reflects mounting pressure on AI vendors to adopt security practices comparable to those in traditional software. Teams evaluating AI vendors for enterprise use should track whether their providers are aligning with SAFE as it gains adoption.
NVIDIA

NVIDIA Releases Alpamayo 2 Super, a Frontier Open Model for Autonomous Vehicles, for Commercial Use
NVIDIA has made Alpamayo 2 Super commercially available, positioning it as a frontier open model specifically designed for robotaxi and autonomous vehicle applications. The release targets the autonomous driving stack, offering developers and AV operators a production-ready, open-weight model they can integrate directly into commercial deployments. This is a significant step for the AV space, bringing frontier-class AI capabilities to an industry that has historically relied on proprietary, closed systems. Developers building on top of NVIDIA's autonomous vehicle platform can now access and customize a state-of-the-art model for their specific use cases without the restrictions of a closed license. The commercial availability lowers the barrier to entry for AV startups and enterprise fleets looking to deploy advanced AI-driven driving systems.
NVIDIA

NVIDIA Cosmos-H-Dreams Brings Real-Time Generative Simulation to Surgical Robotics
NVIDIA has released Cosmos-H-Dreams, a generative simulation model designed specifically for surgical robotics, enabling real-time synthetic environment generation for training and validating robotic surgical systems. The model is hosted and detailed on Hugging Face, making it accessible to researchers and developers working at the intersection of robotics and medical AI. Cosmos-H-Dreams addresses a critical bottleneck in surgical robotics development: the scarcity of high-quality, diverse training data from real surgical environments. By generating photorealistic and physically plausible surgical scenarios, the model allows robotic systems to be trained and stress-tested without requiring access to live operating rooms. Developers building physical AI or robotics systems should examine this as a template for using generative simulation to close the sim-to-real gap in safety-critical domains.
NVIDIA

NVIDIA Deploys Vera CPUs Alongside AI Agents to Accelerate Chip Design
NVIDIA is integrating its Vera CPUs directly into agentic AI workflows for internal chip design, using AI agents to automate and speed up the semiconductor development process. This represents a concrete production use case of agentic AI in one of the most complex engineering domains, with NVIDIA dogfooding its own hardware and AI stack. The Vera CPUs are being used in tandem with AI agents to handle tasks like design verification, layout optimization, and iterative simulation, compressing traditionally long design cycles. For developers and engineers building agentic systems, this is a high-signal example of how multi-agent orchestration with specialized hardware can deliver measurable throughput gains in expert domains. It also hints at NVIDIA's longer-term strategy of positioning Vera as an AI-native CPU platform beyond just data center inference.
NVIDIA

NVIDIA and Microsoft Launch Open Secure AI Alliance for AI Cybersecurity
NVIDIA and Microsoft have co-founded the Open Secure AI Alliance (OSAIA), a new industry coalition aimed at standardizing AI safety and security practices across the ecosystem. The alliance notably excludes OpenAI, Google, and Anthropic, signaling a distinct coalition of infrastructure and enterprise players rather than frontier model labs. The initiative focuses on open frameworks for securing AI deployments, covering model integrity, supply chain security, and adversarial threat mitigation. For developers building production AI systems, this alliance could shape emerging security standards and best practices they will need to comply with or adopt. Watching OSAIA's published frameworks will be important for teams designing secure AI pipelines.
NVIDIA

NVIDIA Uses Vera CPU to Accelerate Next-Gen Chip Design with AI-Driven EDA
NVIDIA has published details on how its Vera CPU is being deployed internally to speed up the electronic design automation (EDA) workflows used to design future CPUs and GPUs. By running AI workloads directly on Vera, NVIDIA is compressing simulation and verification cycles that traditionally take weeks. This represents a meaningful shift in how cutting-edge silicon is designed — AI is now eating into the hardware design process itself, not just inference or training pipelines. For developers building on NVIDIA hardware, this signals faster iteration cycles for future GPU generations. It also underscores the company's vertical integration strategy, where its own hardware accelerates the design of its next hardware generation.
NVIDIA

NVIDIA and Partners Outline South Korea's AI Infrastructure Roadmap at AI Summit
At a dedicated AI summit in South Korea, NVIDIA and local partners outlined a coordinated roadmap for AI infrastructure expansion across the country, covering data center buildout, model deployment, and sovereign AI development goals. The event underscores NVIDIA's strategy of embedding itself into national AI plans as a foundational hardware and software partner, extending beyond individual enterprise deals. For developers in the Asia-Pacific region, this signals meaningful investment in local compute availability and AI services infrastructure over the coming years. The partnership framework also reflects a growing trend of countries pursuing sovereign AI capacity rather than depending entirely on US-based cloud providers. Engineers and teams evaluating infrastructure for regional deployment should watch Korean cloud and compute partnerships as they formalize.
NVIDIA

Nvidia Vera Rubin: Inside the Agentic AI Factory Rewriting the CPU Playbook
A detailed analysis of NVIDIA's Vera Rubin architecture examines how it is purpose-designed for agentic AI workloads, fundamentally rethinking the role of the CPU in AI inference pipelines. Unlike prior GPU generations optimized primarily for training, Vera Rubin is framed as an 'agentic AI factory'—a system designed to handle the orchestration, memory, and throughput demands of multi-step, autonomous AI agents running continuously. The architectural shift has significant implications for how developers design agentic infrastructure, as GPU-CPU balance, memory bandwidth, and job scheduling all behave differently under sustained agentic load versus batch inference. For teams building long-running agents or high-throughput agentic pipelines, understanding Vera Rubin's design constraints and advantages will be important for hardware procurement and system architecture decisions. This piece is essential reading for ML engineers planning infrastructure for next-generation agentic deployments.
NVIDIA

NVIDIA Open Sources GPU-Accelerated Medical Physics Simulation Framework
NVIDIA has released an open-source framework for GPU-accelerated medical physics simulation, making high-performance simulation tools available to the broader research and clinical AI community for the first time. The framework enables dramatically faster simulation of radiation therapy, imaging, and related physical processes by leveraging GPU parallelism, which previously required significant proprietary infrastructure. For AI researchers and developers working in healthcare or scientific computing, this lowers the barrier to building and validating AI models trained on synthetic or simulated medical data. Open sourcing the framework also invites community contributions, benchmarking, and integration with existing medical AI pipelines and datasets. This is a concrete example of NVIDIA extending its ecosystem beyond hardware into domain-specific software infrastructure.
NVIDIA

NVIDIA AI Supercomputer Goes Live at Naval Postgraduate School
NVIDIA has brought a DGX-based AI supercomputer online at the Naval Postgraduate School, marking a significant expansion of AI compute infrastructure into defense education and research settings. The system is purpose-built for AI workloads and will support research, curriculum development, and operational AI experimentation for military and national security applications. This deployment reflects the broader trend of sovereign and institutional AI infrastructure investment, as governments seek to build independent AI capabilities rather than rely solely on commercial cloud providers. For developers and researchers working at the intersection of AI and defense, this represents a new center of gravity for applied AI work with unique operational constraints and requirements. The installation also highlights NVIDIA's continued dominance in purpose-built AI datacenter hardware for high-stakes environments.
NVIDIA

Wistron Opens Advanced NVIDIA AI Systems Manufacturing Plant in Fort Worth, Texas
Wistron has opened a new advanced manufacturing facility in Fort Worth, Texas, dedicated to producing NVIDIA AI systems — part of a broader trend of AI hardware manufacturing being reshored to the United States. This is a supply chain story with direct developer implications: increased domestic manufacturing capacity for NVIDIA systems should improve availability and reduce lead times for GPU clusters that have been constrained for years. For startups and enterprises trying to procure on-premise GPU infrastructure, this could meaningfully shorten the queue. It also signals that the AI hardware supply chain is maturing from a pure Asian manufacturing dependency to a more geographically distributed model. Developers building physical AI or on-premise inference solutions should watch whether this translates to improved hardware access timelines over the next 12-18 months.
NVIDIA

NVIDIA Spectrum-6 Networking Silicon Ships for Gigascale AI Factories
NVIDIA announced that Spectrum-6, its next-generation Ethernet networking silicon designed specifically for the Vera Rubin era, is now arriving in gigascale AI factory deployments. Spectrum-6 is built to handle the extreme east-west bandwidth demands of NVL72 and similar dense GPU cluster configurations, addressing one of the key bottlenecks in scaling distributed training and inference. This matters to developers and MLOps engineers because network fabric is increasingly the hidden constraint in multi-node training runs — a faster, lower-latency interconnect directly reduces step time and gradient synchronization overhead. The shift to Ethernet-based fabric (vs. InfiniBand) also has architectural implications for how AI infrastructure is designed and sourced. Teams planning large-scale training infrastructure should factor Spectrum-6 availability and compatibility into their hardware procurement planning now.
NVIDIA

NVIDIA Vera Rubin Launches with Industry-Leading Performance Per Watt and Lowest Token Cost
NVIDIA officially detailed the Vera Rubin GPU architecture, positioning it as the successor to Hopper and Blackwell with a focus on performance per watt and lowest cost-per-token for inference at scale. The Vera Rubin NVL72 configuration is being highlighted as the target platform for large-scale AI factory deployments, with partners already spinning up infrastructure around it. For developers and infrastructure teams, this represents the next hardware target for optimizing inference pipelines — model quantization, batching strategies, and serving frameworks will all need benchmarking against this new baseline. The emphasis on token cost reduction is directly relevant to anyone running high-volume LLM inference, where hardware efficiency directly maps to API pricing and margin. Teams building on cloud infrastructure should expect Vera Rubin-based instances to begin appearing in provider roadmaps within the next 6-12 months.
NVIDIA

Bristol Myers Squibb Builds Life Science AI Factory on NVIDIA Vera Rubin
Bristol Myers Squibb announced it is building what NVIDIA describes as the life science industry's most advanced AI factory, running on NVIDIA's Vera Rubin GPU architecture. The deployment targets drug discovery, molecular simulation, and large-scale biomedical model training—workloads that require extreme memory bandwidth and interconnect performance that Vera Rubin is specifically designed to deliver. This is a significant enterprise infrastructure signal: Vera Rubin is moving from announcement to production deployment in regulated, high-stakes scientific domains, which will accelerate the ecosystem of frameworks and tooling optimized for that architecture. For ML engineers working in biotech or adjacent fields, this signals that Vera Rubin will become the reference hardware for large-scale scientific AI within the next 12-18 months. It also validates the AI factory model—purpose-built, co-designed compute and software stacks—as the template for serious enterprise AI deployments.
NVIDIA

NVIDIA at SIGGRAPH 2026: Agentic AI and Physical Simulation Take Center Stage
At SIGGRAPH 2026, NVIDIA announced a suite of advances spanning graphics rendering, physical simulation, and agentic AI tooling, signaling a major push to position its platform as the backbone for next-generation interactive and autonomous environments. Key announcements include new capabilities in its Omniverse and simulation stack that integrate agentic workflows, enabling AI agents to operate within physically accurate virtual environments. This is significant for developers building training environments for robotics, game AI, or any system requiring grounded world models—NVIDIA is essentially productizing the sim-to-real pipeline. The agentic simulation tooling in particular could reduce the cost and complexity of generating synthetic training data at scale. Watch the SIGGRAPH session recordings closely if your work touches embodied AI, procedural content generation, or multi-agent simulation.
NVIDIA

NVIDIA Launches Cosmos 3 Edge for On-Device Physical AI and Simulation
NVIDIA released Cosmos 3 Edge, a new model in its Cosmos world foundation model family, optimized for edge deployment in physical AI and robotics simulation scenarios. Unlike its datacenter-focused predecessors, Cosmos 3 Edge is designed to run on constrained hardware, making it practical for embedded systems, autonomous vehicles, and industrial robotics at the inference edge. The model is available via Hugging Face, lowering the barrier significantly for developers to experiment without needing NVIDIA datacenter access. This is a meaningful infrastructure shift for anyone working on sim-to-real pipelines, robot learning, or digital twin applications—Cosmos 3 Edge enables local simulation and planning rather than round-tripping to a cloud endpoint. Developers building in the physical AI space should pull the weights and benchmark latency on their target hardware now.
NVIDIA

NVIDIA Vera Rubin Optimizes Intelligence-per-Dollar for Post-Training and Agentic AI Workloads
NVIDIA's blog details how the Vera Rubin architecture is specifically designed to maximize what they call 'intelligence per dollar' for post-training workloads — the compute-intensive phase covering RLHF, DPO, continued pretraining, and synthetic data generation. This framing signals a strategic shift: as base model training costs plateau, the competitive battlefield is moving to post-training efficiency and agentic inference. Vera Rubin's memory bandwidth and interconnect improvements are positioned to reduce the per-step cost of reinforcement-learning loops and multi-agent orchestration. For ML platform engineers and teams running their own fine-tuning or alignment pipelines, this has direct implications for infrastructure roadmap decisions. It also suggests NVIDIA is anticipating that agentic workloads — with their longer context windows and multi-step reasoning — will drive the next wave of GPU demand.
NVIDIA

NVIDIA Releases Nemotron 3 Embed: Open 8B Embedding Model Ranks #1 on RTEB
NVIDIA AI has released Nemotron 3 Embed, an open embedding model collection whose 8B parameter checkpoint has taken the top spot on the Retrieval Text Embedding Benchmark (RTEB). This is a significant milestone for open-source retrieval tooling, as top-ranked embedding models have historically been proprietary or closed-weight. The release includes multiple checkpoint sizes, giving developers flexibility to trade off latency and cost against quality. For anyone building RAG pipelines, semantic search, or document retrieval systems, this is a drop-in upgrade worth benchmarking immediately. The open license and competitive ranking make it a strong default choice over commercial embedding APIs for cost-sensitive or privacy-constrained deployments.
NVIDIA

NVIDIA Nemotron 3 Embed Takes #1 on RTEB, Targeting Agentic Retrieval Pipelines
NVIDIA has released Nemotron 3 Embed, an embedding model that achieved the top overall ranking on the Retrieval Text Embedding Benchmark (RTEB), which is specifically designed to evaluate models for agentic retrieval scenarios. Unlike general embedding benchmarks, RTEB stress-tests models on multi-hop retrieval, long-context passages, and tool-augmented search — capabilities critical for building reliable RAG pipelines in agentic workflows. The model is available on Hugging Face, making it immediately accessible for developers to drop into existing retrieval stacks. For teams building agents that depend on accurate, context-aware retrieval, this is a meaningful baseline shift — top RTEB performance suggests better real-world grounding compared to previously dominant models. Developers building enterprise RAG or agentic search should benchmark Nemotron 3 Embed against their current embeddings, especially on complex multi-step retrieval tasks.
NVIDIA

NVIDIA and Japan Partner on Full-Stack AI and Robotics Ecosystem Across Industries
NVIDIA announced a broad partnership with Japanese industry and government to deploy its full AI and robotics stack — spanning chips, software, simulation, and deployment frameworks — across manufacturing, logistics, and other sectors in Japan. The initiative goes beyond hardware deals, encompassing NVIDIA's Isaac robotics platform, Omniverse simulation tools, and AI Enterprise software, making it a significant ecosystem play. For developers, this signals that NVIDIA's robotics and edge AI toolchain is being validated at industrial scale, which typically accelerates SDK maturity and third-party integration support. Japan's manufacturing sector is one of the most demanding testbeds for robotics AI, meaning learnings from this deployment will likely feed back into the broader developer ecosystem. This also reinforces NVIDIA's positioning as the dominant infrastructure provider not just for training, but for end-to-end AI deployment in physical environments.
NVIDIA

NVIDIA Launches Jetson Thor T3000 and T2000 for Mainstream Robotics and Edge AI Agents
NVIDIA has announced the Jetson Thor T3000 and T2000 computers, purpose-built for robotics and edge AI agent workloads, expanding the Jetson lineup to target broader industrial and commercial deployment. The new modules are designed to run multimodal AI models and agentic pipelines locally, with significantly increased compute headroom compared to previous Jetson generations. For developers building embodied AI systems, autonomous robots, or edge inference pipelines, these represent a meaningful infrastructure upgrade — enabling more capable on-device reasoning without cloud round-trips. The announcement aligns with a broader NVIDIA push into the full robotics stack, including simulation, training, and deployment tooling. Developers working on real-time agentic systems in constrained environments should evaluate these as a viable substrate for next-generation edge deployments.
NVIDIA

NVIDIA: Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency
NVIDIA has published a detailed technical argument positioning performance-per-watt as the primary lens developers and infrastructure teams should use when evaluating AI hardware and system design choices. The piece argues that raw FLOP counts and even cost-per-token metrics are insufficient without factoring in energy consumption, which increasingly determines total cost of ownership at scale. As AI inference workloads grow, power constraints at the data center level are becoming a real bottleneck affecting how many models can run simultaneously and at what cost. For developers architecting inference infrastructure or advising on hardware procurement, this reframes the evaluation criteria away from peak throughput alone. The post also implicitly positions NVIDIA's own GPU lineup favorably in this metric, but the underlying argument about energy efficiency holds independent of vendor.
NVIDIA

NVIDIA Launches Nemotron Labs: Open Models for Enterprise and Sovereign AI Trust and Customization
NVIDIA has introduced Nemotron Labs, a new initiative and model family centered on open models designed specifically for enterprises and nation-states that require AI they can audit, control, and fine-tune without cloud dependency. The framing around sovereign AI is notable — NVIDIA is explicitly targeting governments and regulated industries that cannot send data to third-party APIs. Nemotron models are positioned as fully customizable, giving developers and enterprise teams the ability to adapt base models to proprietary data and compliance requirements. This extends NVIDIA's role from hardware provider to a full-stack AI platform company competing directly with model providers like Anthropic and OpenAI in the enterprise segment. Developers working on regulated or air-gapped deployments should evaluate the Nemotron family as a credible open alternative.
NVIDIA

NVIDIA Publishes Coding Guide to Tile-Based GPU Programming: cuTile, Triton, and Flash Attention
NVIDIA has released a detailed coding guide covering tile-based GPU programming using cuTile and Triton kernels, with specific treatment of Flash Attention as a canonical example. Tile-based programming is fundamental to writing efficient GPU kernels for transformer inference and training, as it controls how data is staged through shared memory to maximize throughput and minimize bandwidth bottlenecks. The guide bridges the gap between high-level ML framework abstractions and the hardware-level primitives that determine real-world performance, covering both NVIDIA's proprietary cuTile interface and the open-source Triton compiler. For developers working on custom inference engines, model optimization, or deploying large models at scale, understanding these primitives directly impacts latency and cost. This is a must-read resource for ML engineers who have hit the ceiling of framework-level optimization and need to write or audit custom kernels.
NVIDIA

GeForce NOW Expands with RTX 5080-Powered Toronto Servers for Cloud AI Workloads
NVIDIA has expanded its GeForce NOW cloud infrastructure with a new Toronto server cluster powered by GeForce RTX 5080 GPUs, increasing capacity and reducing latency for North American users. While primarily framed as a gaming cloud expansion, RTX 5080 hardware running in cloud servers is directly relevant to developers who use GeForce NOW's API access or NVIDIA's cloud rendering and inference services for GPU-accelerated workloads. The RTX 5080's Blackwell architecture brings improved tensor core throughput and memory bandwidth that benefit both real-time rendering and lightweight inference tasks at the edge. For developers building applications that leverage NVIDIA's cloud GPU fleet — including those using GeForce NOW as an accessible GPU-as-a-service layer — the Toronto expansion improves regional availability and potentially lowers round-trip latency for Canadian and northeastern US users. This also signals continued NVIDIA investment in distributed GPU cloud infrastructure as a complement to its datacenter H/B-series offerings.
NVIDIA

NVIDIA Nemotron Labs 3 Puzzle 75B A9B: Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput
NVIDIA's Nemotron Labs has released Puzzle 75B A9B, a compressed hybrid Mixture-of-Experts language model with 75 billion total parameters but only 9 billion active per token, achieving a reported 2.03x improvement in server throughput over its uncompressed counterpart. The hybrid MoE architecture is the key technical innovation here — combining dense and sparse routing to maximize GPU utilization while preserving model quality at scale. For infrastructure engineers and teams running self-hosted inference, a 2x throughput gain at this parameter scale is significant and could substantially reduce per-token compute costs on NVIDIA hardware. The model is likely optimized for NVIDIA's own GPU stack (H100/H200/B200), so teams not on NVIDIA infrastructure should verify compatibility before planning deployments. This release reinforces NVIDIA's strategy of moving up the stack from silicon into model architecture, making its hardware more competitive by co-designing models that exploit its memory bandwidth and interconnect strengths.
NVIDIA

NVIDIA Nemotron Achieves Benchmark-Leading Performance with LangChain Deep Agents Harness
NVIDIA's Nemotron model has demonstrated benchmark-leading results when paired with LangChain's deep agents evaluation harness, validating the open-stack approach to agent development. This is significant because the benchmark uses a real agentic harness — multi-step tool use, reasoning chains — rather than static Q&A, making the results more representative of production agent behavior. For developers building on LangChain, this provides evidence that Nemotron is worth evaluating as a backbone model for complex agent workflows. NVIDIA's push with open-stack integrations also means this isn't a closed ecosystem win — the components are composable. Engineers can pair this with the NVIDIA open data for agents release (also today) for a more complete agent training and evaluation pipeline.
NVIDIA

NVIDIA Vera CPU Gains Traction for Agentic Workloads Requiring High Single-Thread Performance
NVIDIA's Vera CPU is seeing adoption among AI infrastructure teams specifically because of its strong single-threaded performance at scale — a characteristic that matters more than many developers expect when running agentic workloads with complex orchestration logic, tool dispatching, and sequential decision-making code. Most AI infrastructure discourse focuses on GPU throughput, but the CPU bottleneck in agentic systems — where each step involves branching logic, memory lookups, and API calls — is becoming a real constraint at production scale. NVIDIA's blog highlights AI innovators who have chosen Vera specifically for this single-thread ceiling, suggesting this is an observed production pain point rather than a theoretical one. For infrastructure engineers designing systems for high-concurrency agentic deployments, this is a useful data point when evaluating CPU-GPU co-design tradeoffs. The Vera CPU is part of NVIDIA's broader push to own the full compute stack for AI, not just the GPU layer.
NVIDIA

NVIDIA Open-Sources Audex: A 30B Audio-Text LLM Built on Nemotron
NVIDIA has released Audex (Nemotron-Labs-Audex-30B-A3B), a unified audio-text large language model that integrates audio understanding directly into a text-capable backbone without degrading its language reasoning performance. The model is a 30B parameter mixture-of-experts architecture with only 3B active parameters per forward pass, making inference more practical than the parameter count suggests. A key design goal was preserving the text intelligence of the underlying Nemotron model while adding audio modality — a common failure mode in multimodal fine-tuning that NVIDIA explicitly claims to have addressed. For developers building voice assistants, transcription pipelines, or audio-grounded reasoning applications, Audex offers a production-weight open model worth benchmarking. Its release on Hugging Face makes it immediately accessible for experimentation.
NVIDIA

NVIDIA: Open Models Are Driving AI Research, Highlighted at ICML 2026
NVIDIA has published a blog post timed to ICML 2026 making the case that open models are now central drivers of cutting-edge AI research, not just convenient baselines. The piece highlights how open-weight models enable reproducibility, community-driven improvements, and rapid iteration that closed APIs cannot match. For developers, this signals that NVIDIA is institutionally invested in the open model ecosystem, which has implications for tooling, hardware optimization, and future model releases. The ICML context means this framing is being presented directly to the research community, likely influencing grant directions and academic-industry collaboration. Developers building research infrastructure or fine-tuning pipelines should note that the open model ecosystem is gaining serious institutional momentum.
NVIDIA

nvidia announces next generation H200 GPU for AI training
Nvidia unveiled the H200 GPU promising 2x performance over the H100 for large model training workloads.
Nvidia