NVIDIA

NVIDIA

Posted via Workday

Director, Software Engineering – Cloud Compute and Infrastructure

Posted Oct 10, 2026

Role at a glance

Job function
Software Engineering & IT Cloud & Infrastructure Engineering DevOps & Site Reliability Engineering
Salary
Not Disclosed
Location
Santa Clara, California, United States
Employment
Full-time

Spotted an issue?

We’ll check it against the original posting.

Log in to report

About the role

Original posting provided by NVIDIA

View original

NVIDIA is seeking an engineering director to lead the software teams behind our bare-metal datacenters and cloud compute infrastructure. The organization supports hundreds of megawatts of datacenter capacity already online, with more capacity coming. You will build the software that makes this growing physical infrastructure available as reliable, secure, and efficient compute services for NVIDIA engineering.


Our team provides NVIDIA’s continuous integration (CI) infrastructure: the environment where hardware, firmware, drivers, networking, and system software are integrated and tested on their path to production. You will enable engineering teams to bring up preproduction systems, reproduce failures, validate changes, and move platforms toward production readiness. The fleet spans multiple hardware generations and maturity levels: x86 and Arm servers, GPUs, DPUs, high-speed networking, storage, and rack-scale, liquid-cooled systems. Platforms such as Grace Blackwell and Vera Rubin illustrate the breadth of compute and interconnect technology involved. This role combines hands-on systems judgment with leadership of the teams making that diversity manageable at scale.


What you’ll be doing:

  • Lead software engineering teams responsible for cloud compute, bare-metal fleet management, and CI infrastructure, owning architecture, implementation, deployment, and operations.
  • Set the technical direction for compute control planes, resource provisioning, topology-aware placement, reservations, and capacity management across physical servers, virtual machines, and containers.
  • Automate the bare-metal lifecycle: hardware discovery and inventory, server bring-up, imaging, firmware and driver configuration, out-of-band management, health checks, reprovisioning, and recovery.
  • Build CI services that allocate the right hardware, provision repeatable test environments, record hardware and software configurations, collect diagnostics, and restore systems to a known state. Shorten the path from a hardware or software change to actionable validation results.
  • Keep engineering services dependable while supporting evolving preproduction hardware and software. Define service-level objectives, isolate failures, improve observability, and turn incidents and recurring test-infrastructure failures into engineering fixes.
  • Measure and improve hardware onboarding time, provisioning speed, CI queue time, productive fleet utilization, recovery time, and cost efficiency as capacity expands.
  • Set engineering standards for design and code reviews, automated testing, secure development, release quality, and safe changes to production systems.
  • Hire, coach, and retain engineers and engineering managers. Establish clear ownership, develop technical leaders, and build teams that deliver consistently over multiple release cycles.
  • Partner with hardware, firmware, drivers, networking, storage, security, and validation teams to onboard platforms and diagnose failures across system boundaries. Translate engineering users’ needs into clear priorities and technical decisions.
  • Work with datacenter operations and facilities teams to bring additional capacity online. Connect rack power, cooling, physical topology, and hardware-health telemetry to provisioning, placement, serviceability, and operational readiness.

What we need to see:

  • 15+ overall years of experience in software engineering, distributed systems, or cloud infrastructure, including 7+ years leading engineering teams and experience managing engineering managers.
  • Direct engineering ownership of the underlying services of a public or private compute cloud, such as AWS EC2, Google Compute Engine, Azure Compute, OCI Compute, or a comparable infrastructure-as-a-service platform.
  • Strong technical depth in distributed systems, Linux, virtualization, containers, and the networking and storage services that support large compute fleets.
  • Experience building software and automation for bare-metal infrastructure, including server provisioning, hardware inventory, firmware or operating-system lifecycle management, and recovery across heterogeneous systems.
  • Experience running production infrastructure with demanding availability requirements, including failure isolation, incident response, observability, and reliable deployment and recovery mechanisms.
  • The ability to guide architecture, evaluate implementation choices, and debug complex interactions among hardware, firmware, drivers, operating systems, and distributed services with senior engineers.
  • A record of sustained ownership as platforms evolve, with measurable improvements in reliability, scalability, or engineering delivery. Strong communication, collaboration, and people-development skills.
  • A degree in computer science, computer engineering, or a related discipline, or equivalent experience.

Ways to stand out from the crowd:

  • Experience building or operating OpenStack infrastructure, particularly Nova, Neutron, Cinder, or Ironic, or contributing to related open-source projects.
  • Engineering leadership spanning compute control planes and site reliability, including multi-tenant isolation, scheduling, placement, and capacity allocation.
  • Experience with GPU clusters, rack-scale computing, NVLink, InfiniBand or high-speed Ethernet, and AI or high-performance computing workloads.
  • Experience supporting preproduction hardware, new product introduction, or hardware-in-the-loop CI, including repeatable validation environments and systematic regression isolation.
  • Familiarity with liquid-cooled, high-density datacenters and how power, thermal constraints, cooling, and component health affect fleet availability and scheduling and delivery of compute platforms across multiple regions, datacenters, or hybrid-cloud environments, including fleet upgrades and workload migration with minimal service disruption.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 13, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA

About the company

NVIDIA

Large Enterprise

NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.