Role at a glance
- Job function
-
Software Engineering & IT Software Engineering
- Salary
- Not Disclosed
- Location
- Remote - California, United States
- Work arrangement
- Remote
- Employment
- Full-time
- Experience
- 6+ years building production infrastructure software or distributed systems.
- Education
- BS or MS in Computer Science, Engineering, or equivalent experience.
Spotted an issue?
We’ll check it against the original posting.
Role Summary
This role builds and maintains the supported Kubernetes runtime and software delivery for NVIDIA GPU clusters across public clouds and specialized GPU providers. The primary focus is Runtime, Release Engineering, or Provider Integration, with collaboration across the team.
What You'll Do
- Build Go controllers and APIs to install, upgrade, and validate GPU cluster software; integrate components and evolve Helm and Argo CD...
- Build validation pipelines and systems to allocate GPU capacity across validation runs, accounting for cloud reservations and quotas;...
- Bring new providers and GPU hardware into production, resolve integration failures with partner teams, and automate provisioning,...
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Required: 6+ years building production infrastructure software or distributed systems; strong programming skills in Go or another language for production systems and willingness to work primarily in Go; Kubernetes experience and depth in at least one of controllers and operators, release automation, test and validation systems, or cloud integration; experience delivering engineering projects, diagnosing complex failures, and collaborating across teams; BS or MS in Computer Science or Engineering, or equivalent experience.
Required
- 6+ years building production infrastructure software or distributed systems.
- Strong programming skills in Go or another language to build production systems, with willingness to work primarily in Go.
- Kubernetes experience and depth in at least one area: controllers and operators, release automation, test and validation systems, or...
- Experience delivering engineering projects, diagnosing complex failures, and collaborating across teams.
- BS or MS in Computer Science, Engineering, or equivalent experience.
Preferred
- Experience in any of these areas is valuable, but not required: Go development with controller-runtime, CRDs, and reconcilers.
- Release qualification across multiple environments or platforms.
- GPU infrastructure, accelerated networking, or GPU scheduling.
- Bringing new hardware, regions, or cloud providers into production.
- Resource allocation, leasing, or fair-share scheduling and upstream integration, compatibility, or software supply chain integrity.
Original job description
Content provided by the employer
Original job description
Content provided by the employer
NVIDIA researchers depend on GPU clusters for large-scale AI workloads. Our DGX Cloud Kubernetes Runtime & Release team brings those clusters to life across major public clouds and specialized GPU providers, often on hardware that is new to the world when we get it. We build and maintain the supported Kubernetes runtime, automate its delivery, and bring new providers and GPU platforms into production.
We’re growing quickly and taking on broader ownership of NVIDIA’s cluster software delivery. We’re hiring across Runtime, Release Engineering, and Provider Integration, with each role focused on your strengths. You don’t need experience across every area below.
What you’ll be doing:
Your primary focus will be one of three areas, with collaboration across the team:
- Runtime: Build Go controllers and APIs to install, upgrade, and validate GPU cluster software. Integrate components, define API contracts, and evolve Helm and Argo CD delivery toward controller-driven automation.
- Release Engineering: Build validation pipelines that inform release decisions across providers and GPU platforms. Develop systems to allocate GPU capacity across validation runs and account for cloud reservations and quotas. Make qualification more efficient through reusable tests and clear failure reports.
- Provider Integration: Bring new providers and GPU hardware into production, potentially among the first engineers working with new silicon. Resolve integration failures with partner teams and turn initial provisioning, upgrade, and operational checks into repeatable automation.
What we need to see:
- 6+ years building production infrastructure software or distributed systems.
- Strong programming skills in Go or another language to build production systems, with willingness to work primarily in Go.
- Kubernetes experience and depth in at least one area: controllers and operators, release automation, test and validation systems, or cloud integration.
- Experience delivering engineering projects, diagnosing complex failures, and collaborating across teams.
- BS or MS in Computer Science, Engineering, or equivalent experience.
Ways to stand out from the crowd:
- Experience in any of these areas is valuable, but not required:
- Go development with controller-runtime, CRDs, and reconcilers.
- Release qualification across multiple environments or platforms.
- GPU infrastructure, accelerated networking, or GPU scheduling.
- Bringing new hardware, regions, or cloud providers into production.
- Resource allocation, leasing, or fair-share scheduling and upstream integration, compatibility, or software supply chain integrity.
This role suits an engineer who wants direct influence over what reaches production, and who builds for the hundredth cluster while shipping the first. Join us and help build the next generation of NVIDIA’s GPU cloud infrastructure!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About the company
NVIDIA
Large Enterprise
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.