Role at a glance
- Salary
- $200K – $322K/yr
- Location
- 2 Locations, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 12+ years of experience designing and building distributed systems and cloud infrastructure
- Education
- BS or MS in Computer Science, Engineering, or a related field, or equivalent experience.
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Customer Success Engineer joins NVIDIA's DGX Cloud organization, bridging customer success and cloud infrastructure engineering for internal research and product teams. The role provides architectural guidance and hands-on solutions while contributing to tooling and GPU capacity management strategy.
What You'll Do
- Design and implement distributed cloud infrastructure across compute, storage, networking, and GPU capacity management
- Partner with internal research and product teams to understand workloads and provide architectural guidance
- Contribute code and codify working patterns into tools, playbooks, and reusable building blocks
- Build and maintain agentic tooling to automate operational workflows and infrastructure resource management
- Analyze customer demand and future capacity needs to drive infrastructure efficiency initiatives
- Present technical roadmaps, architecture decisions, and demos to internal stakeholders and NVIDIA leadership
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
BS or MS in Computer Science, Engineering, or a related field, or equivalent experience; 12+ years designing and building distributed systems and cloud infrastructure; GPU capacity management for high-performance computing; production coding in Golang, Java, C, C++, Python, or Rust; Kubernetes and/or distributed task scheduling; Infrastructure, Networking, Storage, and DevOps scripting/tooling; AI/ML workloads at scale; cross-functional communication and consensus-building.
Required
- BS or MS in Computer Science, Engineering, or a related field, or equivalent experience
- 12+ years of experience designing and building distributed systems and cloud infrastructure
- Demonstrated experience in GPU capacity management for high-performance computing
- Production coding in Golang, Java, C, C++, Python, or Rust
- Experience with Kubernetes and/or distributed task scheduling
- Strong background in Infrastructure, Networking, Storage, and DevOps scripting/tooling
- Experience deploying AI/ML workloads at scale
- Strong communication and relationship-building skills
Original job description
Content provided by the employer
Original job description
Content provided by the employer
The DGX Cloud organization bridges customer success and cloud infrastructure engineering, partnering directly with NVIDIA's internal research and product teams to accelerate AI workload development. As a Customer Success Engineer, you'll embed deeply with internal customers — gaining a thorough understanding of their applications and translating that knowledge into architectural guidance, best practices, and hands-on solutions. This role sits at the unique intersection of solutions architecture and platform strategy: you'll write code, build tooling, and help shape NVIDIA's GPU capacity management from the inside. Working across Engineering, Product, Finance, and Operations, you'll connect infrastructure roadmaps to business needs in a way that directly influences how NVIDIA's most advanced AI teams move faster. If you thrive where deep technical work meets high-stakes collaboration, this role was built for you.
What you’ll be doing:
Design and implement distributed cloud infrastructure at scale — spanning compute, storage, networking, and GPU capacity management across IaaS, PaaS, and SaaS models zendesk
Partner with internal research and product teams to understand workloads from both a technology and business perspective, providing architectural guidance that drives their success
Contribute code directly when needed to move projects forward, and codify working patterns into tools, playbooks, and building blocks that others can reuse
Build and maintain agentic tooling to automate operational workflows and infrastructure resource management
Analyze the DGX Cloud ecosystem to understand current customer demand and future capacity needs, driving infrastructure efficiency initiatives in partnership with Engineering, Finance, and Product
Present technical roadmaps, architecture decisions, and demos to internal stakeholders and NVIDIA leadership, driving cross-functional consensus on infrastructure strategy
What we need to see:
BS or MS in Computer Science, Engineering, or a related field, or equivalent experience.
12+ years of experience designing and building distributed systems and cloud infrastructure, with demonstrated experience in GPU capacity management for high-performance computing
Demonstrated ability to write production code in Golang, Java, C, C++, Python, or Rust
Experience with Kubernetes and/or distributed task scheduling
Strong background in Infrastructure, Networking, Storage, and DevOps scripting/tooling
Experience deploying AI/ML workloads at scale
Strong communication and relationship-building skills, with a demonstrated ability to drive cross-functional consensus and align stakeholders across departments
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and dedicated people in the world working for us.
#LI-Remote
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About the company
NVIDIA
Large Enterprise
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.