Role at a glance
- Salary
- $140K – $270.3K/yr
- Location
- Santa Clara, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 5+ years proven experience; or master’s degree.
- Education
- Bachelor’s Degree (or equivalent experience) in a STEM (Science, Technology, Engineering, Math or Physics) field
Spotted an issue?
We’ll check it against the original posting.
Role Summary
This role joins NVIDIA’s platform SWQA team to develop and execute validation and reliability testing for HGX, DGX, and MGX server platforms across operating systems, firmware, and the CUDA software stack. The role supports automation, failure analysis, and production-quality software validation for NVIDIA platform systems.
What You'll Do
- Develop and execute platform test plans for servers, operating systems, firmware, and the CUDA software stack.
- Install and test operating systems, server firmware, and software stacks.
- Perform root cause analysis for reliability and validation test failures and drive mitigations.
- Build, develop, and debug server- and OS-level automation frameworks and tests.
- Review partner and supplier test results and prescribe additional reliability testing.
- Manage the bug lifecycle and collaborate with inter-groups to drive solutions.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Bachelor’s degree or equivalent experience in a STEM field; 5+ years of proven experience or a master’s degree. Experience with OS and server-level automation, CI/CD, DevOps, Python, Shell, Ansible, Jenkins, C/C++, Java, JavaScript, Linux, server troubleshooting, model testing, AI tools/frameworks, NLP, LLM benchmarking, and AI-assisted test development.
Required
- Bachelor’s degree or equivalent experience in a STEM field
- 5+ years of proven experience or a master’s degree
- OS and server-level automation
- CI/CD and DevOps experience
- Python, Shell, Ansible, Jenkins, C/C++, Java, and JavaScript
- Server and Linux troubleshooting and debugging
- Experience in bare-metal and KVM/VMWare/Hyper-V environments
- Model testing, AI tools/frameworks, NLP, and LLM benchmarking
Preferred
- FW, BMC/OpenBMC, network protocols, enterprise storage devices, PCIe buses and devices, IO sub-devices, CPU and memory, ACPI, UEFI, and...
- GitHub/GitLab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker
- AI-related tools, LLM, and NLP
- Experience working with NVIDIA GPU hardware
- Virtualization in Linux, KVM, and Docker orchestrated with Kubernetes
- Parallel programming, ideally CUDA/OpenCL
Original job description
Content provided by the employer
Original job description
Content provided by the employer
NVIDIA is the world leader in GPU Computing. We are passionate about markets include gaming, automotive, vision, HPC, datacenters and networking in addition to our traditional OEM business. NVIDIA is also well positioned as the ‘AI Computing Company’, and NVIDIA GPUs are the brains powering Deep Learning software frameworks, analytics, data centers, and driving autonomous vehicles. We have some of the most experienced and dedicated people in the world working for us. If you are dedicated, forward-thinking, and hard-working technical people across countries sounds exciting, this job is for you. NVIDIA is looking for an outstanding individual who thrives in a diverse work environment, has outstanding interpersonal skills and possesses a strong sense of engagement and continuous process improvement. This candidate must have enterprise server integration, strong Linux experience, reliability testing with various telemetries, scale out cluster, test plan development, track record in developing AI tools and NLP, DevOps, CI/CD experience to join our platform SWQA team.
What you’ll be doing:
Responsible for the development and execution of NVIDIA HGX/DGX/MGX platform test plan on servers, OS, FW and CUDA SW stack from design doc.
Installing and testing various systems OS, server firmware and SW stack.
Drive support for root cause analysis on reliability and validation test failures to identify root cause(s) and achieve mitigation.
Build, develop/debug server and OS level automation front-end and back-end framework and tests
Review partner and supplier test results and prescribe additional reliability testing on components, servers, and packaging as needed.
Work in an agile software development team with very high production quality standards.
Manage bug lifecycle and collaborate with inter-groups to drive for solutions.
What we need to see:
Bachelor’s Degree (or equivalent experience) in a STEM (Science, Technology, Engineering, Math or Physics) field
5+ years proven experience; or master’s degree.
Proven years of OS and server level automation, CI/CD process and DevOps experience using Python, SHELL, Ansible, Jenkins, C/C++, Java, JavaScript
Strong server and Linux(Ubuntu, RedHat, CentOS, SuSE, Fedora and etc…) troubleshooting and debugging experience in a bare-metal and KVM/VMWare/Hyper-V environment.
Good knowledge and hands-on experience in model testing, AI tools/frameworks (TensorFlow, Pytorch, Cursor and etc…), NLP and LLM benchmarking
Experience in using AI development tools for test plans creation, test cases development and test cases automation
Strong experience in FW, BMC/OpenBMC, Network protocol, internal/external enterprise storage devices, PCIe buses and devices, IO sub-devices, CPU and memory, ACPI, UEFI spec, Redfish - huge plus
Proven years of experience in GitHub/Gitlab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker) – huge plus
Ways to stand out from the crowd:
AI related tools, LLM and NLP.
Experience working with NVIDIA GPU hardware is a strong plus.
Good to have solid understanding of virtualization in Linux (KVM, Docker orchestrated with Kubernetes)
Background in parallel programming ideally CUDA/OpenCL is a plus
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About the company
NVIDIA
Large Enterprise
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.