Role at a glance
- Salary
- $168K – $270.3K/yr
- Location
- Santa Clara, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 8+ years of hands-on experience in cluster management and related tools
- Education
- A Master's or Ph.D. in Computer Science or a related field, or equivalent experience.
Spotted an issue?
We’ll check it against the original posting.
Role Summary
This Senior Test Developer / test engineer role is part of NVIDIA's Enterprise Software QA team and focuses on designing, building, optimizing, and testing large-scale infrastructure for unified cloud services and data center offerings. The role supports cloud product quality, performance, scalability, reliability, and continuous testing across infrastructure and software layers.
What You'll Do
- Work with development teams on test plans, test execution, reviews, failure analysis, and quality and risk assessment for cloud...
- Leverage AI skills to expedite test scope, test planning, execution, and automation workflows.
- Lead NVIDIA Cloud and Data Center bring-up activities, including validation, reporting, issue debugging, design input, and coverage...
- Design, develop, and maintain CI/CD pipelines for continuous testing in cloud environments.
- Perform performance, scalability, and reliability testing of cloud services.
- Implement and maintain cloud test environments and supervise infrastructure alerts for significant events.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Master's or Ph.D. in Computer Science or a related field, or equivalent experience; experience with AI development tools for test cases, automation, code coverage, and triaging; 8+ years of hands-on experience in cluster management and related tools; 2+ years of experience with AWS, Azure, Google, or OCI Cloud; experience with network, storage, security, cluster configuration and debugging; expertise administering, operating, and configuring Kubernetes; experience with Gitlab, Jenkins, and GitOps; proficiency with Prometheus, Grafana, Cloudwatch, and Thanos; proficiency debugging networks, DHCP, DNS, HTTP, Linux, and containers.
Required
- Master's or Ph.D. in Computer Science or a related field, or equivalent experience
- Experience with AI development tools used in creating test cases, automating test cases, code coverage, triaging
- 8+ years of hands-on experience in cluster management and related tools, including Docker Containers, Slurm, Kubernetes, and Ansible
- 2+ years strong experience with cloud infrastructure platforms like AWS, Azure, Google, OCI Cloud
- Hands-on experience with network, storage, security, cluster configuration and debugging, cloud infrastructure management tools like...
- Expertise in administering, operating, and configuring Kubernetes
- Experience in CI/CD tools such as Gitlab and Jenkins and the GitOps model
- Proficiency in various monitoring tools: Prometheus, Grafana, Cloudwatch, and Thanos
Preferred
- Familiarity with "Base Command Manager" for managing and monitoring high performance computing
- Experience in writing automation for web application using tools like selenium, playwright
Original job description
Content provided by the employer
Original job description
Content provided by the employer
We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment.
What you'll be doing:
Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency
Leverage AI skills to expedite the test scope, test plan, execution and automation workflows.
Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas.
Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed.
Perform performance, scalability, and reliability testing of cloud services.
Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud.
Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability.
Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues.
Working with teams to ensure quality of the cloud products getting delivered focusing on critical areas like security, storage, workloads, performance on latest SW and FW components.
What we need to see:
A Master's or Ph.D. in Computer Science or a related field, or equivalent experience.
Experience with AI development tools used in creating test cases, automating test cases, code coverage, triaging.
8+ years of hands-on experience in cluster management and related tools, including Docker Containers, Slurm, Kubernetes, and Ansible.
2+ years strong experience with cloud infrastructure platforms like AWS, Azure, Google, OCI Cloud.
Hands-on experience with network, storage, security, cluster configuration and debugging, cloud infrastructure management tools like terraform, ansible.
Expertise in administering, operating, and configuring Kubernetes.
Experience in CI/CD tools such as Gitlab and Jenkins and the GitOps model.
Proficiency in various monitoring tools :Prometheus, Grafana, Cloudwatch, and Thanos.
Proficiency in debugging issues involving networks, DHCP, DNS, HTTP, Linux, and containers.
Ways to Stand Out from the Crowd:
Familiarity with "Base Command Manager" for managing and monitoring high performance computing.
Experience in writing automation for web application using tools like selenium, playwright.
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About the company
NVIDIA
Large Enterprise
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.