Role at a glance
- Salary
- $272K – $431.3K/yr
- Location
- Santa Clara, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 15+ years of systems software development including at least 1 year dedicated to developing/exploring AI.
- Education
- BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Cloud Site Reliability Engineering Architect will join NVIDIA's global IPP Cloud Infrastructure Team and serve on the GPU Private Cloud team. The role supports infrastructure used for interactive development, centralized CI/CD, and QA testing across NVIDIA, helping optimize software development and AI development systems.
What You'll Do
- Evaluate, identify, and develop software solutions to optimize critical software development workflows.
- Architect, implement, and support end-to-end CI/CD systems using open-source and NVIDIA proprietary software.
- Onboard NVIDIA internal development teams to private cloud infrastructure by discovering use cases and available solutions.
- Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.
- Lead software development projects and technically direct a team of engineers.
- Craft and implement critical metrics using analytics methods and dashboards.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Experience maintaining cloud infrastructure and highly available production environments; strong programming skills in Java, Python, and Shell script; understanding of distributed systems and REST APIs; experience with SQL/NoSQL databases, Docker, virtual machines, and cloud technologies including OpenStack and Kubernetes; ability to work across organizational boundaries.
Required
- Cloud infrastructure and highly available production environments
- Java, Python, and Shell script
- Distributed systems and REST APIs
- SQL/NoSQL databases, including MySQL, Cassandra, MongoDB, or Elasticsearch
- Docker containers and Virtual Machines
- Cloud technologies including OpenStack, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, and Kafka
- Ability to work across organizational boundaries in a multi-national, multi-time-zone corporate environment
Preferred
- Depth in AI, Machine Learning and Deep Learning algorithms and techniques.
- Strong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.
- Experience developing large-scale software systems using modular architecture under real-time performance requirements.
- Background in designing high-performance, scalable software systems with a strong focus on hardware cost optimization.
Original job description
Content provided by the employer
Original job description
Content provided by the employer
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Autonomous Vehicles to cater to their infrastructure needs. These cloud services provide almost half a million automated jobs per day on thousands of servers helping with the efficiency of thousands of NVIDIA's software engineers worldwide. The cloud hosts various machines and devices with operating systems like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors. It delivers unified CI/CD solutions and cloud-based software development. Are you passionate about distributed infrastructure and looking for sophisticated, critical issues, ready to build the next generation of cloud services, design creative solutions, mine through data to uncover real problems and fix them?
What you'll be doing:
Serve as an SRE Architect part of GPU Private Cloud team used by thousands of NVIDIANs globally for interactive development, centralized CI/CD, and QA testing.
Evaluating, identifying and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA.
Architecting, implementing, and supporting end-to-end CI/CD system using open-source and NVIDIA proprietary software.
Customer (NVIDIA Internal development teams) onboarding to Private cloud infrastructure with a good discovery of the use case and available solutions within the cloud.
Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.
Leading software development projects and technically direct a team of brilliant engineers and guide them to provide efficient and impactful solutions.
Looking for problems within software systems and resolving the issues
Craft and implement critical metrics using various analytics methods and dashboards.
What we need to see:
BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).
15+ years of systems software development including at least 1 year dedicated to developing/exploring AI.
Experience of maintaining cloud infrastructure and highly available production environment.
Strong programming and software development skills in JAVA, Python, Shell-script along with good understanding of distributed systems and REST APIs.
Experience in working with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB or Elasticsearch.
Excellent knowledge and working experience with Docker containers and Virtual Machines.
Good background of Cloud technologies like: OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.
Ability to work across organizational boundaries effectively to improve alignment and productivity between teams in a multi-national, multi-time-zone corporate environment.
Ways to stand out from the crowd:
Depth in AI, Machine Learning and Deep Learning algorithms and techniques.
Strong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.
Experience developing large-scale software systems using modular architecture under real-time performance requirements.
Background in designing high-performance, scalable software systems with a strong focus on hardware cost optimization.
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About the company
NVIDIA
Large Enterprise
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.