Role at a glance
- Job function
-
Software Engineering & IT Cloud & Infrastructure Engineering DevOps & Site Reliability Engineering
- Salary
- Not Disclosed
- Location
- Santa Clara, California, United States
- Work arrangement
- Hybrid
- Employment
- Full-time
- Education
- Master's
Spotted an issue?
We’ll check it against the original posting.
About the role
Original posting provided by NVIDIA
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables unique creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.
Join our team at NVIDIA as a Senior Storage Platform Engineer responsible for designing, deploying, and operating the storage platforms that power NVIDIA's EDA FARM, software engineering, AI/ML teams, and engineering workflows at scale. You will own the full lifecycle of our multi-vendor storage infrastructure while simultaneously building the automation, pipelines, and integrations that turn storage into a scalable, self-service platform.You will drive Infrastructure as Code adoption across various Storage platforms, and ensure our storage estate is tightly integrated with CMDB, observability, configuration management, and self-service tooling.
What You'll Be Doing:
Lead end-to-end deployment of storage systems across NetApp, Pure Storage, Cloudian, and DDN, owning timelines, configuration quality, and delivery against project milestones.
Define and enforce configuration standards, baselines, and operational runbooks across all platforms; ensure every deployment is consistent, documented, and auditable.
Design and implement Infrastructure as Code frameworks (Ansible, Terraform, or equivalent) to automate provisioning, configuration management, and lifecycle operations across all storage platforms.
Build and maintain CI/CD pipelines for storage configuration deployments — ensuring every change is reviewed, tested, and rolled out in a repeatable, validated manner.
Develop and own integrations between storage platforms and the broader infrastructure ecosystem: CMDB for asset discovery and inventory sync, observability stacks for metrics and alerting, configuration management tools for drift detection, and self-service portals that let engineering teams provision storage on demand.
Manage day-2 operations including capacity planning, firmware and software lifecycle management, performance tuning, and root cause analysis.
Partner with stakeholders like chip design teams, software engineering, research and application teams to deliver integrated, fit-for-purpose storage solutions for engineering workloads
Champion GitOps practices for storage — every configuration tracked, every change reviewed, no manual snowflakes.
What we need see:
8+ years of hands-on experience with enterprise storage systems in large-scale production environments.
Deep working knowledge of at least three platforms from our stack: NetApp ONTAP (NFS/NAS), Pure Storage FlashArray or FlashBlade, Cloudian HyperStore (S3 object), DDN (Lustre/EXAScaler or high-performance NAS).
Strong command of storage protocols — NFS, SMB, iSCSI, NVMe-oF, S3, and Lustre.
Proven experience building infrastructure automation with Ansible, Terraform, or equivalent IaC tools.
Proficiency in Python, Go, or similar, with a track record of building reusable tooling rather than one-off scripts.
Experience integrating storage systems with CMDB platforms (ServiceNow, Nautobot, Cerebro or equivalent) and observability stacks (Prometheus/Grafana, Splunk, Datadog, or equivalent).
Solid experience with CI/CD platforms (GitHub Actions, Jenkins, GitLab CI, or equivalent) and Git-based workflows.
Ability to work across teams and translate infrastructure needs into platform capabilities.
MS Degree in Computer Science or equivalent experience
Way to stand out from the crowd:
You've worked in HPC, AI/ML, or large-scale research computing environments where storage is mission-critical and throughput matters.
You've built self-service storage provisioning workflows — not just automated deployments, but systems that enable other teams to move independently.
You are experienced with containerization technologies, such as Docker, Mesosphere DCOS, Kubernetes (k8s).
You've contributed to or maintained internal developer tools or platform APIs — not just consumed them.
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!
#LI-Hybrid
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About the company
NVIDIA
Large Enterprise
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.