Role at a glance
- Salary
- $152K – $287.5K/yr
- Location
- Santa Clara, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 5+ years of industry software engineering experience
- Education
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent experience
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Senior Software Engineer will join NVIDIA’s DGX Cloud / Fleet Intelligence team to build backend systems for GPU health monitoring, telemetry ingestion, operational automation, and fleet visibility. The role supports high-performance GPU infrastructure through scalable services, APIs, data systems, and cloud operations tooling.
What You'll Do
- Design and develop Go backend services, REST APIs, and data models for GPU Health and Fleet Intelligence.
- Build customer-facing and agent-facing APIs for compute zones, node groups, nodes, alerts, events, metrics, inventory, attestation,...
- Develop high-volume ingestion and persistence paths for in-band agents and out-of-band collectors.
- Work with Aurora PostgreSQL, Kafka/MSK, S3, SQS, Prometheus/AMP, and OpenTelemetry-based observability.
- Build and operate scheduled backend services for liveness tracking, alerting, rollups, cleanup, notification delivery, attestation, and...
- Collaborate with agent, infrastructure, SRE, UI, and cloud operations teams to turn operational workflows into scalable backend systems.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Strong Go backend development; REST APIs and production services; PostgreSQL schema design, query tuning, transactions, migrations, and operational data modeling; distributed systems, event pipelines, telemetry ingestion, or operational analytics; cloud infrastructure, Docker, Kubernetes, and Helm; debugging across APIs, databases, cloud services, and production systems; authentication and authorization patterns; Linux-based development and production environments.
Required
- Strong Go backend development experience
- Experience building REST APIs and production services using frameworks such as Gin or similar
- Strong PostgreSQL experience, including schema design, query tuning, transactions, migrations, and operational data modeling
- Experience with distributed systems, event pipelines, telemetry ingestion, or operational analytics
- Experience with cloud infrastructure, Docker, Kubernetes, and Helm
- Strong debugging skills across APIs, databases, cloud services, and production systems
- Familiarity with authentication and authorization patterns such as JWT, service account keys, trusted-edge proxies, or customer-scoped APIs
- Experience with Linux-based development and production environments
Preferred
- Background with telemetry, monitoring, observability, health, or fleet-management platforms
- Experience with AWS services such as MSK, Aurora, S3, SQS, AMP, or CloudWatch
- Experience with OpenTelemetry, Prometheus, LightStep, Grafana, or production tracing/metrics systems
- Background with NVIDIA GPUs, DGX systems, DCGM, XID analysis, attestation, or AI datacenter operations
- Experience designing APIs and storage systems that support real-time operational workflows at fleet scale
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPUs act as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent.
As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA is widely recognized as one of the most desirable employers, with some of the most versatile people in the world working for us. If you're passionate about building scalable, efficient backend systems to power cloud operations, we invite you to join our team.
We are looking for a Senior Software Engineer to join our DGX Cloud / Fleet Intelligence team and build backend systems that power GPU health monitoring, telemetry ingestion, operational automation, and fleet visibility for NVIDIA’s high-performance GPU infrastructure.
What You'll Be Doing:
This role focuses on backend services for GPU health, Fleet Intelligence, telemetry ingestion, inventory, attestation, alerting, reporting, and cloud operations automation, in addition, you will:
Design and develop Go backend services, REST APIs, and data models for GPU Health and Fleet Intelligence.
Build customer-facing and agent-facing APIs for compute zones, node groups, nodes, alerts, events, metrics, inventory, attestation, retention policies, and reports.
Develop high-volume ingestion and persistence paths for in-band agents and out-of-band collectors.
Work with Aurora PostgreSQL, Kafka/MSK, S3, SQS, Prometheus/AMP, and OpenTelemetry-based observability.
Build and operate scheduled backend services for liveness tracking, alerting, rollups, cleanup, notification delivery, attestation, and XID analysis.
Optimize database schemas, partitioned time-series storage, query performance, CTE-heavy queries, and connection pooling for reliable service behavior.
Collaborate with agent, infrastructure, SRE, UI, and cloud operations teams to turn operational workflows into scalable backend systems.
Improve service reliability, security, observability, testing, and deployment quality across Docker, Kubernetes, Helm, and cloud environments.
What We Need To See:
5+ years of industry software engineering experience with a Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent experience.
Strong Go backend development experience.
Experience building REST APIs and production services using frameworks such as Gin or similar.
Strong PostgreSQL experience, including schema design, query tuning, transactions, migrations, and operational data modeling.
Experience with distributed systems, event pipelines, telemetry ingestion, or operational analytics.
Experience with cloud infrastructure, Docker, Kubernetes, and Helm.
Strong debugging skills across APIs, databases, cloud services, and production systems.
Familiarity with authentication and authorization patterns such as JWT, service account keys, trusted-edge proxies, or customer-scoped APIs.
Experience with Linux-based development and production environments.
Ways To Stand Out From The Crowd:
Background with telemetry, monitoring, observability, health, or fleet-management platforms.
Experience with AWS services such as MSK, Aurora, S3, SQS, AMP, or CloudWatch.
Experience with OpenTelemetry, Prometheus, LightStep, Grafana, or production tracing/metrics systems.
Background with NVIDIA GPUs, DGX systems, DCGM, XID analysis, attestation, or AI datacenter operations.
Experience designing APIs and storage systems that support real-time operational workflows at fleet scale.
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About the company
NVIDIA
Large Enterprise
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.