Role at a glance
- Salary
- $184K – $356.5K/yr
- Location
- Santa Clara, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 10+ years in the software industry with specialization in system software and/or firmware development.
- Education
- BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience.
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Datacenter System Software team is developing fleet-scale debuggability infrastructure for NVIDIA datacenter systems. This Senior Engineer will build solutions that collect, correlate, normalize, and analyze logs across GPUs, CPUs, networking products, components, trays, and racks to help identify actionable root causes for fleet-level issues.
What You'll Do
- Architect and build fleet-wide log collection and analysis solutions across components, trays, and racks.
- Develop tooling to collect, normalize, and time-align kernel, driver, syslog, Redfish, SEL, firmware, and BMC logs through in-band and...
- Build and maintain a log catalog and taxonomy mapping raw log signatures to fault classes, severity, and remediation guidance.
- Develop debugging and root-cause tooling that produces ranked diagnoses from high-volume fleet logs while keeping production-node...
- Partner with developers, SWQA, and product engineering on logging solutions, event schemas, and log-producer contracts.
- Steward the project's open-source release, review community contributions, and represent the tooling in upstream discussions.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
10+ years in software industry, specializing in system software and/or firmware development; proven track record shipping scalable server products or fleet-wide experience; strong Python or RUST skills; deep Linux systems experience; hands-on experience with BMC, Redfish, IPMI, and SEL; skills in log parsing, normalization, structured logging, schemas, and taxonomies.
Required
- System software and/or firmware development
- Scalable server products or fleet-wide experience
- SCM such as Git or Perforce
- Jira or project-management tools
- Python or RUST
- Linux systems, including kernel and driver logs, syslog, and journald
- Out-of-band management and platform interfaces, including BMC, Redfish, IPMI, and SEL
- Log parsing, normalization, structured logging, event schemas, and taxonomies
Preferred
- Leading debuggability solutions on rack-scale compute architectures such as GB200/GB300 NVL72
- Log and telemetry analytics stacks such as OpenSearch/ELK, Loki, Prometheus, Grafana, or PagerDuty
- Time-series databases
- x86/ARM system architecture and C/C++
- Integrating AI/LLM tooling into engineering workflows
- Follow-the-sun support organizations with measurable response SLAs
- Contributing to or maintaining open-source projects
Original job description
Content provided by the employer
Original job description
Content provided by the employer
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world.
We are the Datacenter System Software team, and we are looking for a highly motivated, creative Senior Engineer o drive Fleet Scale Debuggability end to end. You will design, architect, and build infrastructure, tooling, analytics on how to collect multi-rack scale logs. The solution should normalize, correlate, and reason over logs spanning multiple components, trays, or racks including NVIDIA's GPUs, CPUs, Network products. The logs shall be fetched inband or out of band and should help triage fleet level issues seen by our customers. Your work directly shortens the path from a raw, noisy log stream to an actionable root cause. Join us at the forefront of technological advancement.
What you will be doing:
Architect, Design, build, fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
Develop tooling to collect, normalize, and time-align logs from heterogeneous sources — kernel and driver logs, syslog, Redfish event logs, SEL, firmware and BMC logs — over both in-band and out-of-band channels. Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance, so triage is repeatable rather than tribal knowledge.
Develop debug and root-cause tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults. Drive the design for collecting and analyzing logs at fleet scale while keeping overhead on production compute nodes low.
Partner with all matrixed organizations — developers, SWQA, and product engineering — in a fast-moving environment with end-to-end logging solutions, event schemas, and the contract between log producers and your tooling.
Steward the project's open-source release: keep internal and public code paths clean, review community contributions, and represent the tooling in upstream discussions.
Write design docs and own end-to-end delivery, working across teams from definition through implementation, debugging, testing, and early customer support.
Perform code reviews and partner with development and QA to strengthen unit testing, integration coverage, and test plans.
Track work through Jira and bug-management tools and build a realistic end-to-end execution plan in collaboration with other engineers and managers.
What we need to see:
10+ years in the software industry with specialization in system software and/or firmware development.
BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience.
Proven track record of shipping scalable server products or fleet-wide experience.
A self-starter who loves finding creative solutions to complicated problems, with excellent written and oral communication skills — including executive-level reporting — strong work ethic, and dedication to teamwork.
Flexibility to work and communicate effectively across teams, partners, and time zones.
Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira. Strong, demonstrable skills in Python or RUST.
Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms.
Hands-on experience with out-of-band management and platform interfaces — BMC, Redfish, IPMI, SEL — and an understanding of in-band vs. out-of-band trade-offs.
Strong skills in log parsing, normalization, and structured logging, and comfort designing schemas and taxonomies for machine-readable events.
Ways to stand out from the crowd:
Experience leading debuggability solutions on sophisticated rack-scale compute architectures like GB200/GB300 NVL72. Familiarity with log and telemetry analytics stacks (e.g., OpenSearch/ELK, Loki, Prometheus, Grafana, PagerDuty) and time-series databases.
Hands-on experience with x86/ARM system architecture and coding (C/C++, Python). Experience with SCM (Git, Perforce) and project management tools (Jira).
Track record of integrating AI/LLM tooling into engineering workflows — for triage, validation, log analysis, or test generation. Experience standing up follow-the-sun support organizations with measurable response SLAs
Experience contributing to or maintaining open-source projects, including managing the boundary between internal and public code.
NVIDIA is considered one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you are creative and autonomous, we want to hear from you!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About the company
NVIDIA
Large Enterprise
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.
NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.