NVIDIA

NVIDIA

Posted via Workday

Senior Solutions Architect, First Time Deployment Validation - NVIS

Posted Aug 5, 2026

Role at a glance

Salary
$148K – $235.8K/yr
Location
Santa Clara, California, United States
Work arrangement
On-site
Employment
Full-time
Experience
More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.
Education
Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

The Senior Solutions Architect on NVIDIA’s First Time Deployment Team validates AI factories from initial rack power-on through customer handoff. The role runs AI and LLM workloads on Linux-based GPU clusters, improves observability and automation, and partners with engineering, deployment, and customer teams to support product launches.

What You'll Do

  • Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
  • Execute AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.
  • Investigate and resolve failed, hung, or underperforming training jobs and benchmarks.
  • Build and improve observability using metrics, logs, traces, and dashboards.
  • Develop Python and Shell automation for benchmark execution, result collection, and regression checks.
  • Work with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field; more than 6+ years managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings; hands-on multi-GPU or multi-node AI/ML workload experience with NCCL; knowledge of AllReduce and AllToAll; familiarity with PyTorch or TensorFlow; Python and Shell/Bash proficiency; benchmarking experience; observability troubleshooting experience; strong communication and cross-functional collaboration skills.

Required

  • Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field
  • More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings
  • Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL
  • Solid grasp of collective communication patterns, particularly AllReduce and AllToAll
  • Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow
  • Proficiency with Python and Shell/Bash for scripting, automation, and tooling
  • Experience with benchmarking
  • Comfortable working with observability data

Preferred

  • Experience with AI factory or large-scale AI infrastructure build, deployment, or operations
  • Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments
  • Familiarity with observability stacks used for large distributed systems
  • Experience building automation and CI-style pipelines for running and validating benchmarks at scale
  • Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions

Original job description

Content provided by the employer

The First Time Deployment Team owns first-time execution of NVIDIA's latest products and systems; gathering install and bring-up evidence, operationalizing the validation process, documenting blockers and finding solutions to launch AI Factories at scale. Our results are spread across NVIDIA so we can succeed at scale.

We're looking for an ambitious Senior Solutions Architect to drive validation of NVIDIA AI factories from first rack power-on through customer handoff. You will be embedded in launches from the start, running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters using NCCL and collectives (AllReduce, AllToAll) to validate performance and scalability. When workloads or benchmarks fall short, you're the expert who digs in, partners with engineering, and drives resolution.

You will operationalize observability and automation to accelerate validation, capture structured evidence across every bring-up milestone, and work directly with internal deployment teams and external customers to ensure AI factories are ready at launch. Your work directly enables the success of NVIDIA's first external product launches!

What You Will be Doing:

  • Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
  • Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks.
  • Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.
  • Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.
  • Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
  • Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks
  • Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.
  • Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
  • Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.
  • Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.

What We Need to See:

  • Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.
  • More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.
  • Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL.
  • Solid grasp of collective communication patterns, particularly AllReduce and AllToAll, and how they are applied in contemporary ML/LLM training.
  • Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow.
  • Proficiency with Python and Shell/Bash for scripting, automation, and tooling.
  • Experience with benchmarking (crafting, executing, and interpreting performance benchmarks).
  • Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads.
  • Strong communication skills and the ability to work effectively with cross-functional teams.

Ways to Stand Out From the Crowd:

  • Experience with AI factory or large-scale AI infrastructure build, deployment, or operations.
  • Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments.
  • Familiarity with observability stacks (e.g., metrics/monitoring, logging, tracing systems) used for large distributed systems.
  • Experience building automation and CI-style pipelines for running and validating benchmarks at scale.
  • Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 18, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA

About the company

NVIDIA

Large Enterprise

NVIDIA is a leading technology company renowned for its graphics processing units (GPUs) and innovative computing solutions that enhance visual experiences across multiple platforms, including gaming, scientific research, and artificial intelligence. Founded in 1993, the company has expanded its offerings to include powerful AI frameworks and deep learning platforms, making significant contributions to industries such as gaming, data centers, automotive, and healthcare. NVIDIA's commitment to pushing the boundaries of visual computing continues to drive advancements in both hardware and software, positioning the company at the forefront of emerging technologies and digital transformation.