Role at a glance
- Salary
- $180K – $440K/yr
- Location
- Palo Alto, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
Spotted an issue?
We’ll check it against the original posting.
Role Summary
This role designs and operates a large-scale distributed system powering a supercomputing cluster for AI training. The work spans low-level systems performance, hardware-software-algorithm co-design, and maintaining scalable, reliable infrastructure.
What You'll Do
- Design, build, and implement a large-scale distributed system for a supercomputing cluster.
- Profile, debug, and optimize performance across GPUs, the Linux kernel, networking, and filesystems.
- Collaborate on hardware, software, and algorithm co-design for AI training.
- Maintain and innovate on the codebase to ensure scalability and reliability.
- Develop tools to enhance team productivity and streamline workflows.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Systems programming experience in C, C++, or Rust; computer systems fundamentals; hands-on expertise with Kubernetes, including cluster architecture, pod lifecycle, networking, storage, service mesh, and production-grade operations.
Required
- Systems programming experience in C, C++, or Rust
- Computer systems fundamentals
- Hands-on expertise with Kubernetes, including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and...
Preferred
- Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers
- Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives)
- Proficiency in performance analysis, profiling, and low-level optimization techniques
- Solid understanding of computer networks and the TCP/IP stack
- Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark)
- Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows
- Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel
- Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar)
Original job description
Content provided by the employer
Original job description
Content provided by the employer
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
RESPONSIBILITIES:
- Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
- Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
- Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
- Maintain and innovate on our codebase to ensure scalability and reliability.
- Develop tools to enhance team productivity and streamline workflows.
BASIC QUALIFICATIONS:
- Systems programming experience in C, C++, or Rust
- Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
- Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations
PREFERRED SKILLS AND EXPERIENCE:
- Collaborate in a fast-paced, open environment to design and foundational systems.
- Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers
- Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives)
- Proficiency in performance analysis, profiling, and low-level optimization techniques
- Solid understanding of computer networks and the TCP/IP stack
- Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark)
- Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows
- Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel
- Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar)
COMPENSATION AND BENEFITS:
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
About the company
xAI
Large Enterprise
xAI is a cutting-edge technology company focused on developing advanced artificial intelligence solutions to enhance human capabilities and optimize decision-making processes. Founded by a team of leading experts in AI and machine learning, xAI aims to address complex challenges across various industries, including healthcare, finance, and transportation. By prioritizing ethical AI development, the company is committed to creating innovative tools that empower organizations to harness the full potential of artificial intelligence while ensuring transparency and accountability.
xAI is a cutting-edge technology company focused on developing advanced artificial intelligence solutions to enhance human capabilities and optimize decision-making processes. Founded by a team of leading experts in AI and machine learning, xAI aims to address complex challenges across various industries, including healthcare, finance, and transportation. By prioritizing ethical AI development, the company is committed to creating innovative tools that empower organizations to harness the full potential of artificial intelligence while ensuring transparency and accountability.