Role at a glance
- Salary
- $180K – $440K/yr
- Location
- Palo Alto, California, United States Seattle, Washington, United States
- Work arrangement
- On-site
- Employment
- Full-time
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Platform Core team develops the lowest layers of the software stack for AI supercomputers, interfacing directly with core hardware and the network fabric. This role focuses on building reliable Linux-based platform software that exposes the capabilities of supercomputer networking and compute hardware.
What You'll Do
- Build and maintain the Linux-based operating system underpinning the supercomputer network fabric.
- Write and maintain high-performance device drivers for network and compute hardware.
- Manage deployment, operation, and debugging of the platform in production.
- Resolve and root-cause anomalies to improve system reliability.
- Develop and improve observability and configuration management tools.
- Collaborate with hardware teams and external partners on next-generation hardware and bring up the OS and software stack.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Hands-on systems programming experience in C or C++; strong operating systems fundamentals; understanding of computer networking; excellent written and verbal communication skills.
Required
- Hands-on systems programming experience in C or C++
- Strong operating systems fundamentals, including scheduling, memory management, and I/O
- Understanding of computer networking in the host OS stack and common network devices such as switches and routers
- Excellent written and verbal communication skills
Preferred
- Deep knowledge of the Linux kernel and its networking stack
- Proficiency with debugging tools spanning kernel, network, and userspace, including ftrace, perf, Wireshark/tcpdump, eBPF, and gdb
- Experience bringing up new hardware from scratch
- Track record of developing and optimizing device/peripheral drivers
- Understanding of peripheral hardware interface standards including Ethernet, PCIe, I2C, and SPI
- Experience with product security, OS/system hardening, and secure boot schemes
Original job description
Content provided by the employer
Original job description
Content provided by the employer
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
At SpaceXAI, we design, build, and operate the largest AI supercomputers in the world, from the ground up. As a member of the Platform Core team, you will be responsible for the lowest layers of the software stack on the devices setting the pace for how fast we can push the frontier of AI. Our team’s software interfaces directly with the core hardware of the supercomputer and its network fabric. Our mission is to expose the full capability of that hardware — limited only by physics — to the rest of the stack, all while retaining the smallest, most robust footprint possible.
RESPONSIBILITIES:
- Build and maintain the lean, high-reliability Linux-based operating system that underpins our supercomputer network fabric.
- Write (or rewrite) high-performance device drivers to extract the maximum physical capability from our network and compute hardware.
- Manage deployment, operation, and debugging of our platform in production, including:
- Rigorous resolution and root-causing of anomalies to ensure the system is constantly improving in reliability.
- Developing and constantly improving our tools for observability and configuration management.
- Collaborate with our hardware teams and external partners to design the next generation of supercomputer hardware – and then bring up our OS and software stack on it.
BASIC QUALIFICATIONS:
- Hands-on systems programming experience in C or C++.
- Strong operating systems fundamentals (scheduling, memory management, I/O) — you must be able to articulate in extreme detail how a computer works, from the hardware all the way up to high-level applications.
- Understanding of computer networking, both in the host (OS stack) as well as common network devices like switches and routers.
- Excellent written and verbal communication skills.
PREFERRED SKILLS AND EXPERIENCE:
- Deep knowledge of the Linux kernel and its networking stack — transferable experience with another production-grade kernel is okay, too!
- Proficiency with debugging tools spanning kernel, network, and userspace (e.g., ftrace, perf, Wireshark/tcpdump, eBPF, gdb, and more).
- Experience bringing up new hardware from scratch — you have flashed bare-metal software on a microcontroller with JTAG, or bootstrapped a Linux kernel and rootfs on new hardware.
- Demonstrated track record of developing and optimizing device/peripheral drivers — ideally for high-speed interfaces, or with advanced considerations like DMA and cache coherency.
- Understanding of major peripheral hardware interface standards (Ethernet, PCIe, I2C, SPI, etc.) — you can critically review a schematic, and debug these interfaces if needed.
- Experience with product security, OS/system hardening, and secure boot schemes.
COMPENSATION AND BENEFITS:
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
About the company
xAI
Large Enterprise
xAI is a cutting-edge technology company focused on developing advanced artificial intelligence solutions to enhance human capabilities and optimize decision-making processes. Founded by a team of leading experts in AI and machine learning, xAI aims to address complex challenges across various industries, including healthcare, finance, and transportation. By prioritizing ethical AI development, the company is committed to creating innovative tools that empower organizations to harness the full potential of artificial intelligence while ensuring transparency and accountability.
xAI is a cutting-edge technology company focused on developing advanced artificial intelligence solutions to enhance human capabilities and optimize decision-making processes. Founded by a team of leading experts in AI and machine learning, xAI aims to address complex challenges across various industries, including healthcare, finance, and transportation. By prioritizing ethical AI development, the company is committed to creating innovative tools that empower organizations to harness the full potential of artificial intelligence while ensuring transparency and accountability.