xAI

xAI

Posted via Greenhouse

Network Engineer

Posted Sep 15, 2026

Role at a glance

Job function
Software Engineering & IT Systems & Network Administration
Salary
$150K – $250K/yr
Location
Palo Alto, California, United States
Work arrangement
On-site
Employment
Full-time
Experience
Several years designing and/or operating production networks

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

The Network Engineer will design, deploy, and operate the production datacenter, campus, core, edge, and supercompute networks supporting SpaceXAI’s training and inference infrastructure. The role owns network design and operations end to end, including platform qualification, safe change execution, reliability, and performance as the infrastructure scales.

What You'll Do

  • Design, deploy, and operate production datacenter and campus/core networks at scale
  • Own routing and switching configuration standards, including BGP and at least one IGP such as OSPF or IS-IS
  • Qualify new network platforms, optics, and topologies and contribute to architecture and capacity planning
  • Build and improve monitoring, alerting, and operational documentation
  • Troubleshoot Layer 2/Layer 3 incidents end to end and drive root cause and lasting fixes
  • Automate repetitive network tasks with Python, Ansible, or similar tooling

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment; hands-on experience with BGP and at least one interior routing protocol; working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, and optics / high-speed Ethernet; experience troubleshooting live production network incidents and participating in on-call; strong written and verbal communication.

Required

  • BGP
  • At least one interior routing protocol
  • TCP/IP
  • VLANs
  • EVPN/VXLAN or equivalent datacenter overlays
  • Optics / high-speed Ethernet
  • Troubleshooting live production network incidents
  • Participating in on-call

Preferred

  • Modern datacenter vendors, including Arista, Cisco, Juniper, or Nvidia/Mellanox
  • High-performance or supercompute networking, including RoCEv2, congestion control, or GPU cluster fabrics
  • Network automation with Python, Ansible, Terraform, or similar
  • EVPN
  • Leaf-spine
  • Large-scale Ethernet fabrics
  • Rapid datacenter or cluster capacity build-outs

Original job description

Content provided by the employer

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

SpacexAI is building and operating large-scale networks that underpin training and inference infrastructure, including high-performance / supercompute fabrics that connect GPU clusters, plus the core, edge, and datacenter networks that keep that infrastructure reachable and reliable.

We need a Network Engineer who is strong on fundamentals and comfortable owning production network design, deployment, and operations end to end. This is a hands-on engineering seat — not a NOC technician role and not a network-software (telemetry/ZTP platform) SWE role. You will design and build networks, qualify platforms, ship changes safely, and keep availability and performance high as we scale.

Travel to Memphis (and other build sites) may be required for capacity build-outs. You will participate in a team on-call rotation.

RESPONSIBILITIES:

  • Design, deploy, and operate production datacenter and campus/core networks at scale
  • Own routing and switching configuration standards (BGP and at least one IGP such as OSPF or IS-IS), including change design, peer reviews, and execution
  • Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning
  •  Build and improve monitoring, alerting, and operational documentation so issues are caught and fixed quickly
  • Troubleshoot Layer 2/Layer 3 incidents end to end — from link flaps and optics through routing and traffic engineering — and drive root cause and lasting fixes
  • Automate repetitive network tasks with Python, Ansible, or similar tooling where it reduces toil
  •  Partner with compute, facilities, and software teams during cluster build-outs and maintenance windows
  •  Support high-performance / supercompute network environments (Ethernet AI/HPC fabrics, RoCE/RDMA-capable designs) as part of the broader network estate — deep specialist RoCE/NCCL ownership is a plus, not the bar for this seat

BASIC QUALIFICATIONS:

  •  Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment
  •  Solid hands-on experience with BGP and at least one interior routing protocol
  • Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, and optics / high-speed Ethernet
  • Experience troubleshooting live production network incidents and participating in on-call
  • Strong written and verbal communication; clear change docs and incident notes

PREFERRED SKILLS AND EXPERIENCE:

  • Experience with modern datacenter vendors (e.g. Arista, Cisco, Juniper, Nvidia/Mellanox)
  • Familiarity with high-performance or supercompute networking (RoCEv2, congestion control, GPU cluster fabrics) — useful context for our environment, not a hard filter
  •  Network automation (Python, Ansible, Terraform, or similar) used in production
  • Experience with EVPN, leaf-spine, and large-scale Ethernet fabrics
  •  Prior work supporting rapid datacenter or cluster capacity build-outs

ADDITIONAL REQUIREMENTS:

  • Willing to work onsite in Palo Alto

COMPENSATION AND BENEFITS:

$150,000 - $250,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

xAI

About the company

xAI

Large Enterprise

xAI is a cutting-edge technology company focused on developing advanced artificial intelligence solutions to enhance human capabilities and optimize decision-making processes. Founded by a team of leading experts in AI and machine learning, xAI aims to address complex challenges across various industries, including healthcare, finance, and transportation. By prioritizing ethical AI development, the company is committed to creating innovative tools that empower organizations to harness the full potential of artificial intelligence while ensuring transparency and accountability.