xAI

xAI

Posted via Greenhouse

Member of Technical Staff - Post-Training and RL

Posted Aug 5, 2026

Role at a glance

Salary
$180K – $600K/yr
Location
Palo Alto, California, United States
Work arrangement
On-site
Employment
Full-time
Experience
relevant experience is not required

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

This role focuses on critical post-training and reinforcement learning challenges for AI systems, including methods intended to improve reasoning, truthfulness, and real-world capabilities. The work includes reward modeling, preference optimization, and reinforcement learning.

What You'll Do

  • Work on reward modeling challenges
  • Develop preference optimization methods, including RLHF and DPO
  • Apply reinforcement learning to improve reasoning, truthfulness, and real-world capabilities

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

Post-training and reinforcement learning techniques; reinforcement learning and alignment methods; power user of AI models.

Required

  • Post-training and reinforcement learning techniques
  • Reinforcement learning and alignment methods
  • AI models

Preferred

  • Previously worked on post-training or RLHF
  • Trained models used by millions of people

Original job description

Content provided by the employer

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

  • You will work on the most critical post-training and reinforcement learning challenges at any given time — including reward modeling, preference optimization (RLHF/DPO), and RL for improving reasoning, truthfulness, and real-world capabilities.
  • You will get clarity on your first project before an offer.

BASIC QUALIFICATIONS:

  • You believe truth-seeking AI is the most important and challenging problem.
  • You are obsessed about building incredibly useful models through post-training and RL techniques.
  • You are a power user of AI models and eager to push the boundaries of what’s possible with reinforcement learning and alignment methods.
  • If you previously worked on post-training, RLHF, or trained models used by millions of people it’s a big plus, but relevant experience is not required.
  • You take pride in your work and thrive in meritocratic environments.

COMPENSATION AND BENEFITS:

$180,000 - $600,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

xAI

About the company

xAI

Large Enterprise

xAI is a cutting-edge technology company focused on developing advanced artificial intelligence solutions to enhance human capabilities and optimize decision-making processes. Founded by a team of leading experts in AI and machine learning, xAI aims to address complex challenges across various industries, including healthcare, finance, and transportation. By prioritizing ethical AI development, the company is committed to creating innovative tools that empower organizations to harness the full potential of artificial intelligence while ensuring transparency and accountability.