Role at a glance
- Salary
- $144K – $270K/yr
- Location
- Palo Alto, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 4+ years of experience in lieu of a degree
- Education
- Bachelor’s degree in engineering, computer science, or a related STEM discipline
Spotted an issue?
We’ll check it against the original posting.
Role Summary
This hands-on technical role develops Human Data projects that provide engineering teams with high-quality training tasks and evaluations. The role partners with model and engineering teams to improve data yield, evaluation lift, and Grok’s behavior and domain performance through targeted data work.
What You'll Do
- Partner with model and engineering teams to translate needs into data projects and evaluation strategies.
- Own end-to-end delivery of data and evaluation projects supporting model development.
- Build tools and systems to measure data yield, evaluation lift, usage, and quality.
- Research, implement, and evaluate techniques for data collection, annotation, generation, and multimodal integration.
- Design and improve annotation workflows, labeling interfaces, and related tools.
- Coordinate with engineering, technical staff, Human Data teams, and stakeholders to scale projects, share learnings, and report status,...
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Bachelor’s degree in engineering, computer science, or a related STEM discipline, or 4+ years of experience in lieu of a degree; experience collaborating with cross-functional teams; demonstrated experience analyzing datasets to identify trends, anomalies, quality issues, or integrity problems.
Required
- Bachelor’s degree in engineering, computer science, or a related STEM discipline, or 4+ years of experience in lieu of a degree
- Experience collaborating with cross-functional teams
- Demonstrated experience analyzing datasets to identify trends, anomalies, quality issues, or integrity problems
Preferred
- Master’s degree or higher in a relevant technical field
- Direct experience curating, evaluating, or improving training or evaluation datasets for large language models or other AI/ML systems
- Experience designing, supporting, or optimizing annotation tools, labeling interfaces, or data workflows that prioritize factual...
- Familiarity with multimodal data or domain-specific data
- Experience conducting model or dataset evaluations focused on quality, truthfulness, or alignment with product goals
- Experience in a scripting language like Python
- Experience in SQL or other data analysis tools
Original job description
Content provided by the employer
Original job description
Content provided by the employer
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:
You will ensure engineering teams get the right data they need—high-quality tasks and evaluations—by crafting and executing high-value Human Data projects. You sit at the intersection of engineering and data operations: partnering with model teams, designing projects that capture meaningful signals, and driving data yield and eval lift. This is a hands-on technical role for engineers who deeply understand training and evaluation and want to influence data strategy.
RESPONSIBILITIES:
- Partner with model and engineering teams to understand needs and translate them into high-value data projects and evaluation strategies.
- Own end-to-end delivery of critical data and eval projects that capture meaningful training signals and support rapid model development.
- Build tools and systems to measure data effectiveness (data yield, eval lift, usage), and run evaluations yourself to continuously improve quality.
- Maintain rigorous data integrity and truthfulness, including validation processes for factual accuracy; prioritize quality over quantity.
- Research, implement, and evaluate techniques for data collection, annotation, generation, and multi-modal integration.
- Shape Grok’s behavior and domain performance through targeted data work.
- Design and improve annotation workflows, labeling interfaces, and related tools with a focus on data quality and integrity.
- Act as a liaison between engineering, technical staff, and tutoring / Human Data teams to drive alignment and knowledge sharing.
- Collaborate with Human Data Ops and stakeholders to scale projects, share learnings, and contribute to demand forecasting; report status, insights, and blockers for rapid decisions.
BASIC QUALIFICATIONS:
- Bachelor’s degree in engineering, computer science, or a related STEM discipline, or 4+ years of experience in lieu of a degree
- Experience collaborating with cross-functional teams (engineering, research, product, or annotation/operations groups)
- Demonstrated experience analyzing datasets to identify trends, anomalies, quality issues, or integrity problems
PREFERRED SKILLS AND EXPERIENCE:
- Master’s degree or higher in a relevant technical field
- Direct experience curating, evaluating, or improving training or evaluation datasets for large language models or other AI/ML systems
- Experience designing, supporting, or optimizing annotation tools, labeling interfaces, or data workflows that prioritize factual accuracy and data integrity
- Familiarity with multimodal data (text + images, code, or other modalities) or domain-specific data (science, mathematics, programming, recent events, etc.)
- Experience conducting model or dataset evaluations focused on quality, truthfulness, or alignment with product goals
- Experience in a scripting language like Python
- Experience in SQL or other data analysis tools
ADDITIONAL REQUIREMENTS:
- Weekend work may be required.
- Travel to other SpaceXAI sites may be required.
COMPENSATION AND BENEFITS:
$144,000 - $270,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.
About the company
xAI
Large Enterprise
xAI is a cutting-edge technology company focused on developing advanced artificial intelligence solutions to enhance human capabilities and optimize decision-making processes. Founded by a team of leading experts in AI and machine learning, xAI aims to address complex challenges across various industries, including healthcare, finance, and transportation. By prioritizing ethical AI development, the company is committed to creating innovative tools that empower organizations to harness the full potential of artificial intelligence while ensuring transparency and accountability.
xAI is a cutting-edge technology company focused on developing advanced artificial intelligence solutions to enhance human capabilities and optimize decision-making processes. Founded by a team of leading experts in AI and machine learning, xAI aims to address complex challenges across various industries, including healthcare, finance, and transportation. By prioritizing ethical AI development, the company is committed to creating innovative tools that empower organizations to harness the full potential of artificial intelligence while ensuring transparency and accountability.