Role at a glance
- Salary
- $175K – $308.5K/yr
- Location
- Cupertino, California, United States San Diego, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- Three years of relevant industry experience in test automation, software development, or related areas.
- Education
- BS in Computer Science, Mathematics, or a related field (or equivalent practical experience)
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Apple Intelligence Platform Experience Validation team builds tooling and automation to evaluate the quality of Apple Intelligence generative features before release. This senior individual-contributor role leads automated model evaluation and partners with modeling, framework, and infrastructure teams to make model quality a continuously measured signal.
What You'll Do
- Design, build, and maintain LLM-as-a-judge evaluation harnesses and integrate them into CI and automation pipelines.
- Author and curate evaluation sets and rubrics, and run evaluation jobs at scale.
- Triage evaluation results and distinguish model regressions from rubric problems or infrastructure noise.
- Build test coverage and tooling from component tests through end-to-end tests for generative AI models and frameworks.
- Package evaluation tooling as reusable libraries and jobs for team and partner-team adoption.
- Define quality gates and reporting to communicate regressions before human evaluation or population rollout.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
BS in Computer Science, Mathematics, or a related field (or equivalent practical experience); three years of relevant industry experience in test automation, software development, or related areas.
Required
- Python, including data-pipeline fluency (JSON/YAML, REST APIs)
- LLM-as-a-judge evaluation and rubric design
- Software engineering fundamentals for maintainable pipelines and tooling
- Debugging and triage skills
- Software development lifecycle, testing methodologies, and QA processes
- Written and verbal communication
- Cross-functional leadership with modeling, framework, and infrastructure teams
- On-device tooling and device/model eval infrastructure
Preferred
- Experience with Xcode
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Summary
The Apple Intelligence Platform Experience Validation team builds the tooling and automation that keeps Apple Intelligence features high-quality before they ship. We are looking for a Senior SDET to lead the design and implementation of automated model evaluation: standing up LLM-as-a-judge in existing and new pipelines, and building the infrastructure that catches model regressions before they reach human evaluation or the live on population.
This is a hands-on, senior individual-contributor role. You will own eval automation as a discipline across the team, partnering with modeling, framework, and infrastructure teams to make model quality a first-class, continuously measured signal.
Description
You will build and maintain model level, component or end-to-end evaluation coverage for the generative features our team validates. Your job is to leverage LLM judge scoring output quality in automation, ensuring reliable, repeatable eval jobs that run that produce actionable signal.
The kinds of problems you will work on include:
* Image / visual generation: validating model output and its associated classification metadata, and detecting quality or behavior regressions across model updates.
* Natural-language generation: evaluating whether generated artifacts and responses match user intent, moving at-desk LLM judges into a scalable and repeatable automation environment.
* Correctness beyond string matching: replacing exact-match checks for open-ended or factual responses with an LLM-as-judge stage integrated into the pipeline.
* Generated insights and summaries: assessing whether model-generated content is sensible and good enough to surface to users.
You will decide when a component-level check (an API or CLI that exercises the model against its framework) is sufficient and when a full end-to-end user flow is required, and you will build the tooling for both.
Responsibilities
Design, build, and maintain LLM-as-a-judge evaluation harnesses and integrate them into existing and new CI/automation pipelines.
Author and curate eval sets and rubrics; partner with modeling teams whose own eval sets can run hundreds of examples judged by a separate model.
Run eval jobs at scale, triage results, and distinguish real model regressions from rubric problems or infrastructure noise so the signal stays actionable.
Plan and build test coverage and tooling from low level component tests through to end-to-end tests that exercise models and frameworks powering Generative AI features
Package eval tooling for reuse — reusable libraries and jobs that other engineers on the team and partner teams can adopt.
Define quality gates and reporting so model regressions are caught and communicated before human eval or population rollout.
Collaborate with data scientists and modeling engineers on approaches to spot regressions across large output sets or between model updates.
Minimum Qualifications
BS in Computer Science, Mathematics, or a related field (or equivalent practical experience)
Three years of relevant industry experience in test automation, software development, or related areas.
Preferred Qualifications
Strong practical knowledge of Python, including data-pipeline fluency (JSON/YAML, REST APIs).
Hands-on experience with LLM-as-a-judge evaluation and rubric design, or a strong demonstrated ability to ramp into it quickly.
Strong software engineering fundamentals — able to define atomic, composable components and build maintainable pipelines and tooling, not just scripts.
Strong debugging and triage skills; able to separate genuine regressions from infrastructure or rubric noise.
Strong knowledge of the software development lifecycle, testing methodologies, and QA processes.
Excellent written and verbal communication; able to document clearly and describe quality signal to modeling and leadership audiences.
Ability to lead work across varying priorities and partner multi-functionally with modeling, framework, and infrastructure teams.
Experience building on-device tooling and device/model eval infrastructure.
Experience integrating with CI/CD and job orchestration systems, and comfort deploying tooling as reusable libraries.
Familiarity with generative model behavior — image generation, NLP, or LLM output evaluation.
Experience curating and reasoning about large datasets; comfort manually inspecting data (Jupyter or similar) to build intuition and drive next steps.
Awareness of dataset bias and fairness considerations in evaluation.
Experience with database/query tooling (e.g., SQL) and dashboards/visualization for reporting quality trends.
Experience with Xcode is a bonus.
Pay & Benefits — San Diego, California, United States
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $175,000 and $308,500, and your base pay will depend on your skills, qualifications, experience, and location.Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits
Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.
Application Deadline
Apple accepts applications to this posting on an ongoing basis.
Pay & Benefits — Cupertino, California, United States
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location.Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits
Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.
About the company
Apple
Large Enterprise
Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.
Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.