Apple

Apple

Posted via Apple Careers

Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Always Hiring

Posted Aug 14, 2026

Role at a glance

Salary
$175K – $308.5K/yr
Location
Santa Clara, California, United States Seattle, Washington, United States
Work arrangement
On-site
Employment
Full-time
Experience
5+ years of experience leading complex, ambiguous technical projects from end to end.
Education
BS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field.

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

The Foundation Model Inference team builds and optimizes inference systems for Siri AI, Apple Intelligence, and applications powered by large foundation models. The role works across research and production within CloudOS and Private Cloud Compute to bring model architectures from prototype to large-scale deployment while improving performance, privacy, and trust.

What You'll Do

  • Partner with the Foundation Model Research team and external partners to optimize inference for language, vision, and speech model...
  • Design and ship production-grade inference systems serving millions of customers in real time.
  • Build profiling tools and simulators to identify and resolve performance bottlenecks across hardware configurations and use cases.
  • Drive technical decisions on high-throughput, low-latency serving at supercomputing scale.
  • Mentor and grow engineers across the organization.

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

5+ years leading complex technical projects; hands-on LLM inference stack experience; GPU or TPU programming concepts; PyTorch, JAX, or TensorFlow; high-throughput distributed services; cloud deployment using Kubernetes and Docker; bachelor's degree in a specified field.

Required

  • 5+ years of experience leading complex, ambiguous technical projects from end to end
  • Hands-on experience with LLM inference stacks
  • Working knowledge of GPU or TPU programming concepts
  • Proficiency with PyTorch, JAX, or TensorFlow
  • Experience building and operating high-throughput services at large distributed scale
  • Proficiency deploying applications on cloud platforms (AWS, GCP, or equivalent) using Kubernetes and Docker
  • BS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field

Preferred

  • Experience building productions systems in Go or Python
  • Strong knowledge of deep learning architectures including Transformers, encoder/decoder models, and multimodal variants
  • Experience with inference optimization frameworks such as TensorRT-LLM, vLLM, SGLang, TGI, or Nvidia Triton Server
  • Experience authoring custom CUDA kernels using CUDA C++ or OpenAI Triton
  • MS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field

Original job description

Content provided by the employer

Summary

We are the Foundation Model Inference team within Cloud OS and AI Inference organization. We are on a mission to build the most highly performant, secure and private inference stack that powers Siri AI, Apple Intelligence and Apps that are powered with the largest foundation models.

Our systems serve billions of queries daily across Siri AI, Apple Intelligence, Apple Search, Apple Music, Apple TV, App Store, iMessage, Photos, Camera, Spotlight & Safari, at remarkably low latency with every ounce of compute extracted from the hardware beneath them. We optimize language, vision, and speech models with billions of parameters using state-of-the-art techniques and ship them at Apple scale.

This is a rare opportunity to directly shape how AI reaches billions of people worldwide.

Description

You will work at the intersection of research and production, partnering closely with the Foundation Model Research team and our external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment. You will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help set the technical direction for the engineers around you.

This role sits within CloudOS and Private Cloud Compute (PCC) — Apple's purpose-built, privacy-preserving cloud infrastructure for AI workloads. PCC represents a first-of-its-kind approach to running foundation models in the cloud with verifiable privacy guarantees, and CloudOS is the systems foundation that makes it possible. You will be building and optimizing inference systems on top of this infrastructure, working closely with platform and security teams to deliver both performance and trust at scale.

Responsibilities

Partner with the Foundation Model Research team and our external partners to optimize inference for the latest model architectures across language, vision, and speech.
Design and ship production-grade inference systems serving millions of customers in real time.
Build profiling tools and simulators to identify and resolve performance bottlenecks across different hardware configurations and use cases.
Drive technical decisions on high-throughput, low-latency serving at supercomputing scale.
Mentor and grow engineers across the organization.

Minimum Qualifications

5+ years of experience leading complex, ambiguous technical projects from end to end.
Hands-on experience with LLM inference stacks.
Working knowledge of GPU or TPU programming concepts.
Proficiency with PyTorch, JAX, or TensorFlow.
Experience building and operating high-throughput services at large distributed scale.
Proficiency deploying applications on cloud platforms (AWS, GCP, or equivalent) using Kubernetes and Docker.
BS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field.

Preferred Qualifications

Experience building productions systems in Go or Python.
Strong knowledge of deep learning architectures including Transformers, encoder/decoder models, and multimodal variants.
Experience with inference optimization frameworks such as TensorRT-LLM, vLLM, SGLang, TGI, or Nvidia Triton Server.
Experience authoring custom CUDA kernels using CUDA C++ or OpenAI Triton.
MS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field.

Pay & Benefits — Seattle, Washington, United States

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $175,000 and $308,500, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Application Deadline

Apple accepts applications to this posting on an ongoing basis.

Pay & Benefits — Santa Clara, California, United States

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.
Apple

About the company

Apple

Large Enterprise

Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.