Apple

Apple

Posted via Apple Careers

Staff ML Infrastructure Engineer

Always Hiring

Posted Aug 4, 2026

Role at a glance

Job function
AI & Data MLOps & ML Infrastructure
Salary
Not Disclosed
Location
Cupertino, California, United States
Work arrangement
On-site
Employment
Full-time
Education
PhD

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Qualifications

Requires 10+ years of experience in machine learning infrastructure, distributed data systems, or a related field, including building and shipping large-scale production data or ML infrastructure. Requires experience architecting systems used by multiple teams or products and setting technical direction across teams. Requires strong Python and a systems language (Rust preferred; C++ or Go acceptable), hands-on performance engineering for I/O-bound workloads, familiarity with columnar and lakehouse formats, and working knowledge of ML training and inference workflows. Also requires familiarity with transformers, diffusion, retrieval-augmented generation, or fine-tuning; ability to design highly available, easy-to-use systems; mentoring, collaboration, and communication skills; and a B.S., M.S., or Ph.D. in Computer Science or Computer Engineering, or equivalent practical experience.

Required

  • 10+ years of work experience in machine learning infrastructure, distributed data systems, or a related field.
  • 10+ years of experience building and shipping large-scale data or ML infrastructure and platforms in production.
  • Experience architecting and delivering large-scale distributed data or ML infrastructure used by multiple teams or products in production.
  • Track record of setting technical direction and driving it to delivery across teams.
  • Strong Python plus a systems language (Rust strongly preferred; C++ or Go acceptable); hands-on performance engineering for I/O-bound...
  • Familiarity with columnar and lakehouse formats such as Parquet, Iceberg, Delta, or Lance.
  • Working knowledge of end-to-end ML workflows and how training and inference consume data.
  • Familiarity with transformers, diffusion, retrieval-augmented generation, or fine-tuning.

Preferred

  • Experience defining data or ML platform architecture adopted across an organization.
  • Deep experience with the data-loading and dataset-access layer of PyTorch, JAX, or TensorFlow.
  • Experience with distributed ML data-loading frameworks such as Ray Data, NVIDIA DALI, WebDataset, or Mosaic StreamingDataset.
  • Experience feeding data to GPU or TPU fleets at scale and keeping them saturated.
  • Experience with data lineage and governance systems such as DataHub, OpenLineage, or Unity Catalog, or equivalent.
  • Contributions to or operational experience with Spark, Daft, Polars, or DuckDB internals.
  • Experience with Docker and Kubernetes.

About the role

Original posting provided by Apple

View original

Summary

Join a team at the forefront of ML infrastructure and generative AI, where data and model workflows come together to enable the next generation of intelligent experiences on Apple products and services. We build robust systems that connect scalable data pipelines with advanced ML workflows, accelerating the development of real-world AI applications. Our work spans the full ML lifecycle, from experimentation to deployment, and you’ll play a key role in shaping how AI models are built, optimized, and scaled. We develop a platform for ML data and features that powers advanced GenAI applications. This includes embeddings (generation, evaluation, ANN search, multimodal support), AI Ops, efficient inference, and a modern feature platform designed to streamline experimentation and drive innovation. We’re looking for engineers and researchers passionate about generative models, data-centric ML, and intelligent systems across diverse real-world use cases. With the autonomy to experiment, the scale to make an impact, and the support to take ideas from prototype to production, you’ll work alongside a world-class team to build intelligent, flexible systems that make ML development faster, more reliable, and more creative.

Description

The Apple AI Platform team gives Apple's ML engineers and researchers the data systems and large-scale compute they need to build and ship models at Apple's bar for quality and privacy. Our team owns the data layer that large-scale model training depends on: ingestion, versioning, lineage, and governance on the way in, and high-throughput data loading into the training fleet on the way out. As a Staff ML Infrastructure Engineer, you will set the technical direction for that platform and own its hardest system-level problems, the architecture other engineers and teams build on.

Responsibilities

Own the architecture of the platform behind Apple's largest model builds: define how ingestion, immutable versioning, lineage, and governance work across structured, unstructured, and multimodal data at petabyte scale, so every model run is reproducible from a versioned dataset.
Set the technical direction for high-throughput data delivery to Apple's largest GPU and TPU fleets: define the data access and loading architecture that keeps training compute-bound, not I/O-bound.
Make the hard system-level and format calls that the whole platform inherits, columnar and lakehouse strategy, the dataset abstraction spanning structured and multimodal data, the shape of the SDK and core libraries, backed by design and proof, not just opinion.
Drive technical direction and influence across the platform and partner teams (data, embeddings, features, research), and define the interfaces and contracts between them.
Raise the technical bar across the team: mentor senior engineers, lead design reviews, and be the escalation point for the problems no one else can crack.
Partner with research and product leadership to shape the platform roadmap for next-generation workloads: foundation models, multimodal data, and retrieval-augmented systems.
Drive efficiency, reliability, and automation across the data plane and control plane that power Apple's ML fleet.

Minimum Qualifications

10+ years of work experience in machine learning infrastructure, distributed data systems, or a related field.
10+ years of experience building and shipping large-scale data or ML infrastructure and platforms in production.
Extensive experience architecting and delivering large-scale distributed data or ML infrastructure that multiple teams or products depend on in production.
A track record of setting technical direction and driving it to delivery across teams, not just within a single component.
Deep systems engineering: strong Python plus a systems language (Rust strongly preferred; C++ or Go acceptable), and hands-on performance engineering for I/O-bound workloads (Arrow, zero-copy, memory mapping, async I/O, high-throughput object storage).
Deep familiarity with columnar and lakehouse formats (Parquet, Iceberg, Delta, or Lance) and the judgment to choose between them at scale.
Strong working knowledge of the end-to-end ML workflow and how training and inference consume data, enough to architect data systems that serve them.
Familiarity with modern ML and generative techniques (transformers, diffusion, retrieval-augmented generation, fine-tuning) at the level needed to design for those consumers.
Demonstrated ability to design highly available, easy-to-use systems and to mentor and elevate the engineers around you.
Strong collaboration and communication, with the ability to align multiple teams around a technical direction.
B.S., M.S., or Ph.D. in Computer Science, Computer Engineering, or equivalent practical experience.

Preferred Qualifications

Experience defining data or ML platform architecture that was adopted across an organization.
Deep experience with the data-loading and dataset-access layer of a modern ML framework (PyTorch, JAX, or TensorFlow).
Distributed data-loading frameworks for ML: Ray Data, NVIDIA DALI, WebDataset, or Mosaic StreamingDataset.
Experience feeding data to GPU or TPU fleets at scale and keeping them saturated.
Data lineage and governance systems: DataHub, OpenLineage, Unity Catalog, or equivalent.
Contributions to or operational experience with Spark, Daft, Polars, or DuckDB internals.
Containerization and orchestration (Docker, Kubernetes).

Pay & Benefits — Cupertino, California, United States

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Application Deadline

Apple accepts applications to this posting on an ongoing basis.

Apple

About the company

Apple

Large Enterprise

Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.