Role at a glance
- Salary
- $237.6K – $356.4K/yr
- Location
- Cupertino, California, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems
Spotted an issue?
We’ll check it against the original posting.
Role Summary
Apple Service Engineering's Compute team operates and scales large-scale batch compute infrastructure across Apple's data centers. The role leads Infrastructure and Site Reliability Engineers responsible for core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies, with a focus on availability, reliability, performance, and modernization.
What You'll Do
- Champion AI-powered tooling and automation for incident triage, operational toil reduction, capacity efficiency, and engineering workflows
- Lead, mentor, and grow a team of Software and SRE engineers across multiple geographies
- Establish and maintain a 24/7 on-call rotation with escalation paths, severity definitions, and response time SLAs
- Oversee release engineering and deployment automation, including CI/CD pipelines, canary deployments, and zero-downtime rollouts
- Manage infrastructure modernization initiatives including Kubernetes control plane operations, database migrations, and configuration...
- Drive incident management excellence through post-incident reviews, preventive measures, and production readiness reviews
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Minimum: 5+ years managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems; experience building on-call organizations; cloud infrastructure, compute orchestration, bare metal provisioning, Kubernetes, OpenStack, KVM/hypervisor technologies, Infrastructure as Code, SRE principles, communication, and engineering talent development.
Required
- 5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems
- Building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes
- Cloud infrastructure, compute orchestration, and bare metal provisioning at scale
- Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools including Chef, Ansible, Terraform, or Salt
- SRE principles including SLOs, error budgets, capacity planning, and release engineering
- Excellent verbal and written communication skills
- Recruit, develop, and retain high-performing engineering talent
Preferred
- Leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation
- Managing or scaling batch compute, job scheduling, or HPC platforms
- Proficiency in Go or Python
- Observability stacks including Prometheus, Grafana, and distributed tracing, and centralized logging at scale
- Operating large-scale multi-tenant Infrastructure as a Managed Service
- Managing geographically distributed teams and follow-the-sun on-call models
- Capacity efficiency initiatives resulting in measurable cost optimization
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Summary
People at Apple don't just build products — they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it.
The Apple Service Engineering (ASE) team builds and provides systems and infrastructure that power Apple's services (such as iCloud, Apple Music, Apple Intelligence, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. Our services have to scale globally, stay highly available, and "just work." If you love designing, engineering, and running systems and infrastructure that will help millions of customers, then this is the place for you!
Description
Apple Service Engineering (ASE)'s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apple's data centers. You will manage a team that operates core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies — ensuring platform availability, reliability, and performance at Apple scale.
You will drive strategic initiatives spanning multi-datacenter capacity planning, incident management, release engineering, observability, and infrastructure modernization. This role requires a leader who can balance operational excellence with engineering innovation, establishing SLOs, driving production readiness, and building the automation and tooling that enable a growing platform to scale efficiently. You will champion the use of AI to accelerate incident triage, improve operational workflows, drive capacity efficiency, and enhance team productivity across all domains.
Responsibilities
Champion AI-powered tooling and automation to improve incident triage, reduce operational toil, drive capacity efficiency, and accelerate engineering workflows
Lead, mentor, and grow a team of Software and SRE engineers across multiple geographies, fostering a culture of ownership, collaboration, and continuous improvement
Establish and maintain a sustainable 24/7 on-call rotation with clear escalation paths, severity definitions, and response time SLAs across US and UK locations
Oversee release engineering and deployment automation, including CI/CD pipelines, canary deployments, and zero-downtime rollout strategies
Manage infrastructure modernization initiatives including Kubernetes control plane operations, database migrations, and configuration management evolution
Drive incident management excellence — including post-incident reviews, preventive measures, and production readiness reviews for all releases
Minimum Qualifications
5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems
Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes
Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale
Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt)
Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering
Excellent verbal and written communication skills with the ability to influence across teams and levels
Demonstrated ability to recruit, develop, and retain high-performing engineering talent
Preferred Qualifications
Hands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation
Experience managing or scaling batch compute, job scheduling, or HPC platforms
Proficiency in Go or Python with a strong automation-first mindset
Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale
Experience operating large-scale multi-tenant Infrastructure as a Managed Service
Experience managing geographically distributed teams and follow-the-sun on-call models
Track record of driving capacity efficiency initiatives resulting in measurable cost optimization
Pay & Benefits — Cupertino, California, United States
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $237,600 and $356,400, and your base pay will depend on your skills, qualifications, experience, and location.Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits
Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.
Application Deadline
Apple accepts applications to this posting on an ongoing basis.
About the company
Apple
Large Enterprise
Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.
Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.