Role at a glance
- Salary
- $237.6K – $401.7K/yr
- Location
- Cupertino, California, United States
- Work arrangement
- Hybrid
- Employment
- Full-time
- Experience
- 12+ years of experience of progressive engineering leadership experience
- Education
- MS Degree in Computer Science or related degree
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Senior Infrastructure, SRE & AI Platforms Manager will set the long-term technical strategy, organizational structure, and operational roadmap for global, mission-critical infrastructure platforms on the Services Special Projects team. The role leads platforms supporting large-scale consumer and enterprise workloads, with a focus on reliability, performance, cost efficiency, operational excellence, and AI infrastructure.
What You'll Do
- Define and execute the long-term technical vision and capital investment strategy for global compute, storage, network, observability,...
- Lead a growing, multi-tiered team responsible for foundational infrastructure platforms.
- Architect, scale, and optimize environments for AI training and inference, including cluster design, scheduling, interconnect...
- Oversee Kubernetes environments, managed public cloud platforms, and large bare-metal footprints.
- Direct distributed storage, database, data streaming, networking, traffic management, and security platform strategies.
- Establish SRE practices focused on high availability, automated fault recovery, telemetry, logging, metrics, and post-incident analysis.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Minimum: MS degree in Computer Science or related degree; 12+ years of progressive engineering leadership experience; 6+ years managing multi-layered engineering organizations; cloud-native infrastructure, Kubernetes, and hybrid cloud operations across AWS, GCP, and private data centers; large-scale AI/ML training and inference infrastructure; SRE, telemetry, observability, disaster recovery, and 24/7 high-availability operations; executive and technical communication.
Required
- MS Degree in Computer Science or related degree
- 12+ years of progressive engineering leadership experience building, scaling, and operating mission-critical infrastructure platforms...
- 6+ years managing multi-layered engineering organizations (manager-of-managers)
- Cloud-native infrastructure, Kubernetes platform engineering, and hybrid cloud operations (AWS, GCP, private data centers)
- Large-scale systems for AI/ML training and inference workloads, including utilization optimization, scheduling, and high-performance...
- Site Reliability Engineering (SRE) principles, telemetry, observability frameworks, disaster recovery, and managing 24/7...
- Ability to bridge executive strategy and low-level technical trade-offs
Preferred
- Experience leading core infrastructure or foundational platform SRE for a global, tier-1 technology organization operating at massive scale
- Familiarity overseeing diverse open-source and proprietary storage/data ecosystems, including Cassandra, FoundationDB, Kafka, Redis, and...
- Managing large-scale infrastructure investments, capital expenditures, operational budgets, capacity forecasting, and cloud optimization...
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Summary
We are looking to hire a Senior Infrastructure, SRE & AI Platforms Manager to help set the long-term technical strategy, organizational structure, and operational roadmap for global, mission-critical infrastructure platforms on the Services Special Projects team.
This position requires a rare blend of deep technical domain expertise—spanning distributed systems, Kubernetes, and AI workload orchestration—and proven organizational leadership managing large, globally distributed engineering teams.
Description
In this role, you will be responsible for defining and building infrastructure strategy that balances continuous innovation with high reliability, performance, and cost efficiency. You will lead a growing, multi-tiered team of engineers who are responsible for foundational platforms that power large-scale consumer and enterprise workloads.
Beyond operational delivery, you will establish standards for operational excellence, Site Reliability Engineering (SRE), and capacity planning. You will be a key strategic partner, translating complex business imperatives into scalable platform designs while cultivating a strong engineering culture focused on automation, technical ownership, accountability, and continuous improvement.
Responsibilities
Strategic Leadership & Architecture
Multi-Year Roadmap & Strategy: Define and execute the long-term technical vision and capital investment strategy for global compute, storage, network, observability, and AI infrastructure.
Management & Organizational Alignment: Partner with Leadership to align platform capabilities, risk management, capacity investments, and architectural decisions with overarching business goals.
Technical Tradeoffs: Evaluate emerging infrastructure technologies, and make strategic platform trade-off decisions.
AI Compute & Modern Infrastructure Platforms
AI Infrastructure at Scale: Architect, scale, and optimize large-scale environments for training and inference, resolving complex challenges in cluster design, scheduling, interconnect performance, storage throughput, and capacity planning.
Hybrid & Multi-Cloud Compute: Oversee internal Kubernetes compute environments as well as managed public cloud platforms (AWS EKS, GCP GKE) and large bare-metal footprints to provide seamless developer experiences.
Data & Storage Platform Management: Direct the strategy and maintenance for distributed block/object storage alongside managed database and data streaming platforms (e.g., Cassandra, FoundationDB, Redis, PostgreSQL, MongoDB, Kafka).
Networking, Traffic & Security: Ensure reliable global traffic management, load balancing, cloud networking architectures, and enterprise security compliance across all environments.
SRE, Operational Excellence & Engineering Culture
Site Reliability Engineering (SRE): Cultivate a mature SRE culture focusing on high availability, automated fault recovery, telemetry, logging, metrics, and rigorous post-incident analysis.
Global Team & Leadership Development: Build, mentor, and lead a globally distributed organization comprising engineers, managers, and managers-of-managers across all levels (interns through senior principal staff).
Culture of Ownership & Automation: Establish an environment characterized by strong technical ownership, clear accountability, continuous operational refinement, and aggressive automation of manual processes.
Minimum Qualifications
MS Degree in Computer Science or related degree and 12+ years of experience of progressive engineering leadership experience building, scaling, and operating mission-critical infrastructure platforms and global services.
Management & Leadership Scope: 6+ years managing multi-layered engineering organizations (manager-of-managers) with a proven track record of hiring, developing, and retaining top-tier technical talent across global sites.
Cloud & Distributed Compute Expertise: Demonstrated hands-on and architectural mastery of cloud-native infrastructure, Kubernetes platform engineering, and hybrid cloud operations (AWS, GCP, private data centers).
Accelerated Computing & AI Infrastructure: Direct operational and architectural experience running large-scale systems for AI/ML training and inference workloads, including utilization optimization, scheduling, and high-performance storage/networking.
SRE & Production Operations: Deep background in Site Reliability Engineering (SRE) principles, telemetry, observability frameworks, disaster recovery, and managing 24/7 high-availability infrastructure at scale.
Technical Communication: Exceptional ability to seamlessly bridge executive strategy and low-level technical trade-offs—communicating vision to executive stakeholders while driving detailed technical discussions with principal engineers.
Preferred Qualifications
Large-Scale Enterprise Provenance: Experience leading core infrastructure or foundational platform SRE for a global, tier-1 technology organization operating at massive scale.
Multi-Engine Database & Data Infrastructure: Familiarity overseeing diverse open-source and proprietary storage/data ecosystems (e.g., Cassandra, FoundationDB, Kafka, Redis, PostgreSQL).
Financial & Capacity Governance: Proven competency managing large-scale infrastructure investments, capital expenditures, operational budgets, capacity forecasting, and cloud optimization strategies.
Pay & Benefits — Cupertino, California, United States
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $237,600 and $401,700, and your base pay will depend on your skills, qualifications, experience, and location.Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits
Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.
Application Deadline
Apple accepts applications to this posting on an ongoing basis.
About the company
Apple
Large Enterprise
Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.
Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.