Apple

Apple

Posted via Apple Careers

Site Reliability Engineer, AiDP Production Engineering

Always Hiring

Posted Aug 4, 2026

Role at a glance

Salary
Not Disclosed
Location
Austin Metro Area, Texas, United States
Work arrangement
On-site
Employment
Full-time
Experience
4+ years experience
Education
BS/MS in computer science or equivalent experience

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

The Service Reliability Engineer supports Apple’s AI and Data Platform Production Engineering team, which operates real-time, near real-time, and batch analytical platforms across bare-metal and cloud environments. The role focuses on designing, tuning, monitoring, and operating resilient data infrastructure and applications that support business-critical analytics and data processing at large scale.

What You'll Do

  • Assess application requirements and determine appropriate services and topologies on AWS, bare metal, and Kubernetes
  • Build automation for self-healing systems and tools to monitor and alert low-latency applications
  • Troubleshoot application-specific, network, system, and performance issues
  • Partner with engineering teams to prioritize and fix production defects
  • Triage incidents, implement mitigation steps, and conduct root cause analysis
  • Support Java applications and Spark/Flink jobs across bare metal, AWS, and Kubernetes

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

4+ years of experience in cloud-native services, messaging systems, cloud infrastructure and services, modern and distributed databases, and Python or Java programming. BS/MS in computer science or equivalent experience.

Required

  • Cloud-native services, including Apache Spark and Flink
  • Messaging systems, including Kafka
  • AWS, GCP, and Kubernetes
  • Modern and distributed databases, including Snowflake, Cassandra, SingleStore, and SAP HANA
  • Python or Java programming
  • BS/MS in computer science or equivalent experience

Preferred

  • System design, data structures, and incident management best practices
  • Observability tools, including Prometheus, Grafana, and CloudWatch
  • Performance analysis and troubleshooting of large scale distributed systems
  • Troubleshooting complex production issues
  • Root cause analysis and system reliability improvements
  • GenAI or automation tools for issue detection, alerting, or remediation
  • Data visualization tools such as Tableau, Business Objects, and ThoughtSpot

Original job description

Content provided by the employer

Summary

The Production Engineering team within the AI and Data Platform (AiDP) organization manages a wide array of real-time, near real-time, and batch analytical solutions. These platforms are integral to core business functions across Apple. These include sales, operations, finance, AppleCare, marketing, and services, and are instrumental in driving critical, data-driven decisions. To build these solutions, we leverage a combination of proprietary and leading open-source technologies such as Kafka, Spark, Iceberg, and Airflow. A key part of our mission is to enable AI-centric automations that enhance the overall efficiency and intelligence of the platform. We are looking for passionate engineers who thrive on solving complex infrastructure challenges at scale, both on-premises and in the cloud. If you are dedicated to optimizing scalable, maintainable, and user-friendly systems, you will find compelling opportunities to make a significant impact at AiDP.

Description

The Service Reliability Engineer (SRE) role within AiDP Production Engineering is a dynamic position that blends strategic architectural design with hands-on technical execution. As an SRE, you will be responsible for configuring, tuning, and ensuring the resilience of complex, multi-tiered systems to achieve optimal application performance, stability, and availability. Our team manages critical data pipelines and applications across both bare-metal and cloud computing platforms, delivering essential data processing for all of Apple’s key business functions. We operate at an immense scale, handling exabytes of data, petabytes of memory, and tens of thousands of jobs to enable predictable and performance data analytics that power features and inform decisions across the company. If you are passionate about designing, building, and running data infrastructure that has a direct and significant impact on Apple’s global business operations, this is the ideal opportunity for you.

Responsibilities

Ability to understand the application requirements (Performance, Security, Scalability etc.) and assess the right services/topology on AWS, Baremetal & Kubernetes.
Build automation to enable self-healing systems.
Build tools to monitor high performance & alert the low latency applications.
Ability to troubleshoot application specific, core network, system & performance issues.
Involvement in challenging and fast paced projects supporting Apple’s business by delivering innovative solutions.
Partner with engineering teams to prioritize and fix production defects.
Take knowledge transition from engineering teams for changes being rolled out in production.
Triage incidents based on the impact, devise and implement mitigation steps to unblock the business.
Conduct RCA, log defects and partner with engineering team for prioritization.
Support java based applications & Spark/Flink jobs on Baremetal, AWS & Kubernetes.
Share on-call rotation with other team members to support apps and services in scope.

Minimum Qualifications

4+ years experience in cloud-native services, including ETL frameworks like Apache Spark, and Flink.
4+ years experience in messaging systems (Kafka) and cloud infrastructure & services, AWS, GCP, Kubernetes.
4+ years of experience in modern & distributed databases such as Snowflake, Cassandra, SingleStore, and SAP HANA.
4+ years of programming experience in Python or Java.
BS/MS in computer science or equivalent experience.

Preferred Qualifications

Solid understanding of system design, data structures, and incident management best practices.
Should be able to understand complex architectures and be comfortable working with multiple teams.
Observability tools (e.g: Prometheus, Grafana, CloudWatch).
Ability to conduct performance analysis and troubleshoot large scale distributed systems.
Should be highly proactive with a keen focus on improving uptime/availability of our mission critical services.
Strong expertise in troubleshooting complex production issues.
Excellent problem solving, critical thinking, and communication skills.
Proven ability to resolve incidents, perform root cause analysis, and drive system reliability improvements.
Experience using GenAI or automation tools for issue detection, alerting, or remediation.
Experience in data visualization tools such as Tableau, Business Objects, ThoughtSpot.

Application Deadline

Apple accepts applications to this posting on an ongoing basis.

Apple

About the company

Apple

Large Enterprise

Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.