Role at a glance
- Job function
-
Software Engineering & IT DevOps & Site Reliability Engineering Cloud & Infrastructure Engineering
- Salary
- Not Disclosed
- Location
- Austin, Texas, United States
- Experience
- experience developing and maintaining production software, services, tools, or automation.
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Software Engineer builds and operates reliable, secure, and scalable cloud platforms and services. The role develops software and automation supporting production systems, Kubernetes platforms, cloud infrastructure, observability, analytics, and AI/ML workloads.
What You'll Do
- Design, develop, test, and maintain software, services, APIs, tools, and automation.
- Build and improve cloud-native services and platform capabilities, and deploy, operate, and troubleshoot applications on Kubernetes.
- Apply reliability engineering practices and automation to improve availability, performance, scalability, and operational efficiency.
- Use observability data and operational and application analytics to troubleshoot issues, identify trends, and develop dashboards and...
- Support AI/ML infrastructure for training, LLM inference, and GPU workloads, including automation for deployment and operation.
- Participate in capacity, performance, and scale testing and disaster recovery exercises; maintain technical documentation, procedures,...
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Solid understanding of software engineering fundamentals; experience developing and maintaining production software, services, tools, or automation; proficiency in Python and/or Go; hands-on Kubernetes and container experience, including deploying and troubleshooting applications; experience with at least one major cloud platform such as AWS, Google Cloud, or Azure; familiarity with Terraform, Ansible, or similar infrastructure automation technologies; understanding of reliability concepts including monitoring, alerting, SLIs/SLOs, incident management, capacity planning, and automation; experience with or exposure to observability platforms; ability to analyze system and application data; and working knowledge of distributed systems concepts including availability, scalability, networking, fault tolerance, and performance.
Required
- Solid understanding of software engineering fundamentals and experience developing and maintaining production software, services, tools,...
- Proficiency in Python and/or Go (Golang), with the ability to write clean, maintainable, and testable code.
- Hands-on experience with Kubernetes and containers, including deploying and troubleshooting applications.
- Experience with at least one major cloud platform such as AWS, Google Cloud, or Azure.
- Familiarity with Terraform, Ansible, or similar infrastructure automation technologies.
- Understanding of reliability concepts such as monitoring, alerting, SLIs/SLOs, incident management, capacity planning, and automation.
- Experience with or exposure to technologies such as Prometheus, Grafana, Splunk, OpenTelemetry, or similar observability platforms.
- Ability to analyze system and application data to identify trends and troubleshoot issues.
Preferred
- Familiarity with Helm, Kustomize, or similar tools.
- Familiarity with SQL, Python-based data analysis, dashboards, or reporting tools.
- Familiarity with AI/ML concepts, LLMs, model inference, or GPU workloads; prior AI/ML infrastructure experience is beneficial but not...
- Strong analytical and troubleshooting skills with an interest in solving problems through software and automation.
- Strong communication skills and the ability to work effectively within cross-functional engineering teams.
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Summary
We're looking for a motivated Software Engineer to join our team and help build and operate reliable, secure, and scalable cloud platforms and services. In this role, you'll develop software and automation that support production systems, Kubernetes platforms, cloud infrastructure, observability, analytics, and emerging AI/ML workloads.
You'll combine software engineering skills with reliability engineering principles to improve how our platforms are built, deployed, monitored, and operated. You'll work alongside experienced engineers, architects, SREs, and AI/ML teams to solve technical challenges and continuously improve our systems.
The ideal candidate has a strong software engineering foundation, enjoys solving problems through code and automation, and is interested in cloud technologies, Kubernetes, reliability engineering, analytics, and AI.
Description
Software Development: Design, develop, test, and maintain software, services, APIs, tools, and automation using languages such as Python and Go.
Cloud & Platform Engineering: Build and improve cloud-native services and platform capabilities across production and non-production environments.
Kubernetes: Deploy, operate, and troubleshoot applications and services running on Kubernetes.
Develop automation that simplifies deployment and platform operations.
Reliability: Apply reliability engineering practices to improve system availability, performance, scalability, and operational efficiency.
Automation: Identify repetitive operational activities and develop software and automation to reduce manual effort and operational toil.
Observability: Use logs, metrics, traces, dashboards, and alerts to understand system behavior, troubleshoot issues, and identify opportunities for improvement.
Analytics: Analyze operational and application data to identify trends, anomalies, recurring issues, and performance bottlenecks. Develop dashboards and reporting that provide actionable insights.
AI/ML Infrastructure: Support infrastructure and platform capabilities for AI/ML training, LLM inference, and GPU-based workloads. Develop automation to simplify deployment and operation of these environments.
AI-Assisted Engineering: Explore and apply LLMs and AI technologies to improve software development, troubleshooting, analytics, automation, and operational workflows.
Scale & Resilience: Participate in capacity planning, performance testing, scale testing, and disaster recovery exercises.
Continuous Improvement: Identify opportunities to improve platform reliability, developer experience, automation, and operational processes.
Documentation: Create and maintain technical documentation, operational procedures, troubleshooting guides, and runbooks.
Collaboration: Work closely with software engineering, platform, SRE, QA, AI/ML, security, architecture, and program management teams.
Minimum Qualifications
Software Engineering: Solid understanding of software engineering fundamentals and experience developing and maintaining production software, services, tools, or automation.
Programming: Proficiency in Python and/or Go (Golang), with the ability to write clean, maintainable, and testable code.
Kubernetes: Hands-on experience with Kubernetes and containers, including deploying and troubleshooting applications. Familiarity with Helm, Kustomize, or similar tools is preferred.
Cloud: Experience with at least one major cloud platform such as AWS, Google Cloud, or Azure.
Infrastructure as Code: Familiarity with Terraform, Ansible, or similar infrastructure automation technologies.
Reliability Engineering: Understanding of reliability concepts such as monitoring, alerting, SLIs/SLOs, incident management, capacity planning, and automation.
Observability: Experience with or exposure to technologies such as Prometheus, Grafana, Splunk, OpenTelemetry, or similar observability platforms.
Analytics: Ability to analyze system and application data to identify trends and troubleshoot issues. Familiarity with SQL, Python-based data analysis, dashboards, or reporting tools is a plus.
Distributed Systems: Working knowledge of distributed system concepts including availability, scalability, networking, fault tolerance, and performance.
Preferred Qualifications
AI/ML: Familiarity with AI/ML concepts, LLMs, model inference, or GPU workloads is a plus. Prior AI/ML infrastructure experience is beneficial but not required.
Problem Solving: Strong analytical and troubleshooting skills with an interest in solving problems through software and automation.
Collaboration: Strong communication skills and the ability to work effectively within cross-functional engineering teams.
Application Deadline
Apple accepts applications to this posting on an ongoing basis.
About the company
Apple
Large Enterprise
Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.