Apple

Apple

Posted via Apple Careers

ML Software Engineer

Always Hiring

Posted Aug 4, 2026

Role at a glance

Salary
Not Disclosed
Location
Seattle, Washington, United States
Work arrangement
On-site
Employment
Full-time
Experience
2 Years practical experience plus Bachelor's degree in Computer Science, Computer Engineering, or a related field — or equivalent practical experience. Experience building large-scale distributed systems that serve ML inference, reasoning across processes, hosts, and service tiers as well as model behavior under load. A ML performance-centric mindset — able to reason about latency/throughput trade-offs and to distinguish what can be solved at the system level from what requires model–system co-design. Strong in a systems or server language — Python, Go, Rust, Java, C++, or similar. Swift is what we write, but we don't expect it going in.
Education
2 Years practical experience plus Bachelor's degree in Computer Science, Computer Engineering, or a related field — or equivalent...

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

The ML Software Engineer develops and integrates the ML-inference stack for Apple Intelligence's Private Cloud Compute, which runs on Apple Silicon and distributes inference across SoC acceleration hardware and multi-node clusters. The role focuses on stability, performance, new functionality, and reliable service operation as the stack expands to additional platforms.

What You'll Do

  • Improve the stability and performance of Private Cloud Compute.
  • Implement new functionality emerging from the research community.
  • Bring the inference stack up on new generations of SoCs and hardware acceleration IP.
  • Build frameworks to distribute and coordinate inference across SoC acceleration blocks and coordinate work and data across multi-node...
  • Integrate inference code into the service stack so user traffic is served reliably and performantly, with code that is safe to develop,...

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

Minimum qualifications: 2 Years practical experience plus a Bachelor's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience; experience building large-scale distributed systems that serve ML inference across processes, hosts, and service tiers; ability to reason about model behavior under load and ML performance trade-offs; strength in a systems or server language such as Python, Go, Rust, Java, C++, or similar.

Required

  • Experience building large-scale distributed systems that serve ML inference, reasoning across processes, hosts, and service tiers as...
  • A ML performance-centric mindset — able to reason about latency/throughput trade-offs and to distinguish what can be solved at the...
  • Strong in a systems or server language — Python, Go, Rust, Java, C++, or similar. Swift is what we write, but we don't expect it going in.

Preferred

  • Low-level or close-to-the-metal work — systems programming, performance, or hardware/SoC bring-up.
  • Server-side Swift, or RPC/networking stacks (gRPC, Protocol Buffers).
  • Strong debugging and observability instincts — fluent in logs, metrics, and traces, with observability-stack (OpenTelemetry, Splunk),...
  • Depth in ML inference serving optimizations — quantization, sparsity, batching, KV cache, tokenization, GPU acceleration — enough to...
  • Solid grasp of concurrency, async/streaming, resource lifecycle, and error/cancellation handling; bonus for Apple platform experience...

Original job description

Content provided by the employer

Summary

Our team builds the ML-inference stack that powers generative AI for Apple Intelligence's Private Cloud Compute — running on Apple Silicon in the datacenter, distributing work across on-SoC acceleration hardware and multi-node clusters. Built on Private Cloud Compute's privacy guarantees, we're growing the team to scale across more platforms and support a widening set of features.

Description

As part of the team you will help engineer continuous improvements in stability and performance for Private Cloud Compute, help implement entirely new functionality as it emerges from the research community, and help bring our inference stack up on new generations of SoCs and hardware acceleration IP as we extend to more platforms — in collaboration with hardware, product and research teams throughout Apple.

We write performant and scalable frameworks (primarily in Swift, with C++ where we bridge to the hardware) to distribute and coordinate ML inference across the acceleration IP blocks of different SoCs, and to move data and coordinate work reliably across multi-node inference clusters. You will integrate inference code into a full service stack so that user traffic is served reliably and performantly, with a strong focus on code that is easy and safe to develop, update, and monitor in production.

We're a collection of highly skilled and friendly engineers who value each other's opinions and experience. We strive for excellence and believe strongly in the quality of our output. We are a team of domain experts, each specializing in specific core subject areas, with broad collective experience across cloud software services and platforms.

Minimum Qualifications

2 Years practical experience plus Bachelor's degree in Computer Science, Computer Engineering, or a related field — or equivalent practical experience.
Experience building large-scale distributed systems that serve ML inference, reasoning across processes, hosts, and service tiers as well as model behavior under load.
A ML performance-centric mindset — able to reason about latency/throughput trade-offs and to distinguish what can be solved at the system level from what requires model–system co-design.
Strong in a systems or server language — Python, Go, Rust, Java, C++, or similar. Swift is what we write, but we don't expect it going in.

Preferred Qualifications

Low-level or close-to-the-metal work — systems programming, performance, or hardware/SoC bring-up.
Server-side Swift, or RPC/networking stacks (gRPC, Protocol Buffers).
Strong debugging and observability instincts — fluent in logs, metrics, and traces, with observability-stack (OpenTelemetry, Splunk), SLO/error-budget, and on-call experience, and able to drive a production incident to root cause.
Depth in ML inference serving optimizations — quantization, sparsity, batching, KV cache, tokenization, GPU acceleration — enough to optimize the system and reason about the trade-offs and requirements it places on models (you won't be training them).
Solid grasp of concurrency, async/streaming, resource lifecycle, and error/cancellation handling; bonus for Apple platform experience (XPC, Instruments, Swift Concurrency).

Pay & Benefits — Seattle, Washington, United States

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $142,300 and $263,300, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Application Deadline

Apple accepts applications to this posting on an ongoing basis.

Apple

About the company

Apple

Large Enterprise

Apple Inc. is a global technology company known for its innovative products and services, including the iPhone, iPad, Mac computers, and Apple Watch. Founded in 1976, Apple has continuously pushed the boundaries of design and functionality, earning a reputation for high-quality consumer electronics and software solutions like iOS and macOS. With a strong commitment to user experience and privacy, Apple also leads in digital services, offering platforms such as the App Store, Apple Music, and iCloud. The company's focus on sustainability and corporate responsibility further enhances its standing as a leader in the technology sector.