amazon

amazon

Posted via Amazon Jobs

Sr. Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs

Posted Aug 3, 2026

Role at a glance

Job function
Software Engineering & IT Systems Software Engineering
Salary
Not Disclosed
Location
Cupertino, California, United States Seattle, Washington, United States
Work arrangement
On-site
Employment
Internship
Education
Preferred Qualifications - Bachelor's degree in computer science or equivalent

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

This engineer builds and optimizes the low-level data-transfer path for disaggregated inference, moving KV cache between prefill and decode pools across accelerators, servers, and heterogeneous memory. The work focuses on performance over AWS network infrastructure and supports large-scale AI model serving.

What You'll Do

  • Build and optimize low-level data-movement software across accelerators, servers, and heterogeneous memory over AWS's high-performance...
  • Profile performance, find bottlenecks, and improve throughput toward the hardware's theoretical limits.
  • Work across kernel, network-transport, and inference-framework layers with teams building chips, runtimes, and models.
  • Deliver features for large production clusters and AI models.

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

Required

  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience as a mentor, tech lead or leading an engineering team
  • Must have C/C++ Coding Experience

Preferred

  • Bachelor's degree in computer science or equivalent
  • ML Communications (NCCL, NIXL, NVSHMEM)

Original job description

Content provided by the employer

Every token a large language model generates depends on data reaching the right accelerator at the right moment. As AI models outgrow any single chip, the network between accelerators becomes the bottleneck that decides how fast — and how affordably — the world's largest models can serve real users. That network layer is what our team builds.

We're looking for an engineer to work at the frontier of disaggregated inference: splitting LLM serving into separate prefill and decode pools and moving the model's KV cache between them at the absolute limit of what the hardware allows. Get it right and users get answers in milliseconds; get it wrong and the fastest accelerators in the world sit idle waiting on data. You'll own pieces of the high-speed transfer path that make the difference, and you'll measure your success in how close you run to the theoretical peak of the machine.

In this role you will:

Build and optimize the low-level data-movement software across accelerators, servers, and heterogeneous memory — over AWS's highest-performance network fabric.
Push performance to the hardware roofline: profile, find the real bottleneck, and close the gap between "it works" and "it runs as fast as physics permits."
Work across the stack — from kernel and network transport up to the inference frameworks — and partner with teams building the chips, runtime, and models.
Deliver features that run on our largest clusters, for our largest customers, serving the largest AI models in production.
What we're looking for:

Strong C/C++ and a love for low-level, performance-critical systems — solid command of Linux, kernels, memory, and writing fast code.
The instinct to ask "how fast could this possibly go?" and the rigor to measure it.
Experience with high-speed networking or HPC interconnects (RDMA, InfiniBand, libfabric, UCX, NIXL, MPI) is valued highly; embedded-systems experience is a plus.
Prior AI/ML experience is welcome but not required — if you're a great systems engineer, we'll teach you the ML side.
If you like solving genuinely hard problems, working shoulder-to-shoulder with HPC and ML customers, iterating fast, and shipping at a scale few places can offer, come join us. This is a role on the leading edge of AI/ML infrastructure.

About the team: You'd be joining Annapurna Labs, an integral part of AWS. Annapurna designs the hardware and software building blocks behind EC2 — every EC2 instance runs on hardware we designed. We specialize in the chips, systems, and software that optimize the AWS customer experience, and this team sits where cutting-edge AI meets the silicon and the network underneath it.

A day in the life
Annapurna Labs, a crucial part of AWS, is responsible for developing hardware and software components for EC2 infrastructure. Our team focuses on building networking solutions that for Machine Learning (ML) and High-Performance Computing (HPC) workloads on AWS.

We have mixed discipline orgs, you’d be working side by side with infrastructure experts, hardware engineers, RTL engineers, scientists & architects. Our workforce spans the globe and is truly international, you’ll find yourself working side by side with individuals from numerous countries. We take mentorship seriously, you can both expect senior mentorship and will be expected to mentor new and junior engineers.

The pace is fast as we work on the latest advancements of AI/ML, but we take the time to bond as a team and enjoy the successes. We offer flexibility in working hours, and respect WLB as a core org tenet. The team enjoys working with numerous principal-level engineers and closely with directors, career growth opportunities are certainly available. This is a role where you will always be encouraged to keep learning, the AI/ML field is fast moving and constantly evolving.

Basic Qualifications

- 5+ years of non-internship professional software development experience
- 5+ years of programming with at least one software programming language experience
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Experience as a mentor, tech lead or leading an engineering team
- Must have C/C++ Coding Experience

Preferred Qualifications

- Bachelor's degree in computer science or equivalent
- ML Communications (NCCL, NIXL, NVSHMEM)

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.



USA, CA, Cupertino - 193,300.00 - 261,500.00 USD annually
USA, WA, Seattle - 168,100.00 - 227,400.00 USD annually
amazon

About the company

amazon

Large Enterprise

Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.