amazon

amazon

Posted via Amazon Jobs

Sr. Software Engineer- AI/ML, Amazon Neuron Training

Posted Sep 22, 2026

Role at a glance

Job function
Software Engineering & IT Software Engineering
Salary
$193.3K – $261.5K/yr
Location
Cupertino, California, United States
Work arrangement
On-site
Employment
Internship
Education
Bachelor's, Master's

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Qualifications

Required

  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience as a mentor, tech lead or leading an engineering team
  • Bachelor's degree or above in computer science or equivalent
  • * Familiarity with LLM/transformer fundamentals such as attention mechanisms, autoregressive decoding, V-cache behavior, forms of parallelism
  • * Experience in machine learning, data mining, information retrieval, statistics or natural language processing

Preferred

  • Master's degree or above in computer science or equivalent
  • Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience in computer architecture
  • * Hands-on SFT and/or RLHF/PPO/DPO experience
  • * Experience with ML frameworks such as Pytorch/Jax, Distributed libraries and Frameworks, RL frameworks or End-to-end Model Training
  • * Experience with performance engineering: workload profiling, characterization (compute bound, memory bound, network bound), and optimization
  • * Contribution to open source projects

About the role

Original posting provided by amazon

View original
The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on AWS Trainium, Amazon's custom machine learning accelerator. Neuron includes an ML compiler, runtime, collectives library, and application framework that integrate with PyTorch and JAX, so customers can train frontier-scale models on Trainium without rewriting their stack.

The Distributed Training team enables the training of a wide range of models, from large-scale pretraining through post-training and reinforcement learning, on AWS's custom ML accelerators. As more customer workloads shift toward RLHF, PPO/GRPO, and other fine-tuning methods, we are building the distributed training infrastructure, parallelism techniques, numerics, and high-performance kernels that these methods depend on. As part of the broader Neuron organization, we work across frameworks, kernels, compiler, runtime, and collectives — a true hardware and software co-design in practice. We not only optimize current performance but also contribute to future architecture designs, since the gaps we characterize today become requirements for the next generation of Trainium.

We are looking for passionate technical leaders who can help us build and fine tune these distributed training solutions. This role offers a rare opportunity to work at the intersection of machine learning, high-performance computing, and distributed systems, where you will help shape the direction of AI acceleration technology.


Key job responsibilities
You will lead the effort to build distributed training and post-training support into PyTorch and JAX for Trainium accelerators. You will work across PyTorch and Neuron software stack with the Neuron compiler and runtime teams to enable and fine tune large-scale training, post training, and reinforcement learning workloads on the latest Trainium instances. You will own the parallelism strategies these models depend on, spanning data, tensor, pipeline, expert, and context parallelism, and apply reduced-precision formats where they measurably pay off. You will profile end to end to determine whether a workload is bound by compute, memory, collectives, or host overhead, then drive the fix to the layer that owns it, working with compiler, runtime, and collectives engineers to land it. You will translate the performance gaps you characterize into requirements that influence frameworks, and contribute upstream to the open source frameworks our customers train on.

About the team
Inclusive Team Culture
Here at Amazon, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon’s culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust.

Work/Life Balance
Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives.

Basic Qualifications

- 5+ years of non-internship professional software development experience
- 5+ years of programming with at least one software programming language experience
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Experience as a mentor, tech lead or leading an engineering team
- Bachelor's degree or above in computer science or equivalent
- * Familiarity with LLM/transformer fundamentals such as attention mechanisms, autoregressive decoding, V-cache behavior, forms of parallelism
- * Experience in machine learning, data mining, information retrieval, statistics or natural language processing

Preferred Qualifications

- Master's degree or above in computer science or equivalent
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience in computer architecture
- * Hands-on SFT and/or RLHF/PPO/DPO experience
- * Experience with ML frameworks such as Pytorch/Jax, Distributed libraries and Frameworks, RL frameworks or End-to-end Model Training
- * Experience with performance engineering: workload profiling, characterization (compute bound, memory bound, network bound), and optimization
- * Contribution to open source projects

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.



USA, CA, Cupertino - 193,300.00 - 261,500.00 USD annually
amazon

About the company

amazon

Large Enterprise

Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.