amazon

amazon

Posted via Amazon Jobs

Principal Technical Account Manager, ES - NAMER - US-Frontier AI

Posted Sep 19, 2026

Role at a glance

Job function
AI & Data AI Solutions Architecture
Salary
$210.2K – $284.3K/yr
Location
San Francisco, California, United States
Work arrangement
Hybrid
Employment
Full-time
Education
Bachelor's

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Qualifications

Required

  • Bachelor's degree
  • 7+ years of design/implementation/operations/consulting with distributed applications experience
  • 7+ years of technical engineering experience
  • Bachelor’s Degree in Computer Science, Math, or related discipline or 7+ years of related work experience.

Preferred

  • Experience in a 24x7 operational services or support environment
  • Experience in internal enterprise or external customer-facing environment as a technical lead
  • 4+ years of experience in AI/ML, distributed computing, or GPU-accelerated infrastructure (e.g., model training, inference systems, HPC for ML)
  • Familiarity with SageMaker HyperPod for managed distributed training clusters including automated health checks, node replacement, and checkpoint-based recovery
  • Experience with HPC job schedulers (Slurm, PBS, LSF, EKS) for orchestrating multi-node ML training workloads
  • Experience with high-performance parallel file systems (Amazon FSx for Lustre, GPFS/Spectrum Scale) for ML data pipelines
  • Familiarity with AWS Parallel Computing Service (PCS), AWS ParallelCluster, AWS Batch, or equivalent managed HPC/ML cluster services
  • Professional oral and written communication skills, presenting to an audience containing one or more executive team member(s)

About the role

Original posting provided by amazon

View original
Are you ready to transform how businesses leverage artificial intelligence and machine learning at scale? Join the Amazon Web Services (AWS) Support team and become a strategic partner in delivering Amazon AI/ML solutions that empower our Frontier AI customers to innovate, optimize, and achieve unprecedented operational excellence.

Amazon Web Services (AWS) is seeking an experienced Sr. TAM with expertise in AI/ML, HPC, and/or other technologies to join our Frontier AI Technical Account Management (TAM) team.

You'll be at the forefront of solving complex AI/ML model training and inference implementation challenges, guiding Frontier research customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem-solving, you'll help organizations unlock the full potential of artificial intelligence and machine learning technologies — from distributed model training on GPU clusters to production-grade inference at scale.

The TAM role is not directly hands on keyboard within the customer’s environment for troubleshooting customer support issues, rather you will work with appropriate engineers and service teams to see issues through to resolution. You will help our customers design, build, operate, and secure their cloud environments. More importantly, you will work proactively to help craft and execute strategies to drive our customers' adoption and use of AWS services, including EC2, S3, DDB, RDS, and many more.

Your technical acumen and customer-facing skills will enable you to effectively represent AWS within a customer environment, and drive discussions with senior leadership regarding incidents, trade-offs, support and risk management. You will provide advocacy and strategic technical guidance to help plan and build solutions using best practices and proactively keep your customers’ AWS environments operationally healthy and resilient.

The close relationships developed with your customers will allow you to understand their business/operational needs and technical challenges to help them achieve the greatest value from AWS. This position will require the ability to travel 10% or more as needed.

The TAM is the centerpiece of value to our Enterprise Support customers. If you wish to be at the forefront of innovation, come join us!

Key job responsibilities
Deliver Strategic Technical Engagements — Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads, including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster, the latest GPU-accelerated computing (i.e. P6/P6e , G7/G7e instances), AWS Trainium-based training (Trn3 UltraServers), and multi-node NCCL communication tuning over EFA’s SRD protocol.

Architect and Validate Innovative Solutions — Supporting customers who design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, DeepSpeed, Megatron-LM), SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement, high-performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance at scale. Architect solutions using P6e UltraServers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference.

Enable Customer Success — Support customers in implementing business-critical HPC capabilities, including the development of large language model (LLM) (Llama, GPT-class models), physics-informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls for GPU/Trainium-intensive workloads.

Enable Business Critical Outcomes — Partner with service teams to enhance model training throughput, optimize NCCL collective communications, improve GPU/Trainium utilization across multi-node UltraClusters, and drive operational efficiency through proactive monitoring, automated failure recovery (HyperPod health checks), and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR, share reference architecture, performance, and benchmarks with broader Technical communities

Serve as Trusted Advisor and Advocate — Develop and nurture technical partnerships with enterprise stakeholders, serving as the trusted advisor for AI/ML infrastructure decisions spanning compute, networking (Elastic Fabric Adapter with SRD), storage, orchestration, and the HPC-to-AI convergence journey.

A day in the life
Your day will be dynamic and impactful, involving deep technical consultations on distributed training architectures, strategic solution design for GPU and Training cluster deployments, and collaborative problem-solving across multi-node ML environments. You'll engage with technical leaders, architect innovative AI/ML implementations — from Slurm-managed training clusters and SageMaker HyperPod to PyTorch FSDP/DeepSpeed training jobs and Neuron SDK compilation workflows — and provide expert guidance that bridges machine learning infrastructure with business objectives.

You will partner with Solution Architects, Frontline Engineers, and Service teams to provide customers with AWS AI/ML best practice guidance, diving deep into machine learning infrastructure services (PCS, ParallelCluster, HyperPod, Batch), promoting customers' AI/ML workloads to production, developing regional AI/ML strategies, advising on HPC-to-AI convergence patterns (simulation-surrogate loops, physics-informed neural networks), and training field teams on distributed training patterns, GPU/Trainium cluster operations, and the use cases and benefits of artificial intelligence and machine learning at scale. You establish trust with your customers to understand their business/operational needs and technical challenges and help them achieve the greatest value from AWS. This position will require the ability to travel 10% or more as needed.
You will partner with Solution Architects, Frontline Engineers, and Service teams to provide customers with AWS AI/ML best practice guidance, diving deep into machine learning infrastructure services (PCS, ParallelCluster, HyperPod, Batch), promoting customers' AI/ML workloads to production, developing regional AI/ML strategies, advising on HPC-to-AI convergence patterns (simulation-surrogate loops, physics-informed neural networks), and training field teams on distributed training patterns, GPU/Trainium cluster operations, and the use cases and benefits of artificial intelligence and machine learning at scale. You establish trust with your customers to understand their business/operational needs and technical challenges, and help them achieve the greatest value from AWS. This position will require the ability to travel 10% or more as needed.

About the team
The Frontier AI Enterprise Support team supports our complex Startup customers who are building large AI training, inference models, and requires high GPU demand. We help these Frontier research labs scale fast. We are a collaborative group of technical innovators dedicated to pushing the boundaries of cloud computing and artificial intelligence. Our team thrives on solving complex challenges — from optimizing NCCL all-reduce operations across hundreds of GPUs to architecting elastic training clusters that scale with customer demand. We believe in continuous learning, mutual support, and driving technological advancement.

Why AWS?
 Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Inclusive Team Culture Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity and AmazeCon conferences, inspire us to never stop embracing our uniqueness.

Mentorship and Career Growth We’re continuously raising our performance bar as we strive to become

Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

Basic Qualifications

- Bachelor's degree
- 7+ years of design/implementation/operations/consulting with distributed applications experience
- 7+ years of technical engineering experience
- Bachelor’s Degree in Computer Science, Math, or related discipline or 7+ years of related work experience.

Preferred Qualifications

- Experience in a 24x7 operational services or support environment
- Experience in internal enterprise or external customer-facing environment as a technical lead
- 4+ years of experience in AI/ML, distributed computing, or GPU-accelerated infrastructure (e.g., model training, inference systems, HPC for ML)
- Familiarity with SageMaker HyperPod for managed distributed training clusters including automated health checks, node replacement, and checkpoint-based recovery
- Experience with HPC job schedulers (Slurm, PBS, LSF, EKS) for orchestrating multi-node ML training workloads
- Experience with high-performance parallel file systems (Amazon FSx for Lustre, GPFS/Spectrum Scale) for ML data pipelines
- Familiarity with AWS Parallel Computing Service (PCS), AWS ParallelCluster, AWS Batch, or equivalent managed HPC/ML cluster services
- Professional oral and written communication skills, presenting to an audience containing one or more executive team member(s)

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.



USA, CA, San Francisco - 210,200.00 - 284,300.00 USD annually
amazon

About the company

amazon

Large Enterprise

Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.