Role at a glance
- Job function
-
Software Engineering & IT Solutions Architecture
- Salary
- Not Disclosed
- Location
- Arlington, Virginia, United States Austin, Texas, United States
- Work arrangement
- Hybrid
- Employment
- Full-time
Spotted an issue?
We’ll check it against the original posting.
Qualifications
Required
- 3+ years of design, implementation, or consulting in applications and infrastructures experience
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with PyTorch, JIT compilation, and AOT tracing
- 8+ years of specific technology domain areas (e.g. systems engineering, infrastructure, GPU/accelerated computing, networking, security, cloud architecture) experience
- Experience with containerized and distributed computing architectures (Kubernetes, ECS/EKS, or equivalent)
Preferred
- Experience working with end user or developer communities
- Experience with partner sales and/or alliance development in the software/technology industry
- Deep familiarity with the NVIDIA ecosystem: CUDA, TensorRT, Triton Inference Server, NIM microservices, Jetson/IGX edge platforms, and NVIDIA AI Enterprise stack
- Hands-on experience deploying open weight models (Llama 2/3/4, Mistral, Falcon, Phi, Gemma) with optimization frameworks (vLLM, TensorRT-LLM, quantization, LoRA fine-tuning)
- Experience architecting for edge, disconnected, or DDIL environments with hybrid cloud integration (AWS Outposts, ECS Anywhere)
- 5+ years of infrastructure architecture including high-performance networking (InfiniBand, RDMA, EFA), distributed storage, and cluster management
- Experience in the Defense or Intelligence Community sector with understanding of classification levels (IL4/5/6) and compliance frameworks (NIST 800-53, CMMC, FedRAMP, ITAR)
About the role
Original posting provided by amazon
You will work directly with NVIDIA and adjacent ISV partners to architect full-stack AI compute solutions on AWS, from multi-node GPU training clusters to lightweight inference at the disconnected edge. You will guide partners through model optimization, deployment, orchestration, and infrastructure design patterns that enable mission-critical AI in defense, intelligence community, and national security contexts. This includes scoping, sequencing, and architecting solutions that create measurable mission value while driving clarity across internal AWS teams and external agency stakeholders.
This position requires that the candidate selected be a US Citizen and obtain and maintain an active TS/SCI security clearance.
Key job responsibilities
You will own the technical relationship with accelerated computing and edge AI partners adopting AWS, from first whiteboard session to production at scale. Specifically, you will:
• Design and architect GPU-accelerated AI infrastructure on AWS, including multi-node training clusters with EFA networking, high-performance storage, and container orchestration for large-scale model training and inference.
• Guide partners in deploying open weight foundation models (Llama, Mistral, Falcon, Nemotron, Gemma to list a few) across AWS environments, leveraging NVIDIA NIM microservices, Triton Inference Server, TensorRT-LLM, and vLLM for optimized inference at cost and latency targets.
• Architect edge and disconnected compute solutions for DDIL (denied, degraded, intermittent, limited) environments, enabling AI inference on NVIDIA Jetson, IGX, and embedded GPU platforms integrated with AWS hybrid services (Outposts, ECS Anywhere).
• Lead sovereign AI and air-gapped deployment architecture for IL4/IL5/IL6/SCIF environments where partners must run open models entirely on-premise or in isolated AWS regions without reliance on external API endpoints.
• Drive model optimization and right-sizing engagements, including quantization, distillation, pruning, and adapter-based fine-tuning (LoRA/QLoRA) to help partners achieve production-grade performance within infrastructure constraints.
• Serve as the embedded technical advisor for partner engineering teams, conducting architecture reviews, Well-Architected assessments, and proof-of-concept builds that accelerate partner product roadmap delivery on AWS.
• Build executive and working-level relationships with partner CTOs, VP Engineering, and mission-focused government stakeholders (DoD, IC, federal civilian), translating infrastructure capabilities into mission value.
• Develop and publish reusable technical content: reference architectures, blog posts, deployment guides, and benchmark reports for GPU workloads, edge inference patterns, and open model deployment on AWS.
• Identify and drive AWS product feature requests (EC2, Bedrock Custom Model Import, Amazon SageMaker AI, AWS Neuron and AWS Nitro Enclaves) based on partner field signal to influence the service roadmap.
• Establish repeatable benchmarks to evaluate model accuracy, latency, throughput, GPU utilization, reliability, and cost across cloud and edge environments, ensuring solutions meet mission and production requirements.
• Travel approximately 30% of the time for partner site visits, customer engagements, and industry events.
A day in the life
Your morning starts with a design session alongside a partner engineering team, whiteboarding a multi-node GPU training architecture on P5 instances with EFA networking and FSx for Lustre, sizing the cluster for a Llama 3 70B fine-tuning workload that must complete within their sprint cycle.
Mid-morning, you join a call with the partner's edge deployment team to walk through an inference architecture for NVIDIA Jetson AGX Orin devices operating in a disconnected forward-deployed environment. You map out the model optimization pipeline: quantization via TensorRT-LLM, container packaging, and an OTA update mechanism that syncs model weights when connectivity is available through S3 and ECS Anywhere.
After lunch, you draft a reference architecture showing how the partner's threat detection platform can serve a Mistral 7B model via NVIDIA NIM behind an air-gapped EKS cluster in an IL5 environment, with Nitro Enclaves handling sensitive data processing. You coordinate with AWS Networking and Security specialists to validate the VPC design and encryption-at-rest patterns.
Late afternoon, you prep for a joint customer briefing with the partner's CTO and a DoW program office. You build a technical deep-dive showing total cost of ownership: comparing on-demand GPU instances vs. reserved capacity vs. edge inference at the point of need, with latency and throughput benchmarks from your recent proof-of-concept.
You close the day reviewing a product feature request you're submitting to the EC2 team based on partner feedback about GPU memory requirements for serving multiple concurrent open models, and you update your SFDC activities to capture the week's technical engagements.
About the team
You will join the WWPS ISV Partner Solutions Architecture team, a group of senior technologists who serve as the embedded technical bridge between AWS and our most strategic independent software vendor partners. We help partners build, modernize, and scale their solutions on AWS to drive mission outcomes for public sector customers worldwide. Our team operates at the intersection of deep technical architecture, partner engineering, and go-to-market execution.
About the team
Diverse Experiences
AWS values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.
Why AWS?
Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.
Inclusive Team Culture
AWS values curiosity and connection. Our employee-led and company-sponsored affinity groups promote inclusion and empower our people to take pride in what makes us unique. Our inclusion events foster stronger, more collaborative teams. Our continual innovation is fueled by the bold ideas, fresh perspectives, and passionate voices our teams bring to everything we do.
Mentorship & Career Growth
We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.
Work/Life Balance
We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve.
Basic Qualifications
- 3+ years of design, implementation, or consulting in applications and infrastructures experience- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience with PyTorch, JIT compilation, and AOT tracing
- 8+ years of specific technology domain areas (e.g. systems engineering, infrastructure, GPU/accelerated computing, networking, security, cloud architecture) experience
- Experience with containerized and distributed computing architectures (Kubernetes, ECS/EKS, or equivalent)
Preferred Qualifications
- Experience working with end user or developer communities- Experience with partner sales and/or alliance development in the software/technology industry
- Deep familiarity with the NVIDIA ecosystem: CUDA, TensorRT, Triton Inference Server, NIM microservices, Jetson/IGX edge platforms, and NVIDIA AI Enterprise stack
- Hands-on experience deploying open weight models (Llama 2/3/4, Mistral, Falcon, Phi, Gemma) with optimization frameworks (vLLM, TensorRT-LLM, quantization, LoRA fine-tuning)
- Experience architecting for edge, disconnected, or DDIL environments with hybrid cloud integration (AWS Outposts, ECS Anywhere)
- 5+ years of infrastructure architecture including high-performance networking (InfiniBand, RDMA, EFA), distributed storage, and cluster management
- Experience in the Defense or Intelligence Community sector with understanding of classification levels (IL4/5/6) and compliance frameworks (NIST 800-53, CMMC, FedRAMP, ITAR)
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, FL, Miami - 153,600.00 - 207,800.00 USD annually
USA, NY, New York - 169,000.00 - 228,600.00 USD annually
USA, TX, Austin - 153,600.00 - 207,800.00 USD annually
USA, TX, Dallas - 153,600.00 - 207,800.00 USD annually
USA, TX, Houston - 153,600.00 - 207,800.00 USD annually
USA, VA, Arlington - 153,600.00 - 207,800.00 USD annually
USA, VA, Herndon - 153,600.00 - 207,800.00 USD annually
About the company
amazon
Large Enterprise
Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.