Role at a glance
- Salary
- $143.7K – $194.4K/yr
- Location
- Seattle, Washington, United States
- Work arrangement
- On-site
- Employment
- Internship
- Education
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The EKS Node Runtime team builds and operates the foundational compute layer for Amazon Elastic Kubernetes Service, including optimized AMIs, container runtimes, isolation technologies, accelerator support, and node monitoring. This role focuses on reliable and secure infrastructure for containerized and AI/ML workloads across large-scale customer fleets.
What You'll Do
- Design and implement node-level runtime infrastructure, optimized AMIs, GPU and accelerator support, VM isolation runtimes, and node...
- Build AMI release pipelines, integration testing, automated CVE patching, instance qualification testing, and production validation.
- Develop and operate Kubernetes node components including kubelet configurations, device plugins, DRA drivers, containerd, and runc.
- Qualify new GPU and accelerator instance types, manage NVIDIA driver updates, and integrate GPU, EFA, and Neuron capabilities.
- Participate in on-call rotations, respond to SEVs and customer escalations, and improve monitoring, alarming, incident response, and...
- Lead technical design decisions, collaborate with AWS service teams and technology partners, and mentor junior engineers.
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Required
- 3+ years of non-internship professional software development experience
- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field
- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
Preferred
- Experience in Kubernetes, Docker or containers ecosystem, or experience that includes strong analytical skills, attention to detail, and effective communication abilities and experience in managing and troublshooting network
- Knowledge of and experience with cloud infrastructure technologies
- Experience in Linux/RHEL, or experience in Kubernetes, Docker or containers ecosystem and experience with programming/scripting (Batch, VB, PowerShell, Java, C#, Chef, Perl, Ruby and/or PHP)
- Experience with CloudFormation, Chef, Puppet, Salt, or Ansible in production environments
- Experience delivering products against plan in a fast-paced, multi-disciplined, distributed-responsibility and often ambiguous environment
- Experience with enterprise architecture including virtualization technologies and distributed architecture
- Contributions to open-source projects, especially in the Kubernetes/CNCF ecosystem
- Strong operational mindset - experience with monitoring, observability, and incident management at scale
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Key job responsibilities
Design & Build: Architect and implement the node-level runtime infrastructure that powers EKS compute. Design optimized Amazon Machine Images (AMIs), integrate GPU drivers and accelerator support, implement VM isolation runtimes, and build node monitoring systems that operate reliably across millions of customer nodes.
Operate at Scale: Own the operational health of the EKS node runtime layer handling millions of customer workloads across diverse instance types and accelerators. Participate in on-call rotations, respond to SEVs, resolve customer escalations, and drive operational improvements that reduce MTTR and improve fleet stability.
Technical Leadership: Lead the design and implementation of complex node runtime features end-to-end, from requirements through AMI build pipelines, testing frameworks, deployment, and production validation. Drive technical decisions on AMI architecture, container runtime strategies, and GPU/accelerator integration.
Kubernetes & Runtime Expertise: Work deeply with node-level Kubernetes components including kubelet configuration, device plugins, Dynamic Resource Allocation (DRA) drivers, and container runtimes (containerd, runc). Understand Linux internals, systemd, GPU drivers, and isolation technologies to build robust node capabilities.
Cross-Team Collaboration: Partner with EKS control plane teams, EC2 instance teams, NVIDIA, AWS AI/ML services (SageMaker, Trainium, Inferentia), and open-source communities to deliver integrated solutions that enable customer workloads from traditional containers to cutting edge AI training and inference.
Mentorship: Mentor junior engineers on systems programming, Linux internals, and operational best practices. Conduct thorough code reviews, share knowledge on GPU technologies and container runtimes, and contribute to a culture of engineering excellence.
Operational Excellence: Drive improvements in AMI build and release pipelines, automated CVE patching, instance qualification testing, monitoring and alarming for node health, and incident response processes. Build automation that reduces manual toil and accelerates time-to-resolution.
Innovation: Identify opportunities to enable new GPU and accelerator technologies, optimize node startup performance, improve container isolation and security, and push the boundaries of what's possible with AI/ML workloads on Kubernetes. Simplify customer experiences through better defaults, comprehensive monitoring, and self-healing capabilities.
About the team
The EKS Node Runtime team builds and operates the foundational compute layer that powers Amazon Elastic Kubernetes Service (EKS). We own the node-level infrastructure that makes EC2 instances become reliable, secure, and high-performance Kubernetes nodes. This includes designing and building optimized Amazon Machine Images (AMIs), integrating GPU drivers and accelerator support, implementing container runtimes and isolation technologies, building node monitoring systems, and automating security patching across the fleet. Our systems run on millions of customer nodes and enable workloads ranging from traditional microservices to the largest AI training and inference clusters in the world.
Our team tackles challenges at the intersection of Linux systems programming, Kubernetes internals, GPU/accelerator technologies, and distributed systems operations at massive scale. We're building the next generation of node capabilities—from VM-based pod isolation with Nitro partitions to Dynamic Resource Allocation for GPUs and accelerators, from qualifying cutting-edge instance types like P6 (B200/B300) and Trainium2 to enabling AI agent workload patterns with pod pause, resume, and snapshot capabilities. We work closely with NVIDIA, EC2 instance teams, AWS AI/ML services, and the Kubernetes community to deliver the most capable and reliable node platform for containerized workloads.
Inclusive Team Culture
Here at AWS, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences. Amazon's culture of inclusion is reinforced within our 14 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust.
Work/Life Balance
Our team puts a high value on work-life balance. It isn't about how many hours you spend at home or at work; it's about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. While we maintain 24/7 on-call coverage to support our critical infrastructure, we distribute the operational load fairly and invest heavily in automation and operational excellence to minimize toil and ensure sustainable on-call burden.
Mentorship & Career Growth
Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge sharing and mentorship. Whether you're learning about Linux kernel internals, GPU driver architecture, Kubernetes device plugins, or distributed systems operations, our senior members provide one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer - balancing innovation work (new GPU support, AI workload patterns) with operational excellence (automation, monitoring, reliability improvements)—and enable them to take on more complex technical leadership in the future.
Basic Qualifications
- 3+ years of non-internship professional software development experience- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field
- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
Preferred Qualifications
- Experience in Kubernetes, Docker or containers ecosystem, or experience that includes strong analytical skills, attention to detail, and effective communication abilities and experience in managing and troublshooting network- Knowledge of and experience with cloud infrastructure technologies
- Experience in Linux/RHEL, or experience in Kubernetes, Docker or containers ecosystem and experience with programming/scripting (Batch, VB, PowerShell, Java, C#, Chef, Perl, Ruby and/or PHP)
- Experience with CloudFormation, Chef, Puppet, Salt, or Ansible in production environments
- Experience delivering products against plan in a fast-paced, multi-disciplined, distributed-responsibility and often ambiguous environment
- Experience with enterprise architecture including virtualization technologies and distributed architecture
- Contributions to open-source projects, especially in the Kubernetes/CNCF ecosystem
- Strong operational mindset - experience with monitoring, observability, and incident management at scale
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, Seattle - 143,700.00 - 194,400.00 USD annually
About the company
amazon
Large Enterprise
Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.
Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.