Role at a glance
- Salary
- $151.2K – $204.6K/yr
- Location
- Redmond, Washington, United States
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- 5+ years of systems design, software development, operations, automation, and process improvement experience
- Visa support
- candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the...
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Network Reliability team operates a global Service Provider network connecting ground stations, data centers, and Amazon Leo’s satellite constellation. This role focuses on maintaining reliable connectivity across space-based and terrestrial infrastructure and improving customer experience through monitoring, incident response, and operational automation.
What You'll Do
- Build and maintain processes, procedures, and tooling to monitor, troubleshoot, and operate the global Service Provider network
- Develop an end-to-end understanding of network architecture across satellite constellation and terrestrial systems
- Establish and enforce incident management and problem management processes
- Engage with service owners to resolve incidents and coordinate cross-functional efforts
- Prioritize and track initiatives across multiple teams to eliminate recurring problems
- Develop proactive monitoring and alerting solutions to improve network reliability and customer experience
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Required
- 5+ years of systems design, software development, operations, automation, and process improvement experience
- 5+ years of deploying and operating in a Linux/Unix environment experience
- 5+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby experience
- Experience with network troubleshooting tools (telnet, test-netconnection, tracert, tracetcp, iperf, ntttcp, dig, and packet capture tools), or experience with automation and any version control tools and experience in managing and...
- Experience with incident management and troubleshooting in production environments
Preferred
- Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite...
- Experience designing or architecting (design patterns, reliability and scaling) of new and existing systems
- Experience contributing to the definition and implementation of automation opportunities within an operations environment
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Are you passionate about building and operating networks at unprecedented scale? Join our Network Reliability team to shape the future of a global Service Provider network supporting 3,236 satellites and terrestrial infrastructure. You'll be at the forefront of ensuring reliable connectivity that directly impacts customer experience across our constellation.
Export Control Requirement:
Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum.
Key job responsibilities
• Build and maintain processes, procedures, and tooling to monitor, troubleshoot, and operate a global Service Provider network at scale
• Develop comprehensive understanding of end-to-end network architecture, including control and data planes across satellite constellation and terrestrial systems
• Establish and enforce incident management and problem management processes to ensure rapid resolution
• Engage with service owners to resolve incidents and coordinate cross-functional efforts
• Prioritize and track initiatives across multiple teams to systematically eliminate recurring problems
• Develop proactive monitoring and alerting solutions to continuously improve network reliability and customer experience
A day in the life
You'll work at the intersection of space-based and terrestrial networking, tackling unique challenges that few engineers ever encounter. Your day might include analyzing network telemetry from thousands of satellites, collaborating with hardware teams on next-generation designs, or developing automation that prevents issues before they impact customers. You'll have the autonomy to identify problems, propose solutions, and drive them to completion while working with talented engineers across the organization.
About the team
Our Network Reliability team operates one of the most complex networks in existence—connecting ground stations, data centers, and a constellation of satellites to deliver seamless connectivity. We're a collaborative group of curious problem-solvers who thrive on technical challenges and are committed to operational excellence. We value innovation, encourage experimentation, and believe the best solutions come from diverse perspectives and rigorous technical debate.
Basic Qualifications
- 5+ years of systems design, software development, operations, automation, and process improvement experience- 5+ years of deploying and operating in a Linux/Unix environment experience
- 5+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby experience
- Experience with network troubleshooting tools (telnet, test-netconnection, tracert, tracetcp, iperf, ntttcp, dig, and packet capture tools), or experience with automation and any version control tools and experience in managing and troublshooting network
- Experience with incident management and troubleshooting in production environments
Preferred Qualifications
- Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations- Experience designing or architecting (design patterns, reliability and scaling) of new and existing systems
- Experience contributing to the definition and implementation of automation opportunities within an operations environment
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, Redmond - 151,200.00 - 204,600.00 USD annually
About the company
amazon
Large Enterprise
Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.
Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.