Role at a glance
- Salary
- $226K – $369K/yr
- Location
- Mountain View, California, United States
- Work arrangement
- Remote
- Employment
- Contract
- Experience
- 10+ years of experience in software engineering, infrastructure engineering, distributed systems, SRE, production engineering, or...
- Education
- BA/BS degree in Computer Science or related technical field, or equivalent practical experience
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Reliability Infrastructure team defines and drives reliability strategy, standards, and practices for LinkedIn’s critical systems. This Principal Staff Software Engineer serves as a company-wide technical authority, shaping service criticality, resilient architecture, and the safe operation of AI-enabled systems across infrastructure and product engineering.
What You'll Do
- Define and drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering
- Lead adoption and evolution of service criticality models based on business impact and blast radius
- Serve as a technical authority for architecture decisions related to reliability, resiliency, availability, and failure handling
- Partner with infrastructure and product engineering teams to improve system design, reduce incident risk, and strengthen operational...
- Establish and evolve reliability standards including SLOs, SLIs, uptime expectations, redundancy, monitoring, alerting, and failover...
- Drive the strategy for applying LLMs to alert triage, root cause analysis, and incident summarization at scale
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
BA/BS degree in Computer Science or related technical field, or equivalent practical experience; 10+ years of experience in software engineering, infrastructure engineering, distributed systems, SRE, production engineering, or reliability engineering; 5+ years in a technical leadership, architect, or principal-level engineering role; experience with large-scale distributed systems, reliability standards, high availability, redundancy, fault tolerance, failure modes, resiliency patterns, and cross-organizational technical influence; software engineering experience in Java, Go, C++, Python, or similar.
Required
- BA/BS degree in Computer Science or related technical field, or equivalent practical experience
- 10+ years of experience in software engineering, infrastructure engineering, distributed systems, SRE, production engineering, or...
- 5+ years of experience in a technical leadership, architect, or principal-level engineering role
- Experience designing, building, or operating large-scale distributed systems
- Experience defining or driving reliability standards such as SLOs, SLIs, uptime targets, incident reduction, or operational readiness...
- Understanding of high availability, redundancy, fault tolerance, failure modes, and resiliency patterns
- Experience influencing architecture and engineering practices across multiple teams or organizations
- Software engineering experience in one or more languages such as Java, Go, C++, Python, or similar
Preferred
- MS or PhD in Computer Science or related technical field
- Experience operating at company-wide or large org-wide scope as a reliability, infrastructure, SRE, or production engineering technical...
- Experience with tiered service criticality models, priority-based reliability frameworks, or large-scale reliability governance
- Deep expertise in distributed systems reliability, service resilience, and failure isolation at scale
- Experience with incident management, postmortems, operational reviews, and driving long-term corrective actions across organizations
- Experience with observability, monitoring, alerting, capacity planning, disaster recovery, and multi-region failover strategies
- Background in mature SRE, production engineering, platform reliability, or infrastructure resilience environments
- Experience driving reliability transformations across large engineering organizations
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Company Description
LinkedIn is the world’s largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. We’re also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture that’s built on trust, care, inclusion, and fun – where everyone can succeed.
Join us to transform the way the world works.
Job Description
At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. This role may be remote or hybrid. At LinkedIn, hybrid roles are performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team. Remote roles are performed from the designated home work location upon time of hire, and any changes to this home work location requires a review of remote status and approval.
LinkedIn’s Reliability Infrastructure team is responsible for defining and driving the reliability strategy, standards, and practices that keep LinkedIn’s most critical systems stable, resilient, and available at massive scale.
As a Principal Staff Software Engineer, Reliability Infrastructure, you will serve as a senior technical authority for reliability across LinkedIn Engineering. You will help define how critical services are designed, built, operated, and measured, partnering broadly across infrastructure and product engineering teams to improve resiliency, reduce incidents, and raise the reliability bar across the company.
A key focus of this role is driving the adoption and evolution of LinkedIn’s service criticality framework, including reliability expectations for the most business-critical systems. You will help classify services based on criticality and blast radius, define appropriate reliability standards, and influence system architecture to ensure the right levels of availability, redundancy, observability, and failure handling are in place.
As AI-assisted software development, agent-based automation, and autonomous operational systems become more prevalent, this role will also help define how LinkedIn safely builds and operates reliable AI-enabled systems. You will shape standards for evaluating, deploying, monitoring, and governing AI-generated code and agentic workflows, ensuring that automation introduced into critical environments is observable, explainable, auditable, and designed with appropriate safeguards, rollback mechanisms, and human oversight.
This is not a traditional SRE role focused on operating a single service or team. It is a company-wide technical leadership role for someone with deep distributed systems expertise, strong reliability judgment, and the ability to influence architecture and engineering practices across large organizations.
Responsibilities
Define and drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering
Lead adoption and evolution of service criticality models that set reliability expectations based on business impact and blast radius
Serve as a technical authority for architecture decisions related to reliability, resiliency, availability, and failure handling
Partner with infrastructure and product engineering teams to improve system design, reduce incident risk, and strengthen operational readiness
Identify high-risk systems and drive cross-organizational initiatives to improve reliability of critical services
Establish and evolve reliability standards including SLOs, SLIs, uptime expectations, redundancy, monitoring, alerting, and failover patterns
Influence engineering culture by promoting reliability-focused design, incident review rigor, and postmortem-driven improvements
Provide architectural guidance and mentorship to senior engineers and technical leaders across teams
Balance technical strategy, hands-on engineering judgment, and cross-functional influence to drive measurable improvements in site stability
Help shape how LinkedIn builds and operates resilient systems as the platform continues to scale
Drive the strategy for applying LLMs to alert triage, root cause analysis, and incident summarization at scale, ensuring systems are explainable, auditable, and safe to operate autonomously in Ring0/Ring1 environments.
Qualifications
Basic Qualifications
BA/BS degree in Computer Science or related technical field, or equivalent practical experience
10+ years of experience in software engineering, infrastructure engineering, distributed systems, SRE, production engineering, or reliability engineering
5+ years of experience in a technical leadership, architect, or principal-level engineering role
Experience designing, building, or operating large-scale distributed systems
Experience defining or driving reliability standards such as SLOs, SLIs, uptime targets, incident reduction, or operational readiness frameworks
Understanding of high availability, redundancy, fault tolerance, failure modes, and resiliency patterns
Experience influencing architecture and engineering practices across multiple teams or organizations
Software engineering experience in one or more languages such as Java, Go, C++, Python, or similar
Preferred Qualifications
MS or PhD in Computer Science or related technical field
Experience operating at company-wide or large org-wide scope as a reliability, infrastructure, SRE, or production engineering technical leader
Experience with tiered service criticality models, priority-based reliability frameworks, or large-scale reliability governance
Deep expertise in distributed systems reliability, service resilience, and failure isolation at scale
Experience with incident management, postmortems, operational reviews, and driving long-term corrective actions across organizations
Experience with observability, monitoring, alerting, capacity planning, disaster recovery, and multi-region failover strategies
Background in mature SRE, production engineering, platform reliability, or infrastructure resilience environments
Experience driving reliability transformations across large engineering organizations
Executive-level communication skills with the ability to align technical decisions to business impact
Demonstrated ability to influence technical direction without direct authority and drive adoption of standards across teams
Prior work on self-healing or auto-remediation systems at companies with large-scale infrastructure (hyperscalers, large internet companies).
Experience defining reliability, safety, or governance standards for AI-enabled systems, agentic workflows, or AI-assisted software development
Familiarity with LLM and agent evaluation, production monitoring, guardrails, human oversight, and rollback strategies for autonomous systems
Suggested Skills
Distributed systems reliability
Site Reliability Engineering / Production Engineering
High availability and fault tolerance
SLO / SLI design
Incident management and postmortem practices
Observability and monitoring
Resiliency engineering
AI agent reliability and safety
LLM and agent evaluation
AI observability and monitoring
Autonomous remediation and guardrails
Large-scale infrastructure
Cross-organizational technical leadership
Reliability standards and governance
Familiarity with emerging standards and frameworks for agentic AI safety and evaluation (evals pipelines, red-teaming autonomous systems, policy guardrails for production agents).
LinkedIn is committed to fair and equitable compensation practices.
The pay range for this role is $226,000 to $369,000. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor.
The total compensation package for this position may also include annual performance bonus, stock, benefits and/or other applicable incentive compensation plans. For more information, visit https://careers.linkedin.com/benefits.
Additional Information
Equal Opportunity Statement
We seek candidates with a wide range of perspectives and backgrounds and we are proud to be an equal opportunity employer. LinkedIn considers qualified applicants without regard to race, color, religion, creed, gender, national origin, age, disability, veteran status, marital status, pregnancy, sex, gender expression or identity, sexual orientation, citizenship, or any other legally protected class.
LinkedIn is committed to offering an inclusive and accessible experience for all job seekers, including individuals with disabilities. Our goal is to foster an inclusive and accessible workplace where everyone has the opportunity to be successful.
If you need a Reasonable Accommodation to search for a job opening, apply for a position, or participate in the interview process, connect with us and describe the specific Accommodation requested for a disability-related limitation.
Fill out an Accommodation request here: https://app.smartsheet.com/b/form/b660a0327d044969abfd7a4e73d15c36
Reasonable accommodations are modifications or adjustments to the application or hiring process that would enable you to fully participate in that process. Examples of reasonable accommodations include but are not limited to:
- Documents in alternate formats or read aloud to you
- Having interviews in an accessible location
- Being accompanied by a service dog
- Having a sign language interpreter present for the interview
A request for an accommodation will be responded to within three business days. However, non-disability related requests, such as following up on an application, will not receive a response.
LinkedIn will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay or the pay of another employee or applicant. However, employees who have access to the compensation information of other employees or applicants as a part of their essential job functions cannot disclose the pay of other employees or applicants to individuals who do not otherwise have access to compensation information, unless the disclosure is (a) in response to a formal complaint or charge, (b) in furtherance of an investigation, proceeding, hearing, or action, including an investigation conducted by LinkedIn, or (c) consistent with LinkedIn's legal duty to furnish information.
San Francisco Fair Chance Ordinance
Pursuant to the San Francisco Fair Chance Ordinance, LinkedIn will consider for employment qualified applicants with arrest and conviction records.
Pay Transparency Policy Statement
As a federal contractor, LinkedIn follows the Pay Transparency and non-discrimination provisions described at this link: https://lnkd.in/paytransparency.
Global Data Privacy Notice and Compliance Posters for Job Candidates
Please use this link to access documents that provide information about how LinkedIn handles the personal data of employees and job applicants, as well as the E-Verify Participation Notice and the Department of Justice Immigrant and Employee Rights Section Right to Work posters: https://www.linkedin.com/legal/candidate-portal.
About the company
Large Enterprise
LinkedIn is a globally recognized social networking platform designed specifically for professionals to connect, share, and grow their careers. Founded in 2002, it enables users to build professional profiles, network with industry peers, and discover job opportunities across various sectors. LinkedIn also offers a suite of tools for companies, including talent recruitment solutions and branding opportunities, helping organizations to engage with potential candidates and promote their corporate identity. With millions of users worldwide, LinkedIn is a vital resource for career development and professional networking.
LinkedIn is a globally recognized social networking platform designed specifically for professionals to connect, share, and grow their careers. Founded in 2002, it enables users to build professional profiles, network with industry peers, and discover job opportunities across various sectors. LinkedIn also offers a suite of tools for companies, including talent recruitment solutions and branding opportunities, helping organizations to engage with potential candidates and promote their corporate identity. With millions of users worldwide, LinkedIn is a vital resource for career development and professional networking.