Intel

Intel

Posted via Workday

Reliability Engineer

Posted Aug 5, 2026

Role at a glance

Salary
$122.4K – $232.2K/yr
Location
2 Locations, Massachusetts, United States
Work arrangement
On-site
Employment
Full-time
Experience
at least 4-6 yrs experience
Education
BS/MS/PhD in EE/ME Reliability or related

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

The role defines and owns pod-level reliability specifications for large-scale data centers across hardware, thermal, and operational dimensions. It supports AI hardware solutions by translating system requirements into subsystem specifications and improving resilience, serviceability, and availability.

What You'll Do

  • Define and maintain pod-level reliability and availability specifications and targets for compute, memory, storage, network, power, and...
  • Translate system and SLA requirements into pod and subsystem reliability specifications and flow requirements to silicon, platform, and...
  • Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and specification updates.
  • Architect RAS features and graceful degradation or redundancy against pod-level specifications.
  • Partner with facilities on pod power and cooling redundancy, thermal margins, and disaster-recovery readiness.
  • Establish HALT/HASS, burn-in, and qualification processes; track field returns and KPIs against pod specifications.

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

BS/MS/PhD in EE/ME Reliability or related; at least 4-6 yrs experience; experience authoring and owning reliability specs and requirement flow-down; strong RAS, FMEA, statistical reliability (Weibull, FIT) skills; experience with large-scale fleet telemetry and thermal/power redundancy.

Required

  • Experience authoring and owning reliability specs and requirement flow-down
  • Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills
  • Experience with large-scale fleet telemetry and thermal/power redundancy

Preferred

  • AI cluster operations
  • Data analytics (Python/SQL)

Original job description

Content provided by the employer

Job Details:

Job Description: 

Join us to help build the next generation of AI hardware solutions. You will be part of a highly skilled, agile team developing cutting-edge hardware for the AI domain, where we push the boundaries of what silicon can do for emerging AI workloads. With a startup-like culture, we move quickly and give engineers the opportunity to drive significant technical and business impact. 

We are continuously developing modern and effective working methods, including hands-on adoption of AI tools throughout the chip development flow.  

Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions.

Responsibilities:

  • Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems.

  • Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams.

  • Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.

  • Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.

  • Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness.

  • Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.

Qualifications:

Minimum Qualifications:

  • BS/MS/PhD in EE/ME Reliability or related; and/or at least 4-6 yrs experience.

  • Experience authoring and owning reliability specs and requirement flow-down.

  • Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills.

  • Experience with large-scale fleet telemetry and thermal/power redundancy.

Preferred Qualifications:

  • AI cluster operations, data analytics (Python/SQL).

          

Job Type:

Experienced Hire

Shift:

Shift 1 (United States of America)

Primary Location: 

US, Massachusetts, Beaver Brook

Additional Locations:

US, California, Santa Clara

Posting Statement:

All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.

Position of Trust

N/A

Benefits

We offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel.

 

 

Annual Salary Range for jobs which could be performed in the US: $122,440.00-232,190.00 USD

 

 

The range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.

 

 

Work Model for this Role

This role will require an on-site presence. * Job posting details (such as work model, location or time type) are subject to change.

*

ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.
Intel

About the company

Intel

Large Enterprise

Intel Corporation is a global leader in computing innovation, renowned for its advanced semiconductor manufacturing and technology solutions. Founded in 1968, the company is primarily known for developing microprocessors that power a vast range of computing devices, from personal computers to data centers. With a strong commitment to research and development, Intel continuously drives advancements in artificial intelligence, cloud computing, and the Internet of Things (IoT), aiming to enhance connectivity and computing capabilities around the world. Intel's mission is to create technology that enhances lives and enables progress across various industries.