Hardware Reliability Engineer

Meta - Menlo Park, CA

Hiring: Hardware Reliability Engineer Company: Meta Location: Menlo Park, CA Job Posted Time: 2026-09-09 17:58:20 Employment Type: Full-time Target Skills & Keywords: Hardware Reliability Engineering, DFR, DFMEA, Derating, Failure Analysis, Accelerated Life Testing (ALT), HALT, Weibull Analysis, MTBF Modeling, Reliability Statistics, Design Verification Testing, Environmental Stress Testing, Server Hardware, Storage Hardware, Networking Hardware, AI Platforms, Compute Platforms, Silicon Reliability, Custom Silicon, Data Center Environments, Server Rack Hardware, ODM Collaboration, Contract Manufacturers, Root Cause Investigation, Corrective Actions, Design of Experiments, Cross-functional Collaboration Experience: - 6+ years of experience in hardware reliability engineering, including failure analysis and reliability testing of infrastructure hardware - Experience collaborating with hardware suppliers and contract manufacturers to evaluate component reliability and enforce qualification standards - Experience communicating complex reliability findings and technical trade-offs to engineering and operations stakeholders through written reports and presentations Required Skills: - Lead DFR activities such as DFMEA and derating across AI, compute, and storage platforms - Develop reliability tests to expose design weaknesses across compute, storage, server hardware, and networking modules - Establish Design Verification tests to uncover environmental stress weaknesses in server design and ensure designs meet lifetime reliability metrics - Work closely with ODMs to ensure tests are executed as planned and suggest improvements based on lessons learned from previous platforms - Translate test results into meaningful product life metrics and highlight shortcomings in metrics not met - Utilize reliability statistics to support decision making and quantify risk - Lead development of internal reliability test infrastructure to support initiatives and design of experiments - Collaborate cross-functionally with Hardware Engineering, Release To Production, Thermal, and Failure Analysis teams to de-risk design issues - Apply reliability engineering methodologies such as FMEA, HALT, ALT, Weibull analysis, and MTBF modeling to infrastructure hardware - Analyze field failure data and translate findings into actionable root cause investigations and corrective actions Qualifications: - Bachelor's degree in Electrical Engineering, Mechanical Engineering, or a related discipline - Experience applying reliability engineering methodologies such as FMEA, HALT, ALT, Weibull analysis, and MTBF modeling to infrastructure hardware - Experience analyzing field failure data and translating findings into actionable root cause investigations and corrective actions - Experience collaborating with hardware suppliers and contract manufacturers to evaluate component reliability and enforce qualification standards - Experience communicating complex reliability findings and technical trade-offs to engineering and operations stakeholders through written reports and presentations Preferred Qualifications: - Experience in silicon reliability and working on custom silicon is a plus - MSc in Mechanical or Electrical Engineering or related disciplines - Familiarity with data center environments is beneficial - First-hand knowledge of server rack hardware is preferred Compensation: - Base pay range: $144,000.00/yr - $204,000.00/yr - Actual pay will be based on skills and experience Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!