Hardware Reliability Engineer
Meta - Fremont, CA
Hiring: Hardware Reliability Engineer Company: Meta Location: Fremont, CA Job Posted Time: 2026-09-09 17:58:20 Employment Type: Full-time Target Skills & Keywords: Hardware Reliability Engineering, DFR, DFMEA, Derating, Failure Analysis, Accelerated Life Testing (ALT), HALT, Weibull Analysis, MTBF Modeling, Reliability Statistics, Design Verification Testing, Environmental Stress Testing, Server Hardware, Storage Hardware, Networking Hardware, AI Compute Platforms, Silicon Reliability, Custom Silicon, ODM Collaboration, Contract Manufacturers, Data Center Hardware, Root Cause Investigation, Corrective Actions, Design of Experiments, Cross-functional Collaboration Experience: - 6+ years of experience in hardware reliability engineering, including failure analysis and reliability testing of infrastructure hardware - Experience collaborating with hardware suppliers and contract manufacturers to evaluate component reliability and enforce qualification standards - Experience communicating complex reliability findings and technical trade-offs to engineering and operations stakeholders through written reports and presentations Required Skills: - Lead DFR activities such as DFMEA and derating across AI, compute, and storage platforms - Develop reliability tests for compute, storage, server hardware, and networking modules to expose design weaknesses - Establish Design Verification tests to uncover environmental stress weaknesses and ensure designs meet lifetime reliability metrics - Work closely with ODMs to oversee test execution and suggest improvements based on lessons learned - Translate test results into meaningful product life metrics and flag unmet metrics - Apply reliability statistics to support decision-making and quantify risk - Lead development of internal reliability test infrastructure and design of experiments - Collaborate cross-functionally with Hardware Engineering, Release To Production, Thermal, and Failure Analysis teams to de-risk design issues - Apply reliability engineering methodologies such as FMEA, HALT, ALT, Weibull analysis, and MTBF modeling - Analyze field failure data and translate findings into actionable root cause investigations and corrective actions Qualifications: - Bachelor's degree in Electrical Engineering, Mechanical Engineering, or a related discipline - Preferred: MSc in Mechanical or Electrical Engineering or related disciplines - Preferred: Experience in silicon reliability and custom silicon - Preferred: Familiarity with data center environments - Preferred: First-hand knowledge of server rack hardware Compensation: Base pay range: $144,000.00/yr - $204,000.00/yr Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!