Research Program Manager - Model Evals and Safety

Reflection - San Francisco, CA

Hiring: Research Program Manager - Model Evals and Safety Company: Reflection Location: San Francisco, CA Job Posted Time: 2026-09-15 17:59:14 Employment Type: Full-time Target Skills & Keywords: Model Evaluations, AI Safety, Research Program Management, Evaluation Frameworks, Red-Teaming, Alignment Research, ML Engineering, Safety Operations, Technical Program Management, Pre-training, Mid-training, Post-training, Third-Party Assessments, Industry Safety Frameworks, Regulatory Compliance, Blameless Post-Mortems Experience: - 7+ years of experience in technical program management, research operations, or ML engineering - Demonstrated experience standing up new functions, teams, or programs from scratch - Familiarity with the landscape of model evaluation and AI safety, including evaluation methodologies, red-teaming, alignment research, and the evolving regulatory and industry landscape Required Skills: - Ability to build foundational infrastructure for model evals and safety, including evaluation frameworks, tooling requirements, and operational processes - Experience standing up model safety operations, including workflows, review cadences, and decision frameworks - Ability to partner with research and engineering leads across pre-training, mid-training, and post-training to embed safety and evaluation checkpoints - Experience driving scoping and prioritization of eval science and eval infrastructure investments - Ability to establish engagement with external safety ecosystem, including third-party assessments, academic partnerships, and industry safety frameworks - Strong communication skills to represent safety posture to external stakeholders - Ability to create visibility and reporting structures for leadership on model safety status, evaluation coverage, and open risks - Commitment to championing a culture of blameless post-mortems and continuous learning Qualifications: - First-responder mentality with the ability to jump in, assess situations, cut through noise, align stakeholders, and drive resolution - High-leverage leader and operator who can bring clarity to ambiguity and drive decisions when the path forward is unclear - Force multiplier who ensures work across multiple teams connects into a coherent whole - 0-to-1 builder mindset, capable of defining and standing up a new function from the ground up - Not a project tracker, but a strategic operator embedded directly with research and infrastructure teams Compensation: Not specified in the provided job text. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!