AI Evaluation Engineer

Zof AI - San Francisco, CA

**Hiring: AI Evaluation Engineer** **Company:** Zof AI **Location:** San Francisco, CA **Job Posted Time:** 2026-09-02 17:57:42 **Employment Type:** Full-time, On-site **Target Skills & Keywords:** Evals, Verification, AI Quality, Testing, QA **Experience:** Mid to Senior **Required Skills:** - Experience testing, evaluating, or QA-ing complex software systems - Understanding of how LLM and agent systems fail - Strong analytical rigor and skepticism - Ability to write code to build harnesses and automation - Attention to detail and a high-quality bar - Clear written and verbal communication - Comfort operating in a fast-moving environment - High ownership **Qualifications:** - Experience building LLM evals, benchmarks, or test infrastructure (Nice to have) - QA, SDET, or test automation background (Nice to have) - Domain expertise in a vertical where correctness matters (Nice to have) - Experience with statistical evaluation methods (Nice to have) **Compensation:** Competitive salary + meaningful equity **About the Role:** Zof AI is looking for an AI Evaluation Engineer to design and implement tests that validate AI systems beyond demos. This role is central to ensuring AI products meet high standards by creating eval suites, verification harnesses, and quality gates. The ideal candidate is detail-oriented, skeptical by nature, and driven to turn "it seems to work" into measurable evidence. **Perks & Benefits:** - Premium AI development tools (MacBook Pro, Cursor Ultra, Claude Code Ultra, OpenAI Codex Max) - High-performance AI product environment - Close collaboration with leadership, engineering, and customers - Opportunity to work in the San Francisco AI ecosystem - Wellness and productivity support Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!