Member of Technical Staff - Post-Training and RL

SpaceXAI - Palo Alto, CA

Hiring: Member of Technical Staff - Post-Training and RL Company: SpaceXAI Location: Palo Alto, CA Job Posted Time: 2026-09-10 12:23:53 Target Skills & Keywords : Reinforcement Learning About the job Required Skills: •Work on the most critical post-training and reinforcement learning challenges at any given time — including reward modeling, preference optimization (RLHF/DPO), and RL for improving reasoning, truthfulness, and real-world capabilities. •Get clarity on your first project before an offer. Qualifications: •You believe truth-seeking AI is the most important and challenging problem. •You are obsessed about building incredibly useful models through post-training and RL techniques. •You are a power user of AI models and eager to push the boundaries of what’s possible with reinforcement learning and alignment methods. •If you previously worked on post-training, RLHF, or trained models used by millions of people it’s a big plus, but relevant experience is not required. •You take pride in your work and thrive in meritocratic environments. Compensation: •$180,000 - $600,000 / year Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!