ML Lead, AI Data Labeling
NewtonX - United States
Hiring: ML Lead, AI Data Labeling Company: NewtonX Location: United States Job Posted Time: 2026-09-10 14:07:14 Target Skills & Keywords : LLM, Project Management, Python, Reinforcement Learning, Stripe About the job Experience: •2 years working hands-on with RL, including how tool-use trajectories are rewarded and evaluated. If you're not fluent in RL, this isn't the role — it's the foundation of the core judgment you'll be making. •At least 2 years working hands-on with RL, including how tool-use trajectories are rewarded and evaluated. If you're not fluent in RL, this isn't the role — it's the foundation of the core judgment you'll be making. Required Skills: •This role is the technical owner of data quality. You will partner directly with the Program Lead and act as technical lead in communicating with clients, interpreting their core AI model testing goals and assisting the Program Lead in creating concrete technical specs that will accomplish these goals. •You are hands-on and close to the work. This is a foundational role — what it covers will grow as the business does. •In This Role You Will •Own the judgment of whether task designs and rubrics produce a useful training signal for the consuming method (SFT, RLHF, RLVR, agentic RL, CoT, evals). Catch mismatches between what the data rewards and what the customer is actually training for. •Design the task, environment, and rubric structure for agentic workflows. •Own the ML-level validation of the dataset before delivery — not just whether individual submissions meet spec. •Technical Feedback Loop with Operations •Partner with the Program Lead to convert customer requirements into concrete technical specs: expert profiles, screener trees, task interfaces, task templates, QC rubrics, statistical thresholds. Qualifications: •Deep applied ML experience centered on post-training human data — you've owned a human-data workstream as an applied scientist or ML engineer. •Working fluency across modern LLM post-training and evaluation: SFT, RLHF/preference data quality, RLVR, chain-of-thought, eval harness construction, contamination handling, statistical significance, and agentic/tool-use evaluation. •Genuine understanding of how training data becomes model behavior — you can reason about what a model will learn from a given dataset, not just whether the data meets spec. •Strong programming foundation: read and reason about an eval harness, write Python comfortably, work with model APIs, prototype scoring pipelines. Not a production engineer, but not hands-off. •Statistical fluency: you know when an effect is real vs. noise and can defend a sample size or significance threshold. •Client-facing presence: you've defended technical design choices in real time to skeptical audiences and adjusted scope without losing rigor. Range matters — you can talk to a Series B CTO and a Fortune 100 AI lead in the same week. •Strong written communication: methodology sections, technical reports, and specs that hold up to expert review. •If the profile above describes you and your passions, we'd love to hear from you! Compensation: •$180,000 - $260,000 / year •Competitive benefits and rewards package •Comprehensive Benefits: Excellent medical, dental, and vision insurance Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!