Machine Learning Engineer, Evaluations

Brahma Consulting Group - New York City Metropolitan Area

Hiring: Machine Learning Engineer, Evaluations Company: Brahma Consulting Group Location: New York City Metropolitan Area Job Posted Time: 2026-09-15 18:06:15 Employment Type: Full-time / On-site Target Skills & Keywords : Event-Driven, HBase, LLM, Move, Python, SQL About the job Required Skills: •An early-stage AI infrastructure company is hiring its first dedicated evaluation engineer. You would own evaluation end to end for a system whose outputs change over time and have no clean ground truth to compare against. The core problem is definitional before it is technical: deciding what "better" actually means for a representation that evolves, then building the machinery that measures it and keeps measuring it as the methods shift underneath you. •The system itself is a multi-agent harness with genuine black-box behavior. Measurement here means constructing instruments rather than picking metrics off a shelf. •This is a strong fit if you like reading traces until you find what is actually broken, and then shipping the fix yourself rather than filing it. •What you'll do •Design the evaluations. Define the scoring, decide what improvement means for a changing representation, and revise that definition as the product and research move. •Build the pipelines and harnesses. Data ingestion, labeling, versioning, reruns, judge models. The infrastructure that lets the team ask a new question on Monday and have an answer by Friday. •Run them and extract insight. Separate what is genuinely broken from what is merely easy to measure, and propose changes that raise output fidelity. •Build simulation agents at scale. Iteration speed is capped by how many entities can be modeled and measured simultaneously. Raising that ceiling is part of the job. •Own the full loop. The question, the pipeline, the rerun, the writeup. Nothing gets scoped and handed to someone else, and nothing sits waiting on another team's sprint. •What we're looking for •Three or more years in LLM evaluation, agentic optimization, or ML research. The range is wide and seniority is calibrated to depth rather than years. •Strong Python. Comfort with modern ML tooling, experiment configuration frameworks, streaming or event-driven data infrastructure, and SQL. •Experience building measurement from scratch in domains where labeled data does not exist. •High autonomy. The team is small and flat, structure is minimal, and engineers operate closer to founding engineers than to specialists. Curiosity across adjacent fields such as cognitive science, linguistics, or philosophy is genuinely valued rather than decorative. Compensation: •$240,000 •$300,000 Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!