Research Engineer - LLM Post-Training & Agents
Kaon (prev. FlowGPT) - San Francisco Bay Area
Hiring: Research Engineer - LLM Post-Training & Agents Company: Kaon (prev. FlowGPT) Location: San Francisco Bay Area Job Posted Time: 2026-09-11 17:59:20 Employment Type: Full-time Target Skills & Keywords: LLM Post-Training, SFT, DPO, GRPO, Reinforcement Learning, Preference Optimization, Reward Modeling, Agent Training, Tool Use, Memory Architectures, Long-Context Evaluation, Python, C/C++, A/B Testing, Data Quality, Conversational AI Experience: - Hands-on experience training or post-training large language models, including practical familiarity with reinforcement learning or preference optimization. - Experience building or evaluating LLM-based agents, or a strong interest backed by relevant projects. - Ability to turn an open-ended research problem into a clear hypothesis, an experiment, and a measurable result. - Ownership from implementation and debugging through evaluation and deployment. Required Skills: - Strong Python skills and solid software engineering fundamentals; C/C++ experience is a plus. - Build and iterate on LLM post-training pipelines, including supervised fine-tuning, preference optimization, and reinforcement learning (SFT, DPO, GRPO, and related methods). - Develop training and preference datasets from user interactions, with careful data quality controls and held-out evaluations. - Design reward models and feedback signals for narrative quality, instruction following, character consistency, personalization, and memory accuracy. - Build agent training environments and evaluations for tool use, memory, and long-horizon interactions. - Explore memory architectures, context management, and memory consolidation that improve an agent's behavior over time. - Run controlled experiments, analyze failures, and validate improvements through offline evaluations and online A/B tests. - Work with the engineering and product teams to bring research improvements into production. Qualifications: - Strong Python skills and solid software engineering fundamentals. - Hands-on experience training or post-training large language models. - Practical familiarity with reinforcement learning or preference optimization. - Experience building or evaluating LLM-based agents, or a strong interest backed by relevant projects. - Ability to turn an open-ended research problem into a clear hypothesis, an experiment, and a measurable result. - Ownership from implementation and debugging through evaluation and deployment. - Nice to have: Work on agent memory, personalized generation, continual learning, reward modeling, or long-context evaluation. - Nice to have: Experience with roleplay, virtual characters, interactive storytelling, or conversational AI. - Nice to have: Research publications or open-source contributions in relevant areas. Compensation: $200,000–$500,000 total compensation (base + equity), depending on experience and impact. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!