Senior Research Engineer, Computer Vision (LFV/WFM)
Toyota Research Institute - Los Altos, CA
Hiring: Senior Research Engineer, Computer Vision (LFV/WFM) Company: Toyota Research Institute Location: Los Altos, CA Job Posted Time: 2026-09-10 10:29:45 Target Skills & Keywords : AI, AWS, Data Pipeline, Deep Learning, Diffusion Models, Docker, EC2, Kubernetes, Linux, Machine Learning, Node.js, PyTorch, Python, SageMaker About the job Experience: •3 years of relevant experience and strong software engineering skills. Required Skills: •Collaborate directly with research scientists to implement, iterate on, and evaluate new architectures, objectives, datasets, and training strategies. Translate research prototypes into clean, maintainable, reusable code that will be shared across multiple TRI teams and the broader Toyota ecosystem. •Build and maintain scalable pipelines for ingesting, converting, validating, and serving heterogeneous datasets, across robotics and autonomous driving, into unified training-ready formats. Track and integrate new public and internal datasets as they become available. •Support and optimize large-scale distributed training of world foundation models on multi-GPU and multi-node clusters. Manage experiment workflows, profiling, debugging, and hyperparameter sweeps to ensure optimal performance in a timely manner. •Develop tools for dataset inspection, experiment tracking, model evaluation, GPU resource management, and visualization. Automate repetitive workflows to improve team velocity. •Interface directly with other TRI teams and Toyota affiliates to set up shared pipelines, onboard their data, and support joint training and evaluation efforts. •Produce maintainable, well-documented code. Contribute to internal tooling and open-source releases to the scientific community. Qualifications: •Master’s or PhD in Computer Science, Electrical Engineering, Machine Learning, or a related field, with a minimum of 3 years of relevant experience and strong software engineering skills. •Deep proficiency in Python, PyTorch, and the Unix/Linux toolchain. Comfort working in terminal-heavy, SSH-based workflows on shared GPU clusters. •Applied hands-on capability in large-scale deep learning training, including distributed training (DDP, FSDP, DeepSpeed, or similar), GPU profiling, and debugging training failures at scale. •You are proactive, self-directed, and comfortable operating with ambiguity in a research-driven environment that spans multiple divisions. •You are a reliable teammate who communicates clearly and takes ownership of problems end-to-end. •Operational familiarity with standard data formats and collection pipelines as well as simulation environments. •Proficiency with modern AI-assisted development tools (e.g., Copilot, Cursor, Claude Code) for accelerating engineering workflows. •Track record of contributions to open-source projects or publications at top venues is a plus but not required. •Please include links to any relevant open-source contributions or technical project write-ups with your application. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!