Senior Data Engineer

Toyota Research Institute - Los Altos, CA

Hiring: Senior Data Engineer Company: Toyota Research Institute Location: Los Altos, CA Job Posted Time: 2026-09-10 10:29:45 Employment Type: Hybrid Target Skills & Keywords : AI, AWS, C++, CI/CD, ETL, Feature Store, GCP, Kubeflow, Machine Learning, Protobuf, Python, S3, SQL, SageMaker, Spark About the job Experience: •8+ years of experience building data-intensive software systems, ideally in robotics, autonomous driving, or large-scale ML environments Required Skills: •Design and implement scalable, production-grade pipelines for data ingestion, transformation, storage, and retrieval from vehicle fleets and simulation environments •Build internal tools and services for data labeling, curation, indexing, and cataloging across large and diverse datasets •Partner cross-functionally with ML researchers, autonomy engineers, and data scientists to design schemas and APIs that power model training, evaluation, and debugging •Develop and maintain feature stores, metadata systems, and versioning infrastructure for structured and unstructured data •Actively facilitate and support generation and integration of synthetic datasets with real-world logs to enable hybrid training and simulation workflows •Optimize pipelines for cost, latency, and traceability, ensuring reproducibility and consistency across environments •Partner with simulation and cloud platform teams to automate workflows for closed-loop testing, scenario mining, and performance analytics Qualifications: •Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related field •Proficient in Python, SQL, and familiar with C++ •Strong knowledge of cloud-native architectures, including AWS services (e.g., S3, or equivalents (Google Cloud platform) •Operational familiarity with sensor data types (camera, lidar, radar, GPS/IMU) and common data serialization formats (e.g., protobuf. ROS2bag, MCAP) •Comprehensive expertise in data quality, observability, and lineage in high-volume systems •Track record of building reliable and performant infrastructure that supports both ad-hoc exploration and repeatable production workflows •Operational familiarity with ML pipeline orchestration frameworks (e.g. Kubeflow, SageMaker, etc) •Exposure to synthetic data generation, simulation logging, or scenario replay pipelines •Strong software engineering fundamentals, CI/CD, testing, code review, and service deployment best practices •Please include links to any relevant open-source contributions or technical project write-ups with your application. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!