Infrastructure Reliability Engineer
Bright Vision Technologies - Leander, TX
Hiring: Infrastructure Reliability Engineer Company: Bright Vision Technologies Location: Leander, TX Job Posted Time: 2026-09-15 19:56:55 Employment Type: Full-time, Direct W2 Target Skills & Keywords: Data Engineering, AI/ML Infrastructure, Large-Scale Data Pipelines, Petabyte-Scale Systems, Distributed Systems, Python, Spark, Ray, Beam, Dataset Versioning, Data Quality, High-Throughput Data Loading, GPU Utilization, Multimodal Data, Data Privacy, Observability, CI/CD Experience: - 6+ years of data engineering experience, with significant work supporting ML or AI workloads - Hands-on experience operating petabyte-scale storage and pipeline systems - Experience with dataset versioning, lineage, and reproducibility for ML workflows - Familiarity with high-throughput data loading for accelerator-based training Required Skills: - Strong proficiency in Python and at least one JVM or systems language - Deep experience with modern data processing frameworks such as Spark, Ray, or Beam - Strong understanding of distributed systems, data modeling, and storage formats - Strong software engineering practices including testing, CI/CD, and code review - Excellent communication and cross-functional collaboration skills - Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement workflows - Build ingestion systems for diverse modalities including text, image, audio, video, and structured signals - Implement data cleaning, deduplication, filtering, and quality assurance at petabyte scale - Develop dataset versioning, lineage, and provenance tracking systems suitable for reproducible training - Build high-throughput data loading systems that maximize GPU utilization during training - Implement labeling workflows, active learning pipelines, and human-in-the-loop data improvement systems - Design storage architectures balancing cost, throughput, and latency across data tiers - Build evaluation dataset construction pipelines with strict integrity and contamination controls - Implement data privacy, redaction, and consent enforcement throughout the pipeline - Drive observability of data quality, drift, and pipeline health across the AI data estate - Optimize cost and performance through compression, format selection, and caching strategies Qualifications: - Bachelor's or Master's degree in Computer Science or a related field - Experience with multimodal datasets at large scale (preferred) - Familiarity with data quality tooling and dataset evaluation methodology (preferred) - Exposure to privacy-preserving data systems and regulated data handling (preferred) - Open-source contributions to data infrastructure projects (preferred) - Experience supporting frontier model training pipelines (preferred) Compensation: $125,000–$170,000 Annually Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!