Principal ML Engineer
Mimecast - Columbus, OH
Hiring: Principal ML Engineer Company: Mimecast Location: Columbus, OH Job Posted Time: 2026-09-17 00:04:05 Employment Type: Hybrid Target Skills & Keywords : AWS, Compliance, FastAPI, Hugging Face, IAM, Kinesis, Kubernetes, LLM, Lambda, Linear, Machine Learning, Make, Move, NLP, ONNX, PyTorch, Python, S3, SageMaker, Terraform, Transformers, Triton, containerd About the job Required Skills: •Develop and own ML systems end to end, from data sourcing, cleaning, and labeling strategy through feature engineering, model development, deployment, and monitoring. •Set the ML architecture across model design, serving, and surrounding systems, optimizing accuracy, latency, and throughput for highly imbalanced threat-detection data. •Set technical direction for production model serving using AWS SageMaker, NVIDIA Triton Inference Server, ensemble/KServe patterns, hardened container images, and integration with enrichment and gateway layers. •Benchmark and prototype alternatives to de-risk major decisions, then give leadership defensible technical recommendations. •Establish reproducible ML standards, including versioned datasets, region-partitioned data, and shared experimentation workflows. •Make model observability and efficacy measurement first-class concerns through distributed tracing, threshold-independent metrics, raw-payload capture, and monitoring for real regressions. •Own capacity planning and rollout strategy, including throughput per core or GPU, utilization headroom, peak-load provisioning, and phased regional canary or shadow deployments. •Diagnose production incidents, close the structural gaps they expose, and act as a primary reviewer and mentor across the ML codebase. Qualifications: •Breadth across transformer architectures, RNNs, CNNs, generalized linear models, and gradient-boosted trees, with the judgment to select the right approach for the problem rather than defaulting to the largest model. •Deep Python proficiency and strong command of PyTorch, Hugging Face transformers, and NLP tooling, plus working knowledge of ONNX Runtime, quantization such as FP16, and inference acceleration. •Practical experience utilizing dense and lexical retrieval, including embeddings, vector indexes and approximate nearest-neighbor search, BM25, TF-IDF, and hybrid approaches. •Experience working with datasets exceeding two million examples and highly imbalanced data, using rigorous evaluation methods for precision and recall trade-offs, threshold selection, and test-set leakage prevention. •A track record of owning production ML systems on AWS, including SageMaker, S3, Athena, Lambda, Glue, Kinesis, and Bedrock, with Terraform, IAM, containers, and Kubernetes-based deployment. •Solid functional working knowledge of model-serving frameworks such as TorchServe, FastAPI, and NVIDIA Triton Inference Server/KServe, and the trade-offs among throughput, GPU efficiency, flexibility, and speed of iteration. •Hands-on experience running CUDA workloads in production, including driver, toolkit, and runtime alignment; GPU passthrough in containers; debugging GPU failures; and improving GPU utilization. •Fluency with AI-native development tools and modern LLM application patterns, including OpenAI-style chat-completion and structured tool/function-calling APIs, MCP, and agent frameworks. •Demonstrated technical leadership as an individual contributor, including setting direction, mentoring engineers across seniority levels, and communicating technical decisions and their business implications to technical and executive audiences. •An understanding of handling sensitive data in accordance with Master Service Agreements and compliance requirements. Compensation: •$172,000 - $258,000 / year •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!