Senior Lab Reliability Engineer

VAST Data - New York, United States

Hiring: Senior Lab Reliability Engineer Company: VAST Data Location: New York, United States Job Posted Time: 2026-09-02 18:01:31 Target Skills & Keywords : Ansible, Bash, Docker, Infrastructure as Code, Kubernetes, Linux, Python, REST, Subnetting About the job Experience: •4+ years of professional experience in a systems engineering, storage engineering, customer support engineering, or related role Required Skills: •Own the operational reliability of VAST clusters in the lab environment, including proactive health monitoring, upgrade planning, and issue resolution •Serve as the primary technical escalation point for complex cluster issues, working hands-on-keyboard to resolve them and partnering with VAST engineering when deeper investigation is needed •Shape the automation and tooling strategy for the lab environment, including provisioning scripts, CLI utilities, monitoring dashboards, and internal tooling that the rest of the team builds on •Help establish standards for infrastructure-as-code and configuration management (Ansible or equivalent) across the lab environment •Reproduce and isolate difficult issues in controlled lab environments, producing high-quality diagnostic data and reports for engineering •Own broader systems infrastructure supporting the lab, including virtualization platforms (VMware vSphere, Proxmox), compute, networking, and storage •Mentor and support other members of the lab operations team, raising the collective technical bar •Partner with pre-sales SEs, professional services, and engineering to reproduce customer-relevant scenarios and validate solutions in the lab Qualifications: •4+ years of professional experience in a systems engineering, storage engineering, customer support engineering, or related role •Deep hands-on experience with enterprise storage systems (VAST, Pure, NetApp, Isilon, Ceph, or similar) at an operational or reliability level •Strong Linux systems administration skills, including networking, storage, filesystems, systemd, and CLI tooling •Strong scripting/programming experience (Python, Bash, or similar), sufficient to build and maintain automation and tooling other engineers rely on •Experience with Docker and Kubernetes in operational environments •Experience with infrastructure-as-code or configuration management tools (Ansible or equivalent) •Solid networking foundations (VLANs, routing, subnetting, network troubleshooting) •Experience with virtualization platforms (VMware vSphere, ESXi, Proxmox, or equivalent) •Methodical approach to troubleshooting complex issues across storage, networking, and compute layers •Comfort working across time zones with distributed team members •Excellent written and verbal communication, including ability to produce clear technical documentation and reports •Preferred Qualifications •Existing hands-on experience with VAST Data clusters •Prior experience as a Customer Support Engineer at a storage or infrastructure company OR Reliability Engineer / SRE supp Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!