HPC Systems Engineer

NorthMark Strategies - Dallas-Fort Worth Metroplex

Hiring: HPC Systems Engineer Company: NorthMark Strategies Location: Dallas-Fort Worth Metroplex Job Posted Time: 2026-09-17 01:14:28 Target Skills & Keywords : Ansible, CI/CD, Kubernetes, Systems Engineering, Terraform About the job Experience: •8+ years of experience in systems engineering, product engineering, HPC, cloud infrastructure, distributed systems, or data center infrastructure. Required Skills: •Drive cross-domain technical alignment across the compute, storage, networking, Kubernetes, automation, and data center infrastructure teams. •Lead integrated design reviews for new HPC deployments, platform expansions, and major infrastructure initiatives. •Identify technical dependencies, design gaps, scalability constraints, and integration risks throughout the delivery lifecycle. •Partner with domain engineering teams to define end-to-end architecture, engineering standards, and implementation readiness criteria. •Ensure infrastructure designs meet requirements for performance, resiliency, scalability, and long-term operational supportability. •Develop system-level engineering specifications that define how infrastructure components integrate and interoperate. •Participate in failure analysis, resiliency reviews, and operational readiness assessments ahead of production turnover. •Partner cross-functionally with automation teams to improve deployment consistency, validation, lifecycle management, and operational efficiency. Qualifications: •Bachelor's degree in Engineering, Computer Science, or equivalent experience. •Deep expertise in at least one infrastructure domain — compute, networking, storage, Kubernetes, or data center infrastructure — with broad working knowledge across adjacent domains. •Working familiarity with the components of a modern AI/HPC platform, such as GPU compute (NVIDIA H100/H200 class), high-speed fabrics (InfiniBand NDR/HDR, RoCEv2), scale-out storage (VAST Data, WekaFS, Lustre, or GPFS), and Kubernetes or Slurm-based scheduling. •In-depth knowledge of system-level dependencies, architectural tradeoffs, and downstream operational impacts. •Proven ability to lead technical discussions and drive alignment across multiple engineering teams without direct authority. •Excellent communication skills, with the ability to present complex technical concepts to both engineering and leadership audiences. •Demonstrated capacity to thrive in a fast-paced environment with evolving requirements and aggressive delivery timelines. •Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future. Compensation: •Company-Paid Lunch Stipend: Lunch is provided via GrubHub Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!