HPC Systems Engineer

University of California, San Francisco - San Francisco, CA

Hiring: HPC Systems Engineer Company: University of California, San Francisco Location: San Francisco, CA Job Posted Time: 2026-09-11 14:20:36 Target Skills & Keywords : HIPAA, Systems Design About the job Experience: •6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience with large-scale or HPC systems •5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband based clusters Required Skills: •The CoreHPC team at UCSF is seeking an HPC Systems Engineer to play a key role in the development, maintenance, and day-to-day operations of the Institute’s HPC clusters. •The HPC Systems Engineer Will •Apply their engineering and design skills to develop new CI solutions, to develop and enhance monitoring to maintain the integrity of CI systems. •Select methods, techniques and evaluation criteria to develop new CI solutions to address complex research problems. •Be an active member of the support and maintenance efforts for the CoreHPC cluster, resolving user issues, fixing technical problems, resolving outages, patching, and maintaining systems' uptime and availability. •Provides consultation, support, and guidance to researchers on how to address computational problems using standard tools, packages, and approaches. •Develop enhancements of monitoring to maintain the integrity of CI systems. •Participate in multiple technical projects simultaneously. Qualifications: •Bachelor's degree in a related area such as computer science or engineering, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience with large-scale or HPC systems •Expert knowledge of HPC systems infrastructure design •Strong knowledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc. •Advanced knowledge of computer security best practices and policies including demonstrated experience securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requirements •Demonstrated testing and test planning skills. Demonstrated ability to create automated testing. •Knowledge of HPC job scheduler system design and operation such as SLURM or PBS, Demonstrated skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband based clusters •Demonstrated capacity to elicit and communicate technical and non-technical information in a clear and concise manner. •Self-motivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn effectively and meet deadlines. •Understanding of system performance monitoring and actions that can be taken to improve or correct performance. •Demonstrated advanced knowledge, skills and abilities associated with system problem identification and resolution. Experience with design, configuration, operation, repair, and tuning of technology systems. Compensation: •In addition to our PRIDE values, UCSF is committed to equity – both in how we deliver care as well as our workforce Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!