Network Operations Center Specialist - Memphis

SpaceXAI - Memphis, TN

Hiring: Network Operations Center Specialist - Memphis Company: SpaceXAI Location: Memphis, TN Job Posted Time: 2026-09-10 11:56:35 Target Skills & Keywords : Linear, Node.js, SOC About the job Experience: •Prior NOC, data center operations, or campus reliability experience in a high-performance computing, AI/ML infrastructure, or large-scale production environment. •Operational familiarity with Linear or similar work-tracking tools for corrective action programs. •Participation in game days, tabletop exercises, or runbook improvement programs. •Prior work in a fast-paced startup or tech company like SpaceXAI. Required Skills: •Staff the console per shift schedule and watch the designated signal surface: cluster health, node availability, network health, facility trend panels, storage alarms, and threshold breaches. •Acknowledge every page within SLA; classify (actionable / known / noise) and log disposition; feed noise patterns back to SRE so signal quality keeps improving. •Detect, verify, and escalate within time budgets; operate the escalation matrix (NOC → on-call SRE → domain owners) and page correctly the first time. •Open and run incident bridges; own stakeholder communications (first update within SLA, then fixed cadence); maintain the incident timeline in real time; call out ownership stalls. •Produce first-pass RCA framing (what happened, when, what’s impacted, who’s engaged) and hand it to SRE / Hardware Failure Analysis for depth — the NOC does not publish root cause. •Run structured shift handoffs and durable shift logs; maintain cross-site awareness. •Write major-incident reports; open corrective projects in Linear and chase them to closure — the NOC is the nag of record. •Maintain and continuously improve NOC runbooks, escalation matrices, and communications templates; participate in SRE-run game days. Qualifications: •Proven ability to acknowledge, classify, and escalate incidents under SLA in a high-signal environment. •Excellent written and verbal communication skills; able to write clear updates while an incident is in progress. •Demonstrated pattern recognition across multiple domains (compute, network, storage, and/or facilities signals) and curiosity about how those systems interact. •Willingness and ability to work a rotating shift schedule, including nights and weekends, as part of continuous campus coverage. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!