Senior Systems Reliability Engineer II

ThoughtSpot - Mountain View, CA

Hiring: Senior Systems Reliability Engineer II Company: ThoughtSpot Location: Mountain View, CA Job Posted Time: 2026-09-10 13:05:10 Target Skills & Keywords : AI, AWS, Azure, Bash, C, C++, Datadog, GCP, Grafana, Java, LLM, Linux, Prometheus, Python, SaaS, Splunk About the job Required Skills: •Technical & Customer Support •Act as the primary point of contact for customer-facing technical issues related to our SaaS platform, including data connectivity, report errors, performance concerns, access problems, data inconsistencies, software bugs, and integration challenges. •Understand and empathize with the challenges ThoughtSpot users face, offering tailored solutions to improve their experience. •Provide timely, accurate, and clear updates to customers, consistently meeting SLAs and driving issues through to full resolution via tickets and calls. •Translate complex technical issues into clear, concise updates for both technical and non-technical stakeholders. •Create and maintain knowledge-base articles to empower customer self-service and improve support efficiency. •Maintain, monitor, and troubleshoot ThoughtSpot cloud infrastructure using tools like Grafana, Prometheus, Datadog, and Splunk. •Monitor system health and performance through metrics, logs, and dashboards to detect and prevent issues proactively. Qualifications: •B.S. in Computer Science or equivalent relevant experience. •Proven experience troubleshooting complex Linux systems and managing virtualization and cloud platforms (VMware, AWS, Azure, GCP). •Applied hands-on capability in monitoring tools such as Grafana, Prometheus, Datadog, or Splunk. •Demonstrated experience and a keen interest in leveraging AI/ML principles to address SRE challenges — including AIOps, predictive maintenance, and intelligent automation. •Prior experience in enterprise customer support, including on-call rotations and incident management, with the ability to lead root cause analyses. •Strong problem-solving and algorithmic thinking with a solid understanding of system internals. •Excellent verbal and written communication skills with the ability to work independently and cross-functionally in fast-paced environments. •Operational familiarity with scripting and programming languages such as Python, Go, Bash, or Java. •Exposure to infrastructure and service monitoring frameworks with the ability to analyze data to ensure high availability. •Operational familiarity with C/C++ or other low-level systems languages. Compensation: •Flexible work environment (work from home / hybrid options) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!