Site Reliability Engineer

Schonfeld - New York, NY

Hiring: Site Reliability Engineer Company: Schonfeld Location: New York, NY Job Posted Time: 2026-09-17 02:55:33 Target Skills & Keywords : AWS, CI/CD, Datadog, DynamoDB, Elasticsearch, Embedded Systems, Event-Driven, Fixed Income, GitHub Actions, Kubernetes, LLM, MySQL, OpenSearch, PostgreSQL, Python, RAG, REST, S3 About the job Experience: •5+ years of professional cloud automation or site reliability engineering experience, or similar Required Skills: •Set the reliability standards for Enterprise AI, defining Service Level Objectives (SLOs), error budgets, and custom incident response runbooks. •Own the observability, incident response, reliability, and scalability of our AI platform. •Ensure that our agents, gateways, LLM proxies, and RAG pipelines operate with high availability, accuracy, and financial efficiency. •Support users in a dedicated help channel when issues arise — investigating root causes, providing solutions, monitoring the status of upstream dependencies, and communicating updates back to users. •You’ll also be involved in the team’s code quality and best practices, identify gaps in development lifecycles, and continuously improve both the core platform and how developers build with it. •5+ years of professional cloud automation or site reliability engineering experience, or similar •Operational familiarity with software development best practices and experience creating applications in Python •A passion for AI, creative ideas for how it can be a tool for problem solving and productivity Compensation: •Flexible work environment (work from home / hybrid options) •Competitive benefits and rewards package Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!