Software Development Engineer, Infrastructure Reliability Engineering

Amazon - Arlington, VA

Hiring: Software Development Engineer, Infrastructure Reliability Engineering Company: Amazon Location: Arlington, VA Job Posted Time: 2026-09-17 04:57:52 Employment Type: Full-time Target Skills & Keywords : AWS, C#, C++, Embedded Systems, Event-Driven, Java, LLM, Machine Learning, Perl, Serverless About the job Experience: •3+ years of non-internship professional software development experience •2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience •1+ years of software development engineer or related occupational experience •1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience •1+ years of Object Oriented Design experience Required Skills: •Design, build, test, deploy, and operate production services and AI agents on AWS (Amazon Bedrock AgentCore, serverless compute, event-driven pipelines) that automate incident triage, call intelligence, communications, and post-incident documentation and reporting •Own features end-to-end: from discovery with Incident Managers and resolver teams, through design, implementation, evaluation, deployment, and production operation •Build the foundations that gate agent autonomy: LLM output evaluation, observability and alerting for agents in production, and identity and access controls aligned with Amazon standards •Design the feedback loops through which agents learn: capturing human reviews, corrections, approvals, and incident outcomes as evaluation signal, and turning resolved incidents into structured history that improves recommendations over time •Integrate with the incident lifecycle (ticketing, chat, telemetry, detection feeds, and live call transcription) and model consistent incident state across those systems •Raise the bar on software quality, security, testing, and operational excellence for AI systems acting inside production incident workflows •Medical, Dental, and Vision Coverage •Maternity and Parental Leave Options Qualifications: •Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field •Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques Compensation: •Medical, Dental, and Vision Coverage Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!