Site Reliability Engineer - Video Platform - USDS
TikTok USDS Joint Venture - San Jose, CA
Hiring: Site Reliability Engineer - Video Platform - USDS Company: TikTok USDS Joint Venture Location: San Jose, CA Job Posted Time: 2026-09-11 13:25:01 Employment Type: Hybrid Target Skills & Keywords : AWS, C, C#, C++, Change Management, ELK Stack, Java, Linux, Microservices, MongoDB, MySQL, Python, Redis About the job Required Skills: •TikTok video system is a world-leading video platform that provides multimedia storage, delivery, transcoding services. As part of the USDS, the Video Platform team is responsible for building the next generation video processing platform which provides excellent experiences for billions of users around the world. •The USDS Video Platform team is seeking an experienced Site Reliability Engineer to help us continue improving TikTok's video system. If you are passionate about ensuring software reliability, love problem-solving, and are prepared for exciting challenges, we would like you on our team. •Lead and oversee overall reliability of TikTok's video system, including video publishing and distribution. •Perform lifecycle management of production systems including change management, service deployment, operations and emergency response. •Monitor the system and respond to incidents to maintain system service level agreement (SLA), review and follow up all production incidents. •Perform capacity management of compute, storage and network bandwidth resources to ensure system stability and save infrastructure costs. •Provide strong support during big events to ensure the system is capable of consuming a large volume of Internet traffic. •Build tools, automations, visualizations and monitors to facilitate the operation and optimization of the global infrastructure. Qualifications: •Bachelor's degree in Computer Science or a related technical background involving software/system engineering, or equivalent working experience. •Programming experience with at least one of the following languages: C, C++, Java, Python, C# or Go. •Extensive knowledge of networking, operation system, database system and container technology. •Good understanding of every aspect of microservice architecture, and hands on experience in troubleshooting in large scale distributed systems. •Applied hands-on capability in common opensource systems such as Linux, MySQL, MongoDB, Redis and ELK. •Passionate, self-motivated and good teamwork skills. •On-site presence across teams allows the company to operate with greater speed, alignment, and agility — especially in areas like real-time decision-making, team development, and integrated execution. As such, the company is shifting from a hybrid work model to a fully in-person schedule up to 5 days a week. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!