Site reliability engineer
Description
A site reliability engineer keeps large-scale systems running reliably, applying software engineering to operations problems. They build monitoring and alerting, automate operational work, manage incidents when systems fail, set and defend reliability targets, and engineer systems to handle scale, failure, and recovery gracefully.
The role originated at Google and blends deep systems knowledge with coding, focused on uptime, performance, and resilience. SREs work at companies running significant infrastructure, often with on-call responsibilities.
The job suits people who think about failure modes and scale, enjoy automating operations rather than doing them by hand, and stay calm and methodical during the high-pressure incidents when critical systems are down.
Dimensions
How this role scores across 12 work traits — creativity, structure, autonomy, and more
Career path
Labor statistics estimates from 2023–2033