Talent.com
prosum
Principal Site Reliability Engineerprosum • Scottsdale, AZ, US
Search for other jobs
No longer accepting applications
Principal Site Reliability Engineer

Principal Site Reliability Engineer

prosum • Scottsdale, AZ, US
1 day ago
Job type
  • Full-time
  • Quick Apply
Job description

Principal Site Reliability Engineer Position Overview
We are seeking an experienced Principal Site Reliability Engineer (SRE) to provide technical leadership across highly available, large-scale production environments. This role combines software engineering, systems engineering, cloud infrastructure, automation, DevOps, and production reliability to improve the resilience, scalability, performance, observability, and operational health of critical services.

The Principal SRE will partner closely with Software Engineering, Platform Engineering, Cloud Infrastructure, DevOps, and other technology teams to ensure reliability and operational readiness are incorporated throughout the software development lifecycle.

This is a senior individual contributor position with enterprise-level influence. The successful candidate will identify systemic reliability risks, establish technical direction, influence architecture and engineering practices, and help improve reliability capabilities across multiple engineering teams.
Key Responsibilities


  • Apply Site Reliability Engineering (SRE), software engineering, automation, and DevOps principles to improve how production services are built, tested, deployed, monitored, operated, and recovered.

  • Establish and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, availability metrics, and service-health measurements.

  • Design and enhance observability capabilities using metrics, logging, distributed tracing, monitoring, alerting, dashboards, and service-health instrumentation.

  • Drive continuous improvement across CI/CD pipelines, Infrastructure as Code (IaC), cloud infrastructure, deployment practices, automation, testing, incident management, capacity planning, resilience, disaster recovery, and operational readiness.

  • Analyze production environments to identify systemic reliability risks, performance bottlenecks, recurring incidents, and opportunities for automation.

  • Translate production and operational experience into improvements in application code, architecture, infrastructure, tooling, automation, and engineering standards.

  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.

  • Lead or participate in production incident response, troubleshooting, root cause analysis, service restoration, and blameless post-incident reviews.

  • Provide technical leadership during critical production incidents and help improve incident response, escalation procedures, service restoration, and sustainable on-call practices.

  • Reduce operational toil and manual intervention through software development, scripting, automation, reusable tooling, platforms, and engineering patterns.

  • Apply data-driven analysis, experimentation, and engineering principles to validate assumptions and guide technical decisions.

  • Establish and influence enterprise-level SRE, DevOps, cloud, reliability, and operational engineering standards and best practices.

  • Mentor engineers and technical leaders while promoting knowledge sharing and sustainable engineering capabilities across the organization.

  • Operate independently across complex, business-critical reliability and infrastructure challenges.

Required Qualifications


  • 15+ years of relevant professional experience in one or more of the following areas:

    • Site Reliability Engineering (SRE)

    • Software Engineering

    • Systems Engineering

    • Cloud Engineering

    • Platform Engineering

    • DevOps Engineering

    • Infrastructure Engineering

    • Systems Architecture

  • Strong experience with software development and/or scripting using one or more modern programming languages.

  • Advanced understanding of software engineering principles, distributed systems, production environments, troubleshooting, automation, and observability.

  • Experience designing, supporting, or improving highly available, scalable production systems and distributed applications.

  • Experience with public cloud platforms and cloud-native architectures, preferably AWS.

  • Strong knowledge of Linux/Unix systems, networking, infrastructure, application architecture, and production operations.

  • Demonstrated experience diagnosing complex production issues and implementing sustainable technical solutions.

  • Strong analytical, troubleshooting, problem-solving, communication, and cross-functional collaboration skills.

  • Ability to provide technical direction and influence engineering practices across multiple teams and organizational boundaries.

Preferred Qualifications


  • Extensive hands-on experience with Amazon Web Services (AWS) or another major cloud platform, including Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).

  • Experience with CI/CD pipelines and software delivery automation.

  • Experience with Infrastructure as Code (IaC) technologies and practices.

  • Experience with containers and container orchestration technologies.

  • Strong experience with monitoring, logging, distributed tracing, dashboards, alerting, and observability platforms.

  • Experience defining and managing SLIs, SLOs, error budgets, availability targets, and reliability metrics.

  • Experience with incident management, root cause analysis, performance engineering, capacity planning, resilience testing, disaster recovery, and operational readiness.

  • Experience creating reusable automation, tooling, platforms, frameworks, engineering patterns, or standards that improve engineering productivity and system reliability.

  • Experience influencing architecture and technical strategy for large-scale or business-critical production systems.

  • Demonstrated ability to mentor senior engineers and improve technical capabilities across engineering organizations.

  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical discipline, or equivalent practical experience.

Key Technical Skills / ATS Keywords
Site Reliability Engineering (SRE), AWS, Cloud Computing, DevOps, Software Engineering, Systems Engineering, Platform Engineering, Distributed Systems, Production Reliability, High Availability, Scalability, Resilience, Observability, Infrastructure as Code (IaC), CI/CD, Automation, Linux, Unix, Networking, Containers, Container Orchestration, Monitoring, Logging, Distributed Tracing, Alerting, SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis, Production Support, Performance Engineering, Capacity Planning, Disaster Recovery, Resilience Testing, Operational Readiness, Cloud Architecture, Production Operations, Software Development, Scripting, Troubleshooting, Technical Leadership.

Create a job alert for this search

Principal Site Reliability Engineer • Scottsdale, AZ, US