Mercor connects exceptional technical talent with leading organisations working on ambitious technology and AI initiatives. We are looking for experienced Infrastructure / Site Reliability Engineers (SREs) to join a full-time engagement focused on building and operating complex, enterprise-grade infrastructure. We are seeking engineers with strong hands-on experience building, operating, debugging, and scaling sophisticated production systems. The ideal candidate has worked extensively with Kubernetes, AWS, observability platforms such as Datadog, and modern infrastructure tooling . This is a full-time opportunity , and candidates must be able to commit to full-time engagement. ## What You'll Do - Build, operate, and improve highly available and scalable production infrastructure. - Manage and optimise Kubernetes-based production environments . - Design and maintain cloud infrastructure, primarily across AWS . - Improve system reliability, availability, scalability, and operational efficiency. - Build and maintain observability across infrastructure and applications using Datadog or similar platforms. - Investigate production incidents, perform root-cause analysis, and implement durable fixes. - Improve monitoring, alerting, logging, tracing, and overall production visibility. - Develop automation and internal tooling to reduce manual operational work. - Partner closely with software engineering teams on deployments, infrastructure, and production reliability. - Contribute to infrastructure architecture and technical decisions for complex distributed systems. ## Ideal Background - Professional experience in Infrastructure Engineering, Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Production Engineering . - Hands-on experience operating complex, enterprise-grade production systems . - Strong production experience with Kubernetes . - Strong experience with AWS and cloud-native infrastructure. - Experience with Datadog , Prometheus, Grafana, or comparable observability platforms. - Experience with Infrastructure as Code using Terraform, Pulumi, or equivalent technologies . - Strong understanding of distributed systems, networking, containers, Linux, and cloud architecture. - Experience building or maintaining CI / CD and production deployment infrastructure. - Strong debugging, troubleshooting, and incident-response capabilities. - Proficiency in at least one programming or scripting language, such as Python, Go, or Bash . ## Strong Signals - Experience operating Kubernetes and cloud infrastructure at significant production scale. - Experience supporting high-traffic or mission-critical applications. - Experience building infrastructure or platform tooling used by large engineering organisations. - Ownership of production reliability, on-call operations, incident response, or capacity planning. - Experience working within sophisticated, large-scale distributed systems. - Demonstrated improvements to SLOs / SLIs, observability, deployment reliability, infrastructure performance, or operational efficiency . ### Why Join - Solve challenging reliability, scalability, and performance problems across enterprise-grade production systems. - Work extensively with technologies such as Kubernetes, AWS, Datadog, Terraform / Pulumi , and modern cloud-native tooling. - Take meaningful ownership of production reliability, observability, infrastructure architecture, and operational improvements. - Competitive hourly compensation reflecting your experience and technical expertise. - Join a network of highly skilled engineers working on ambitious projects with leading technology and AI organisations.
Create a job alert for this search
Remote Infrastructure / Site Reliability Engineer (SRE) - AI Trainer ($200-$200 per hour) • Glendale, California, US