Job Description
Site Reliability Engineer (SRE) for Google Cloud Platform (GCP), focused on building, maintaining, and improving the reliability, scalability, and performance of cloud infrastructure using Infrastructure as Code (IaC) and Terraform Enterprise. Supports the delivery of secure, compliant, and highly available cloud environments aligned with enterprise standards and regulatory requirements.
Works closely with engineering and platform teams to develop and maintain reusable IaC modules, Terraform configurations, and automated cloud services, enabling consistent and efficient infrastructure provisioning. Contributes to the implementation of standardized platform patterns, including networking, identity, logging, and monitoring capabilities.
Participates in the end-to-end lifecycle of cloud infrastructure, including deployment, monitoring, incident response, and continuous improvement. Helps implement and maintain CI/CD pipelines, policy-as-code frameworks, and automation solutions to ensure reliable and repeatable deployments.
Applies SRE principles and practices, including monitoring, alerting, incident management, and root cause analysis, to improve system reliability and reduce operational risk. Supports the definition and tracking of service performance through metrics such as availability and latency.
Collaborates with architecture, security, and engineering teams to ensure infrastructure is secure, compliant, and operationally resilient. Contributes to DevSecOps practices by integrating security and compliance controls into automated workflows.
Continuously identifies opportunities to improve system reliability, reduce manual effort, and enhance automation. Leverages emerging tools and technologies, including AI/ML where applicable, to support proactive operations, observability, and platform stability.
Key Responsibilities
Design, develop, and maintain Google Cloud Platform (GCP) infrastructure using Infrastructure as Code (IaC) with Terraform Enterprise
Contribute to the implementation of scalable, secure, and compliant cloud solutions aligned with enterprise standards
Develop and maintain reusable Terraform modules and standardized infrastructure patterns to enable consistent and automated provisioning of GCP resources
Follow and contribute to code quality standards, design patterns, and peer review practices to ensure reliable and maintainable infrastructure code
Support the adoption and use of Terraform Enterprise for automated provisioning, policy enforcement, and infrastructure governance
Implement and maintain cloud automation workflows, including provisioning, configuration management, and environment setup
Build and enhance CI/CD pipelines for infrastructure delivery, ensuring automated testing, validation, and compliance checks
Implement policy-as-code and security controls, ensuring infrastructure meets regulatory and enterprise compliance requirements
Participate in the end-to-end lifecycle of infrastructure delivery, including deployment, monitoring, and continuous improvement
Collaborate with architecture, security, and engineering teams to ensure secure, resilient, and compliant cloud configurations
Apply DevSecOps and cloud-native practices to improve automation, security, and deployment efficiency
Contribute to observability, logging, and monitoring solutions to support proactive incident detection and response
Execute testing and validation of IaC modules, including integration and deployment verification
Identify opportunities to automate manual processes and improve operational efficiency
Support reliability, scalability, and performance of cloud platforms through automation and standardization
Troubleshoot and resolve infrastructure and platform issues, contributing to root cause analysis and continuous improvement
Work with stakeholders to implement infrastructure solutions that meet technical and business requirements
Evaluate and adopt emerging tools and technologies to enhance automation, reliability, and platform performance
Conduct performance testing and capacity planning to ensure systems scale reliably under load.
Optimize system performance, latency, and resource utilization across cloud environments
Design and implement observability solutions, including metrics, logs, traces, and alerting strategies.
Reduce alert fatigue and improve signal quality through meaningful alert design and tuning.
Develop dashboards and monitoring frameworks aligned to SLOs.
Required Qualifications
4-8+ years of experience in cloud infrastructure engineering, platform engineering, or cloud operations, with exposure to Google Cloud Platform (GCP)
Strong hands-on experience with Infrastructure as Code (IaC), including practical use of Terraform or Terraform Enterprise for infrastructure provisioning
Solid understanding of software engineering fundamentals, including version control, code quality, and basic testing practices for infrastructure code
Experience developing and maintaining Terraform modules and infrastructure configurations to support automated cloud environments
Familiarity with CI/CD pipelines for infrastructure deployment, including automated build, test, and release processes
Working knowledge of DevSecOps practices, including integrating security and compliance checks into automated workflows
Good understanding of GCP services and cloud architecture fundamentals, including networking (VPCs, IAM, load balancing)
Exposure to policy-as-code, governance, and compliance requirements in enterprise environments
Experience supporting automation and standardization efforts to improve consistency and efficiency in cloud deployments
Understanding of monitoring, logging, and observability tools to support system reliability and performance
Hands-on experience with incident response, troubleshooting, and root cause analysis in cloud or distributed systems
Ability to collaborate effectively with engineering, architecture, and security teams to support reliable and secure platform operations
Strong problem-solving and analytical skills, with the ability to diagnose and resolve infrastructure issues
Effective communication skills, with the ability to work within cross-functional teams and document technical solutions clearly
Interest in learning and applying emerging technologies and automation techniques (including AI/ML where applicable) to improve platform reliability