Platform Engineer / SRE (Site Reliability Engineer) Cloud & AI Automation
Location: Colorado (Mon to Thu wfo)
Role Overview:
- We are seeking a highly skilled and proactive **Senior Platform Engineer / SRE** to lead the design, automation, and operational excellence of cloud platforms and AI-driven systems.
- The ideal candidate will be a seasoned engineer with hands-on experience in **AWS, Terraform, Kubernetes, CI/CD, Infrastructure as Code (IaC), and Observability**, who thrives in fast-paced, high-impact environments.
- You will play a pivotal role in building and maintaining scalable, secure, and self-healing platforms that support enterprise-grade applications and AI-powered workflows.
- This role sits at the intersection of **platform engineering, reliability, security, and AI automation**, where you'll drive innovation through agentic AI, infrastructure automation, and cloud cost optimization.
Key Responsibilities:
- Design, implement, and manage **cloud-native platforms** on AWS using **Terraform, CloudFormation, and IaC best practices**.
- Lead **infrastructure automation** initiatives to enable zero-touch provisioning, self-service environments, and CI/CD pipeline integration.
- Architect and operate **highly available, scalable, and secure systems** using **Kubernetes (EKS/ECS), Docker, and microservices**.
- Implement and maintain **robust observability stacks** using **Grafana, Datadog, Prometheus, ELK, and OpenTelemetry** for real-time monitoring, alerting, and performance tuning.
- Own **incident management lifecycle**: lead on-call rotations, conduct **RCA (Root Cause Analysis)**, and implement preventive measures to improve system reliability.
- Drive **cloud cost optimization** strategies across multi-cloud environments (AWS, Azure, GCP) through right-sizing, auto-scaling, tagging policies, and usage analytics.
- Enforce **security compliance standards** including **PCI, PII, HIPAA, GDPR, ISO 27001/27701**, and SOC 2, ensuring IAM policies, encryption, and audit readiness.
- Integrate **Okta, AWS IAM, and identity federation** for secure access control and role-based access management.
- Spearhead **AI automation and agentic AI** use cases for infrastructure operations-automating deployments, incident response, policy enforcement, and resource provisioning.
- Collaborate with DevOps, SRE, Security, and Product teams to deliver **platform capabilities** that accelerate time-to-market and improve developer experience.
- Mentor junior engineers and contribute to **engineering best practices, documentation, and knowledge sharing**.
Required Qualifications:
- Bachelor's or Master's degree in Computer Science, Engineering, or related field.
- 6+ years of hands-on experience in **DevOps, SRE, or Platform Engineering** roles.
- Expertise in **AWS cloud services** (EC2, S3, Lambda, RDS, VPC, IAM, CloudFront, etc.) and **Terraform/CloudFormation** for infrastructure provisioning.
- Proven experience with **Kubernetes (EKS, AKS, GKE)**, **Docker**, and container orchestration.
- Strong proficiency in **CI/CD pipelines** using Jenkins, GitHub Actions, GitLab CI, or similar tools.
- Deep understanding of **observability tools**: Grafana, Datadog, Prometheus, Kibana, and logging frameworks.
- Experience with **incident management, RCA, and post-mortem processes** in production environments.
- Solid knowledge of **security compliance frameworks**: PCI-DSS, HIPAA, GDPR, CCPA, ISO 27001/27701.
- Experience integrating **identity providers (Okta, AWS IAM)** and managing RBAC across cloud environments.
- Hands-on experience with **Python** for automation, scripting, and tooling.
- Familiarity with **Agentic AI, AI automation, and intelligent operations (AIOps)** for infrastructure and platform management.
- Strong communication, collaboration, and leadership skills.
Preferred Qualifications:
- Experience with **multi-cloud environments** (AWS, Azure, GCP).
- Knowledge of **serverless architectures** (AWS Lambda, Azure Functions).
- Experience with **data platforms** (Snowflake, Redshift, BigQuery) and ETL pipelines (Airflow, Glue).
- Exposure to **AI/ML platforms** and **GenAI integration** in DevOps workflows.
- Certification in AWS, Google Cloud, or Kubernetes (e.g., AWS Certified DevOps Engineer, CKAD, CKA).
Donato Technologies Inc. is a trusted IT staffing, consulting, and software development partner headquartered in Dallas, Texas. We support clients across industries by understanding their unique business needs and delivering tailored technology and workforce solutions. Our focus is on connecting the right talent with the right opportunity-ensuring clients receive dependable, skilled professionals and candidates receive meaningful career growth and support. We work closely with small to mid-sized organizations to provide flexible, high-quality services that drive performance, innovation, and long-term success.