Job type
- Full-time
- Quick Apply
Job description
Job Title: Platform Engineer Sr
Location: Irvine, CA
Hire Type: Full time
JOB DESCRIPTION
Must Have Technical/Functional Skills
- Cloud Platforms: Hands-on administration and support experience across AWS and Microsoft Azure, basic knowledge on compute, storage, networking, identity, access, and monitoring.
- Containers: Practical experience with Kubernetes and Docker, including deployments, services, configuration, logs, health checks, and basic cluster troubleshooting.
- Data Platforms: Working knowledge of Databricks, dbt, Apache Airflow, AutoSys, and enterprise data platform operations. Exposure to LASR and Caspian is preferred.
- Infrastructure as Code and CI/CD: Experience with Terraform and CI/CD tools such as Harness, Azure DevOps, GitHub Actions, Jenkins, or equivalent.
- Scripting: Ability to automate operational tasks using Python, Shell, Bash, or PowerShell.
- Observability: Experience with cloud-native monitoring and log analysis using Azure Monitor, AWS CloudWatch, Datadog, Splunk, or similar tools.
- IT Service Management: Proficiency with ServiceNow and Jira for incident, problem, change, service request, and backlog management.
- Security and Reliability: Understanding of IAM/RBAC, secrets management, vulnerability remediation, patching, platform resilience, and disaster recovery controls.
- Professional Skills: Strong troubleshooting, documentation, collaboration, and written and verbal communication skills in an enterprise environment.
Roles & Responsibilities
- Administer, monitor, and support hybrid cloud infrastructure and data platform services across AWS and Microsoft Azure.
- Deploy, configure, monitor, and troubleshoot containerized workloads running on Kubernetes and Docker.
- Support platform components including Databricks, dbt, Apache Airflow, AutoSys, LASR, and Caspian, including scheduled jobs, dependencies, and integrations.
- Investigate incidents and service degradation across cloud, network, compute, storage, container, orchestration, and data platform layers.
- Manage incidents, service requests, problems, changes, and engineering backlog items through ServiceNow and Jira in accordance with established SLAs and change controls.
- Automate provisioning, configuration, deployment, health checks, and recurring support activities using scripting, Infrastructure as Code, and CI/CD practices.
- Monitor platform availability, performance, capacity, job execution, alerts, and logs; improve dashboards and operational visibility.
- Support upgrades, patching, vulnerability remediation, access controls, backup, disaster recovery readiness, and production release activities.
- Create and maintain runbooks, technical documentation, troubleshooting guides, and knowledge articles.
- Collaborate with application, data engineering, security, network, DevOps, and service management teams to resolve dependencies and improve platform reliability.
- Participate in on-call or production-support rotation as required by the engagement.