Job Title: Data Engineer
Job Location: Glendale, CA
Job Type: Full-Time
Job Description:
Databricks experience (primary requirement)
Apache Airflow
Advanced SQL skills
Python
Spark / PySpark
Scala
Experience building and maintaining data pipelines and workflows
Interview Process
Expected to follow the team's recent hiring process for similar data engineering positions.
Typically consists of:
Initial screening/vetting round
Technical interview
Final interview (potentially onsite)
Total process generally spans 2-3 interview rounds.
Key Responsibilities:
Design, write, test, and deploy data pipelines using PySpark, Scala, SQL, Python
Meet with stakeholders to gather requirements and translate them into scalable data platform solutions
Understanding of Databricks platform and developer tooling to diagnose errors, audit platform activity, and automate updates across pipelines, objects, and integrations
Ability to explain Spark architecture and pipeline behavior to stakeholders to diagnose root causes and recommend solutions
Provide solution architecture across AWS, Databricks, Kubernetes, and Airflow (MWAA), including cross-platform integrations
Manage Databricks platform governance, including Unity Catalog, ACLs, lineage, and data discovery and privacy tooling
Build and maintain Kubernetes containers and containerized utilities supporting deployed data platform services
Apply networking knowledge to troubleshoot connectivity and integration errors across platform components
Perform platform administration: provision and remove access, assess resource utilization, monitor platform health and cost, and evaluate stakeholder requests
Collaborate with engineers, architects, and product managers to drive Core Data platform success; participate in agile/scrum ceremonies
Maintain documentation of platform changes, standards, and pipeline configurations to support data quality and governance
Qualifications:
5+ years of data engineering experience developing and operating large-scale data pipelines
Deep hands-on experience with Databricks and Apache Spark (batch and streaming), including pipeline development in PySpark and/or Scala
Strong understanding of Spark architecture-executors, stages, partitioning, shuffle, and performance tuning-with ability to explain tradeoffs to technical and non-technical stakeholders
Proficiency with Databricks platform tooling (API, SDK, CLI) for automation, auditing, governance, and operational troubleshooting
Proficient in SQL with advanced performance tuning capabilities
Hands-on production experience with Airflow (MWAA) for orchestrating data pipelines
Experience managing Databricks platform governance: ACLs, Unity Catalog, lineage, and access provisioning
Proficiency in Python and at least one additional language (Scala, Kotlin, or SQL-driven pipeline tooling)
Experience designing and optimizing scalable ETL/ELT pipelines integrating diverse structured and unstructured data sources
AWS-primary experience (compute, storage, networking, IAM); experience with other cloud providers is transferable
Proficiency with Docker and Kubernetes for building and maintaining containerized data platform services
Working knowledge of networking concepts to diagnose cross-platform integration and connectivity issues
Familiarity with Snowflake and comparable tooling relative to the Databricks ecosystem
Experience designing and implementing CI/CD and DevOps practices (Git-based workflows)
Experience implementing data quality checks, monitoring, and logging for pipeline reliability
Self-starting problem solver with strong analytical and communication skills; willingness to learn new tooling and trends
Familiar with Scrum and Agile methodologies
Experience with Snowflake is a plus
Bachelor's Degree in Computer Science, Information Systems, or a related field, or equivalent work experience, Master's Degree is a plus