Job Title
Data Pipeline Developer
Job Description
The following are the duties of this position at the full working level. If this vacancy includes more than one grade and you are selected at a lower grade level, you will have the opportunity to learn to perform these duties and receive training to help you grow in this position.
Data Pipeline Development
Designs, develops, and maintains high-performance streaming, batch, and realtime data pipelines using PySpark and Delta Live Tables on Databricks. Integrates structured, semi-structured, and unstructured data sources into the enterprise data lake.
Process Optimization
Identifies, designs, and implements process improvements, including infrastructure redesign for scalability, data delivery optimization, and automation of manual workflows.
Cloud Integration
Integrates Databricks with cloud services (AWS, Azure, GCP) for storage, compute, and security. Utilizes technologies such as Python, SQL, Spark, and AWS S3 to build infrastructure for efficient data extraction, transformation, and loading.
ETL/ELT Workflows
Develops and manages ETL/ELT workflows to ensure data quality, reliability, and adherence to Federal guidelines. Contributes to transitioning legacy data systems into modern cloud-native platforms like AWS and Azure.
Analytical Tool Development
Creates analytical tools to leverage data pipelines, delivering actionable insights into business performance metrics such as operational efficiency and customer acquisition.