Job type
- Full-time
Job description
MalaceHR is seeking an experienced Data Engineer to join a fast-paced technology and enterprise data environment. This position will be responsible for designing, developing, and optimizing scalable data solutions using modern cloud and big data technologies.
The ideal candidate is a self-starter with strong hands-on experience in Azure Databricks, Python, PySpark, Spark SQL, Azure Functions, Delta Lake, Azure DevOps CI/CD, and agent-driven workflows leveraging MCP frameworks. This individual will work closely with cross-functional technical and business teams to develop data pipelines, modernize data platforms, and support enterprise-level data initiatives.
Key Responsibilities
- Design, architect, and develop scalable data solutions using cloud-based big data technologies.
- Ingest, process, transform, and analyze large and diverse datasets to support business and technical requirements.
- Design and develop data management and persistence solutions using relational and non-relational databases.
- Build scalable data lake solutions capable of storing structured, semi-structured, and unstructured data from internal and external sources.
- Develop proofs of concept (POCs) to validate new technologies, architectures, and solution proposals.
- Provide technical guidance supporting migrations to modern cloud-based data platforms.
- Develop, maintain, and optimize ETL and ELT workflows using Azure Databricks, Python, PySpark, and Spark SQL.
- Extract, manipulate, and transform data from databases, data lakes, APIs, files, streaming platforms, and other data sources.
- Develop efficient data-processing pipelines with appropriate validation, error handling, monitoring, and performance tuning.
- Build systems that ingest, cleanse, normalize, and structure large and diverse datasets.
- Develop event-driven and streaming data pipelines supporting real-time and near-real-time processing.
- Contribute to and follow CI/CD processes and Data Engineering development best practices.
- Perform unit testing, system integration testing, regression testing, and support user acceptance testing.
- Translate business requirements into scalable technical solutions that can be designed and engineered.
- Partner with business stakeholders to develop technical documentation and communication materials supporting proper data usage and interpretation.
- Implement data security best practices, including encryption, access controls, data governance, and compliance requirements.
- Maintain data privacy, confidentiality, quality, and integrity throughout the data engineering lifecycle.
- Perform data analysis and troubleshooting to identify and resolve data-related issues.
- Build agent-driven workflows used to load, process, and transform data.
- Develop data solutions leveraging MCP frameworks and agent orchestration.
- Collaborate with software engineers, architects, product teams, analysts, and other cross-functional stakeholders.
Qualifications
- Minimum of 5 years of professional Data Engineering or Data Development experience.
- Strong hands-on experience with:
- Python
- PySpark
- Spark SQL
- ETL/ELT development
- SQL Server
- Data pipeline development
- Bachelor’s degree in Computer Science, Information Science, Mathematics, Statistics, Engineering, Data Science, or another related quantitative discipline preferred.
- Experience working with Microsoft Azure cloud technologies.
- Strong experience with Azure Databricks and Azure Storage/Data Lake technologies.
- Strong understanding of data engineering architecture, data modeling, data integration, and ETL concepts.
- Excellent analytical, technical, troubleshooting, and organizational skills.
- Strong written and verbal communication skills, including technical documentation.
Technical Skills & Competencies
- Hands-on experience working with structured, semi-structured, and unstructured data.
- Experience developing solutions within data lake environments.
- Experience building event-driven data pipelines utilizing queues, streaming platforms, or similar technologies.
- Strong understanding of real-time and near-real-time data processing.
- Advanced hands-on experience with PySpark, Databricks, and Spark SQL.
- Experience working with file and serialization formats including:
- JSON
- Parquet
- CSV
- Other structured and semi-structured data formats
- Knowledge of NoSQL database technologies such as:
- Cosmos DB
- MongoDB
- HBase
- Similar distributed database platforms
- Experience with cloud environments such as Microsoft Azure or AWS, with Azure preferred.
- Experience with technologies including:
- Azure SQL Server
- Azure Data Lake Storage
- Azure Event Hubs
- Azure Functions
- Azure Search
- Cosmos DB
- MongoDB
- Spark Streaming
- Delta Lake
- Azure DevOps
- CI/CD pipelines
- Experience developing or working with AI agent orchestration and MCP-based workflows is highly desirable.
- Ability to troubleshoot complex data-processing and pipeline performance issues.
- Ability to manage multiple projects and priorities simultaneously in a fast-paced environment.
- Reliable, self-motivated, organized, and capable of working independently.
- Strong team player with the ability to collaborate effectively with cross-functional technical and business teams.