- Full-time
- Permanent
- Quick Apply
Senior Data Engineer with Generative AI expertise to design, build, and optimize scalable cloud-native data platforms and AI-powered data solutions.
The ideal candidate will have strong experience in Python, AWS Data Services, PySpark, Data Warehousing, and Modern Data Engineering practices, combined with hands-on exposure to Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Vector Databases, and AI-driven data pipelines.
This role requires a combination of data engineering excellence, cloud architecture expertise, and the ability to enable enterprise AI/ML initiatives through robust and scalable data foundations.
Good understanding of sales, marketing, distribution functions for an asset management firm is needed.
Good understanding of Client, account, sales, marketing performance and investment data is needed.
Preferred Qualifications
- 8+ years of experience in Data Engineering and Cloud Data Platforms.
- 2+ years of experience working with Generative AI, LLMs, or AI Engineering solutions.
- AWS Certifications such as:
- AWS Certified Data Engineer
- AWS Solutions Architect
- AWS Machine Learning Specialty
- Experience with MLOps and AI Operations frameworks.
Required Technical Skills
Data Engineering & Cloud
- Strong programming expertise in Python, including data engineering frameworks and APIs.
- Hands-on experience with AWS Glue, AWS Lambda, Amazon S3, AWS Step Functions, and Event-Driven Architectures.
- Advanced experience with PySpark and distributed data processing frameworks.
- Expertise in designing and optimizing ETL/ELT pipelines for large-scale data movement and transformation.
- Strong experience with Snowflake, Amazon Redshift, or similar cloud data warehouse platforms.
Deep knowledge of SQL, including:
- Complex query development
- Data modeling
- Query optimization and performance tuning
- Data analytics and troubleshooting
Experience with data orchestration tools such as Apache Airflow, AWS MWAA, Prefect, or similar frameworks
Strong understanding of:
- Data Warehousing concepts
- Dimensional Modeling (Star/Snowflake Schema)
- Data Lakes and Lakehouse Architectures
- Cloud-native architecture patterns
- Data Governance and Metadata Management
Generative AI & AI Engineering:
- Hands-on experience with Generative AI solutions and Large Language Models (LLMs) such as:
- OpenAI GPT / Anthropic Claude/Amazon Bedrock/Azure OpenAI
- Experience building Retrieval-Augmented Generation (RAG) solutions.
- Knowledge of Vector Databases such as:
- Pinecone / Amazon OpenSearch
- Experience developing AI-powered applications using:
- LangChain /AI Agent Frameworks/ AWS AgentCore
Familiarity with:
- Prompt Engineering
- Embedding Models
- Model Evaluation and Monitoring
- AI Governance and Responsible AI Practices
Experience integrating enterprise data platforms with AI/ML and GenAI ecosystems.
DevOps & Engineering Best Practices
Experience with CI/CD pipelines using GitHub Actions, Jenkins, Azure DevOps, or similar tools.
Strong understanding of Infrastructure as Code (IaC) using Terraform or AWS CloudFormation.
Experience with containerization technologies such as Docker and Kubernetes.
Knowledge of data security, compliance, and cloud governance best practices.
Key Responsibilities:
- Design, develop, and maintain scalable, secure, and high-performance data pipelines on AWS.
- Build and optimize enterprise-grade ETL/ELT solutions for structured and unstructured data.
- Develop data architectures that support analytics, machine learning, and Generative AI workloads.
- Design and implement RAG-based solutions leveraging vector databases and enterprise knowledge repositories.
- Collaborate with Data Scientists, AI Engineers, Product Owners, and Business Stakeholders to deliver AI-enabled data products.
- Build reusable data services, APIs, and frameworks that accelerate AI adoption across the organization.
- Ensure data quality, lineage, governance, observability, and operational excellence across the data ecosystem.
- Optimize Snowflake/Redshift environments for performance, scalability, and cost efficiency.
- Support AI model deployment, monitoring, and integration with modern cloud platforms.
- Mentor junior engineers and contribute to data engineering best practices, architecture reviews, and technical guidance.
Nice to Have
- Experience with Knowledge Graphs.
- Exposure to Agentic AI architectures and autonomous AI workflows.
- Experience with Databri cks, Delta Lake, and Apache Iceberg.
- Familiarity with Microsoft Fabric, Azure OpenAI, and Microsoft Copilot ecosystem.
- Experience implementing enterprise AI governance and compliance frameworks.