About the job
We are seeking a highly skilled Senior Data / ML Engineer to join our Data
Science & Engineering team. In this role you will design, build, and
maintain robust, scalable data pipelines and infrastructure on AWS, powering
analytics, machine learning, and business-critical reporting. You will work
closely with data scientists, analysts, and product teams to ensure reliable,
performant data delivery across the organization.
The ideal candidate combines deep AWS expertise with modern development
practices - including AI-assisted coding, containerization, and infrastructure
as code - to accelerate delivery without sacrificing quality.
Key Responsibilities
-
Design, develop, and optimize large-scale ETL/ELT pipelines using AWS
services such as EMR (Spark), Glue, Lambda, and Step Functions.
-
Orchestrate complex data workflows with Apache Airflow (Amazon MWAA or
self-managed), ensuring reliability, observability, and SLA adherence.
-
Architect and manage data storage solutions across Amazon S3 (data lake),
Redshift (data warehouse), and RDS (relational databases), applying best
practices for partitioning, compression, and cost optimization.
-
Build and maintain containerized data applications and microservices using
Docker and Amazon ECS/Fargate, including CI/CD automation.
-
Develop event-driven and serverless data processing solutions with AWS
Lambda, SQS, SNS, and EventBridge.
-
Leverage AI-powered coding assistants and IDE integrations (e.g., Kiro,
Cursor, Claude Code) to accelerate development, code review, and
documentation.
-
Implement data quality frameworks, monitoring, and alerting to ensure data
integrity across all pipelines.
-
Collaborate with Data Scientists to productionize ML models and feature
pipelines.
-
Define and enforce data governance, security, and access-control policies in
line with organizational and regulatory standards.
-
Contribute to infrastructure-as-code initiatives using Terraform,
CloudFormation, or CDK.
Required Qualifications
-
Bachelor's or Master's degree in Computer Science, Engineering, Data
Science, or a related field.
-
5+ years of professional experience in data engineering, with at least 3
years of hands-on AWS production workloads.
-
3+ years of experience with AWS Services: EMR (Spark/Hadoop), Apache
Airflow, S3, Redshift, RDS, Lambda, Sagemaker and ECS.
-
Solid experience with Docker (building, optimizing, and deploying
containers) and container orchestration.
-
Proficiency in Python and SQL; experience with Typescript is a plus.
-
Demonstrated ability to use AI-assisted development tools within modern IDEs
for rapid prototyping, code generation, testing and debugging.
-
Strong understanding of data modeling, data governance, and data security
best practices.
-
AWS certifications such as AWS Certified Data Analytics - Specialty or AWS
Certified Solutions Architect are a plus.
-
Experience with infrastructure-as-code tools (Terraform, CloudFormation,
CDK).
-
Exposure to ML Ops workflows, feature stores, or model serving pipelines
(e.g., SageMaker).
-
Knowledge of cost-optimization strategies for large-scale AWS data
workloads.
Tech Stack At a Glance
-
Compute & Processing: EMR (Spark), Lambda, ECS/Fargate, Glue, Sagemaker
-
Orchestration: Apache Airflow (MWAA), Step Functions
-
Storage & Warehousing: S3, Redshift, RDS (PostgreSQL / MySQL)
-
Containers & DevOps: Docker, ECS, GitHub Actions
-
AI-Assisted Dev: Cursor, Kiro, Claude Code
-
Languages: Python, SQL, Bash