Job Description – Data Engineer
We are looking for an experienced
Data Engineer
with strong expertise in
Python, PySpark, SQL, and AWS Cloud
. The candidate will be responsible for building scalable data pipelines,
developing ETL workflows, and working with AWS data engineering services.
Key Responsibilities
-
Design, develop, and maintain scalable
data pipelines and ETL workflows
.
-
Develop data processing solutions using
Python and PySpark
.
-
Write complex and optimized
SQL queries
for data extraction, transformation, and analysis.
-
Build and manage ETL pipelines using
AWS Glue
.
-
Work with
AWS Glue Data Catalog
for metadata and data discovery.
-
Use
Amazon S3
for scalable and secure data storage.
-
Develop serverless data-processing solutions using
AWS Lambda
.
-
Implement event-driven workflows using
Amazon EventBridge
.
-
Orchestrate data pipelines and workflows using
AWS Step Functions
.
-
Use
Amazon Athena
for querying and analyzing data stored in Amazon S3.
-
Monitor, troubleshoot, and optimize data pipelines for performance and
reliability.
-
Collaborate with data analysts, data scientists, and other engineering
teams to deliver data solutions.
Required Skills
-
Strong hands-on experience with
Python
-
Strong experience with
PySpark
-
Advanced
SQL
-
Hands-on experience with
AWS Cloud
-
Amazon S3
-
AWS Glue – ETL & Data Catalog
-
AWS Lambda
-
Amazon EventBridge
-
AWS Step Functions
-
Amazon Athena
-
Good understanding of
ETL/ELT concepts, data pipelines, and data processing
-
Strong problem-solving and analytical skills
Preferred
-
Experience with AWS-based data engineering architectures
-
Knowledge of data warehousing and data lake concepts
-
Experience with performance optimization and production data pipelines