Location:
Bangalore
Experience:
8–12 Years
Work Mode:
Hybrid
Joining:
Mid to Last Week of September
Job Description
We are looking for an experienced
Hadoop Data Engineer
with strong hands-on expertise in
Python, PySpark, Hadoop, and SQL
. The ideal candidate should have experience in developing and optimizing
large-scale data pipelines and working with distributed data processing
technologies.
Key Responsibilities
-
Design, develop, and maintain scalable
data pipelines
using Python and PySpark.
-
Work with
Hadoop ecosystem
for large-scale data storage and processing.
-
Develop complex and optimized
SQL queries
for data extraction and transformation.
-
Perform data cleansing, transformation, and integration from multiple
sources.
-
Optimize PySpark and Hadoop jobs for better performance and scalability.
-
Troubleshoot data pipeline and production issues.
-
Collaborate with data engineers, analysts, and business teams to deliver
reliable data solutions.
-
Ensure data quality, accuracy, and consistency across pipelines.
Must-Have Skills
-
8–12 years
of experience in Data Engineering.
-
Strong hands-on experience with
Python
.
-
Strong expertise in
PySpark
.
-
Hands-on experience with
Hadoop / HDFS
.
-
Strong
SQL
skills.
-
Experience in developing and optimizing
ETL/ELT data pipelines
.
-
Good understanding of distributed data processing and large-scale datasets.
Good to Have
-
Experience with
Hive, YARN, MapReduce, Spark
.
-
Experience with cloud platforms such as
AWS / Azure / GCP
.
-
Experience with data warehousing and Big Data technologies.
-
Knowledge of data quality and performance optimization.