Hadoop Data Engineer
Job Summary:
The Hadoop Data Engineer plays a critical role in managing and optimizing our
big data infrastructure to ensure efficient data processing and analytics
capabilities. You will contribute to the design, implementation, and
maintenance of scalable ETL pipelines that support analytical initiatives
within the organization, driving impactful data-driven decisions.
Key Responsibilities:
-
Design and develop robust ETL pipelines using Python, PySpark, and Hadoop to
process large datasets efficiently.
-
Optimize data workflows and improve data processing speeds while ensuring
data quality and integrity.
-
Collaborate with data scientists and analysts to understand data
requirements and provide comprehensive solutions.
-
Monitor and maintain the Hadoop ecosystem, ensuring high availability and
performance of data storage and processing.
-
Implement best practices for data governance and security, ensuring
compliance with industry standards.
-
Participate in troubleshooting and incident response related to data
pipeline failures and performance issues.
-
Continuously explore and evaluate new technologies and tools to enhance the
overall data environment.
Requirements:
-
Bachelor's degree in Computer Science, Engineering, or a related field.
-
8-12 years of hands-on experience in building and managing data pipelines
using Hadoop ecosystem tools.
-
Proficiency in Python and PySpark for data manipulation and ETL development.
-
Strong SQL skills for data querying and transformation within relational and
non-relational databases.
-
Extensive experience with big data technologies, including HDFS, Hive, and
Kafka.
-
Demonstrated knowledge of data warehousing concepts and data modeling.
-
Excellent problem-solving skills and the ability to work collaboratively in
a hybrid environment.
Preferred Qualifications:
-
Familiarity with cloud services such as AWS, Azure, or Google Cloud Platform
(GCP).
-
Experience with containerization and orchestration technologies like Docker
and Kubernetes.
-
Knowledge of machine learning frameworks and their application in data
engineering.
-
Strong communication skills with the capability to present complex
information clearly to non-technical stakeholders.
Benefits:
- Competitive salary with performance-based bonuses and incentives.
- Flexible hybrid work model with options for remote work.
-
Comprehensive health and wellness benefits, including medical, dental, and
vision coverage.
- Retirement plan with company matching contributions.
-
Continuous learning and professional development opportunities, including
certifications.
- Dynamic work culture that promotes innovation and collaboration.
-
Access to cutting-edge technologies and projects in the big data space.