Data Engineer
Job Summary:
As a Data Engineer, you will be at the forefront of building and optimizing
our data infrastructure, playing a crucial role in shaping how we leverage big
data for business insights. Your expertise in ETL processes and reporting
tools will empower our teams to make data-driven decisions efficiently and
effectively.
Key Responsibilities:
-
Design and implement scalable ETL pipelines using technologies such as
PySpark, DataBricks, and Genie to ensure seamless data flow and
transformation.
-
Create and maintain efficient SQL queries and scripts for data retrieval,
manipulation, and storage across various data sources.
-
Optimize data processing and storage solutions within Hive to enhance
performance and reduce costs.
-
Collaborate with data analysts and data scientists to understand
requirements and deliver actionable data insights through Power BI and
Tableau.
-
Develop and maintain documentation for data architecture, data models, and
ETL processes to ensure clarity and knowledge transfer within the team.
-
Conduct data quality checks and troubleshooting to ensure data accuracy and
reliability throughout the pipeline.
-
Stay updated on industry trends and emerging technologies to drive
innovation in data engineering practices and methodologies.
Requirements:
-
Minimum of 5 years of experience in data engineering or related fields,
demonstrating a strong understanding of data architecture principles.
-
Proficiency in SQL for managing and querying relational databases.
-
Experience with big data technologies such as Hive and PySpark for data
processing.
-
Hands-on experience with ETL tools and frameworks, particularly in creating
and optimizing pipelines with DataBricks and Genie.
-
Familiarity with data visualization tools such as Power BI or Tableau for
reporting and analytics.
-
Strong analytical and problem-solving skills, with a keen attention to
detail.
-
Excellent communication and collaboration skills to work effectively in
cross-functional teams.
Preferred Qualifications:
-
Experience with cloud platforms like AWS, Azure, or Google Cloud for data
engineering solutions.
-
Knowledge of programming languages such as Python or Scala to enhance data
processing capabilities.
-
Familiarity with data governance and data security best practices to protect
data integrity.
-
Experience in developing real-time data processing solutions and
event-driven architectures.