GCP Data Engineer
Job Summary:
We are seeking an experienced GCP Data Engineer to join our dynamic data team.
This role will focus on designing and implementing robust data pipelines using
Python and PySpark, ensuring seamless integration and transformation of data
across our cloud infrastructure.
Key Responsibilities:
-
Develop and optimize data extraction, transformation, and loading (ETL)
processes utilizing Google Cloud Platform’s tools and services.
-
Design data models and implement data pipelines using Python and PySpark for
efficient data processing.
-
Manage and orchestrate workflows with Apache Airflow to automate data tasks
and maintain data reliability.
-
Collaborate with cross-functional teams to understand data requirements and
provide effective data solutions.
-
Monitor and troubleshoot data pipelines to ensure data integrity and
performance.
-
Document data processes and provide training to team members on best
practices and tools.
Requirements:
-
Bachelor’s degree in Computer Science, Engineering, or a related field.
-
2+ years of hands-on experience in data engineering roles, specifically with
GCP.
-
Proficiency in Python and PySpark for data manipulation and analysis.
-
Strong SQL skills for querying and managing relational databases.
-
Experience with data pipeline orchestration tools, particularly Apache
Airflow.
-
Familiarity with cloud services (GCP) and data storage solutions.
-
Excellent problem-solving skills and a keen attention to detail.
Preferred Qualifications:
-
Experience with Google BigQuery and data lake architectures.
-
Knowledge of machine learning algorithms and data modeling techniques.
-
Familiarity with data governance and compliance standards.
-
Strong communication skills and ability to work effectively in a
team-oriented environment.