Lead Data Engineer (Databricks)
Job Summary:
We are seeking a highly skilled Lead Data Engineer with exceptional expertise
in Databricks, PySpark, and SQL to drive the design and implementation of
scalable data pipelines and modern lakehouse architectures. This pivotal role
enables organizations to harness valuable insights from extensive datasets
while collaboratively delivering high-performance data solutions that empower
analytics and machine learning initiatives.
Key Responsibilities:
-
Design, build, and maintain ETL/ELT pipelines on Databricks using PySpark,
Spark SQL, and Delta Lake.
-
Develop and optimize data ingestion frameworks, data transformations, and
end-to-end workflows for both batch and streaming use cases.
-
Implement Delta Lake-based architectures, including versioning, schema
evolution, and ACID-compliant pipelines.
-
Engage with stakeholders to gather data requirements and translate them into
scalable data engineering solutions.
-
Manage and optimize Databricks clusters, jobs, and notebooks for performance
and cost efficiency.
-
Ensure data quality, reliability, and observability through validation
frameworks and monitoring processes.
-
Contribute to data modeling, metadata management, and the establishment of
best practices within the data platform.
Requirements:
-
5+ years of experience in data engineering, with 3+ years of hands-on
expertise in Databricks.
-
7+ years of overall experience in Data Engineering development.
-
Strong understanding of DWH concepts and ETL processes/standards.
-
Proven ability to lead a team of 2-3 members effectively.
-
Hands-on experience with Spark (PySpark/Spark SQL) and distributed data
processing.
-
Solid SQL knowledge and experience working with large-scale datasets.
-
Strong understanding of Delta Lake, medallion architecture, and scalable
lakehouse patterns.
Preferred Qualifications:
-
Experience with CI/CD, Git, and modern DevOps practices for data pipelines.
-
Familiarity with both structured and unstructured data, as well as data
quality frameworks.
-
Knowledge of performance tuning methodologies for data environments.
-
Experience collaborating with data scientists and business stakeholders.