Job Description – Senior Data Engineer / Platform Re-Engineering Lead
(Azure Synapse & Databricks Migration)
Role Overview
We are looking for a highly experienced Senior Data Engineer to lead the
re-engineering of an existing enterprise data platform built on Azure
Synapse Analytics. The role requires deep technical seniority to audit,
understand, and validate a complex end-to-end data architecture spanning
source ingestion through to consumption — and to drive a future migration
of validated workloads to Databricks. This is not a greenfield role: it
demands the ability to reverse-engineer existing implementations, assess
their correctness, and own the technical migration strategy.
Key Responsibilities
-
Lead the technical assessment and re-engineering of an existing
enterprise data platform, spanning all layers from source ingestion
through to data consumption.
-
Reverse-engineer, document, and validate existing pipeline logic, data
models, transformation frameworks, and data governance controls.
-
Identify gaps, defects, and technical debt across the platform and
remediate where implementations are incorrect or sub-optimal.
-
Ensure correctness of data processing patterns including change data
capture, slowly changing dimensions, deduplication, and business
reconciliation.
-
Design and implement target-state architectures aligned to modern
lakehouse principles, ensuring feature parity and business logic
fidelity during transitions.
-
Manage platform evolution initiatives, including parallel-run phases
where multiple implementations operate simultaneously, validating output
consistency before cutover.
-
Define and execute migration strategies for existing workloads to modern
data platforms, preserving existing governance and control framework
semantics.
-
Re-implement ingestion, transformation, and orchestration pipelines on
target platforms, maintaining audit, quality, and reconciliation
standards.
-
Collaborate with business, data governance, and architecture
stakeholders to validate embedded business rules and data quality
requirements.
-
Provide technical leadership across re-engineering and migration
workstreams, contributing to decommission planning for legacy
components.
Core Technical Skills
Azure Synapse & Data Platform
Mandatory hands-on expertise with:
-
Azure Synapse Analytics (Pipelines, Spark Pool, Dedicated SQL Pool)
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Delta Lake on Azure (Synapse Lakehouse patterns)
- Oracle Golden Gate Replication for real-time source integration
- Azure Analysis Services and Power BI consumption layer patterns
-
Deep understanding of medallion architecture: Raw / Harmonized /
Conformed / Consumption layers
-
Strong knowledge of SCD Type 0/1/2, CDC patterns, soft/hard delete, and
retroactive change processing
-
Experience with Synapse SQL Pool — stored procedures, control tables,
and data quality validation patterns
-
Experience with audit, balance, and control frameworks — parameterized,
modular pipeline governance at enterprise scale
-
Familiarity with config-driven and automation-first pipeline patterns
(YAML, PySpark, SQL-driven generation from mapping documents)
Databricks & Lakehouse
Hands-on experience with Azure Databricks (Delta Live Tables, Unity
Catalog preferred)
- Strong Apache Spark skills (PySpark / Spark SQL)
-
Experience migrating workloads from legacy data warehouse or Synapse
environments to a Databricks Lakehouse
-
Ability to re-implement governance and control frameworks natively in
Databricks (audit logging, reconciliation, DQ checks)
-
Experience with Delta Lake features: MERGE, CDC, time travel, schema
enforcement
Data Engineering & Development
- Strong Python and SQL programming skills
-
Experience with ETL/ELT at scale: denormalization, surrogate keys,
directory tables, curated data models
-
Experience integrating complex data sources: Oracle DB, SQL Server,
Azure SQL DB, file systems, Salesforce, APIs
-
Strong data modelling skills: relational, dimensional, and
lakehouse-oriented
DevOps & Automation
-
CI/CD pipelines for data engineering (Azure DevOps / GitHub Actions)
- Infrastructure as Code (Terraform or ARM)
- Containerization (Docker)
-
Experience with automated testing frameworks for data pipelines (unit
testing, reconciliation-based validation)
Nice to Have
- Experience with Unity Catalog for data governance and lineage
-
Familiarity with Azure Purview for data cataloguing and governance
-
Exposure to real-time and streaming pipelines (Event Hub / Kafka /
Kinesis)
-
Experience with GenAI or ML platform integration (MLOps, feature
engineering pipelines)
-
Familiarity with monitoring and observability tools (e.g., Dynatrace)
- Exposure to BI tools (Power BI, Tableau)
Experience & Profile
-
10+ years of experience in Data Engineering, with significant platform
migration or re-engineering experience
-
Proven track record auditing and taking ownership of existing, complex
enterprise data platforms — not just building from scratch
-
Deep knowledge of enterprise data governance patterns: audit trails,
reconciliation, data quality controls, SCD versioning
-
Strong analytical mindset: ability to read existing implementations,
identify intent versus defect, and make sound re-engineering decisions
-
Comfortable operating across both hands-on engineering and technical
architecture
-
Strong communication skills — able to engage business, governance, and
engineering stakeholders with clarity
-
Experience working in regulated or enterprise-scale environments
(financial services a plus)
© 2024 Company Name. All rights reserved.