Loading open roles
Loading open roles
Loading role

GreenTree Advisory Services Pvt. Ltd. · posted 1 month ago
Job Description – Senior Data Engineer / Platform Re-Engineering Lead (Azure Synapse & Databricks Migration)
Role Overview We are looking for a highly experienced Senior Data Engineer to lead the re-engineering of an existing enterprise data platform built on Azure Synapse Analytics. The role requires deep technical seniority to audit, understand, and validate a complex end-to-end data architecture spanning source ingestion through to consumption — and to drive a future migration of validated workloads to Databricks. This is not a greenfield role: it demands the ability to reverse-engineer existing implementations, assess their correctness, and own the technical migration strategy.
Key Responsibilities
· Lead the technical assessment and re-engineering of an existing enterprise data platform, spanning all layers from source ingestion through to data consumption
· Reverse-engineer, document, and validate existing pipeline logic, data models, transformation frameworks, and data governance controls
· Identify gaps, defects, and technical debt across the platform and remediate where implementations are incorrect or sub-optimal
· Ensure correctness of data processing patterns including change data capture, slowly changing dimensions, deduplication, and business reconciliation
· Design and implement target-state architectures aligned to modern lakehouse principles, ensuring feature parity and business logic fidelity during transitions
· Manage platform evolution initiatives, including parallel-run phases where multiple implementations operate simultaneously, validating output consistency before cutover
· Define and execute migration strategies for existing workloads to modern data platforms, preserving existing governance and control framework semantics
· Re-implement ingestion, transformation, and orchestration pipelines on target platforms, maintaining audit, quality, and reconciliation standards
· Collaborate with business, data governance, and architecture stakeholders to validate embedded business rules and data quality requirements
· Provide technical leadership across re-engineering and migration workstreams, contributing to decommission planning for legacy components
Core Technical Skills
· Azure Synapse & Data Platform Mandatory hands-on expertise with:
· Azure Synapse Analytics (Pipelines, Spark Pool, Dedicated SQL Pool)
· Azure Data Lake Storage Gen2 (ADLS Gen2)
· Delta Lake on Azure (Synapse Lakehouse patterns)
· Oracle Golden Gate Replication for real-time source integration
· Azure Analysis Services and Power BI consumption layer patterns
· Deep understanding of medallion architecture: Raw / Harmonized / Conformed / Consumption layers
· Strong knowledge of SCD Type 0/1/2, CDC patterns, soft/hard delete, and retroactive change processing
· Experience with Synapse SQL Pool — stored procedures, control tables, and data quality validation patterns
· Experience with audit, balance, and control frameworks — parameterized, modular pipeline governance at enterprise scale
· Familiarity with config-driven and automation-first pipeline patterns (YAML, PySpark, SQL-driven generation from mapping documents)
Databricks & Lakehouse
· Hands-on experience with Azure Databricks (Delta Live Tables, Unity Catalog preferred)
· Strong Apache Spark skills (PySpark / Spark SQL)
· Experience migrating workloads from legacy data warehouse or Synapse environments to a Databricks Lakehouse
· Ability to re-implement governance and control frameworks natively in Databricks (audit logging, reconciliation, DQ checks)
· Experience with Delta Lake features: MERGE, CDC, time travel, schema enforcement
· Data Engineering & Development
· Strong Python and SQL programming skills
· Experience with ETL/ELT at scale: denormalization, surrogate keys, directory tables, curated data models
· Experience integrating complex data sources: Oracle DB, SQL Server, Azure SQL DB, file systems, Salesforce, APIs
· Strong data modelling skills: relational, dimensional, and lakehouse-oriented
DevOps & Automation
· CI/CD pipelines for data engineering (Azure DevOps / GitHub Actions)
· Infrastructure as Code (Terraform or ARM)
· Containerization (Docker)
· Experience with automated testing frameworks for data pipelines (unit testing, reconciliation-based validation)
Nice to Have
· Experience with Unity Catalog for data governance and lineage
· Familiarity with Azure Purview for data cataloguing and governance
· Exposure to real-time and streaming pipelines (Event Hub / Kafka / Kinesis)
· Experience with GenAI or ML platform integration (MLOps, feature engineering pipelines)
· Familiarity with monitoring and observability tools (e.g., Dynatrace)
· Exposure to BI tools (Power BI, Tableau)
Experience & Profile
· 10+ years of experience in Data Engineering, with significant platform migration or re-engineering experience
· Proven track record auditing and taking ownership of existing, complex enterprise data platforms — not just building from scratch
· Deep knowledge of enterprise data governance patterns: audit trails, reconciliation, data quality controls, SCD versioning
· Strong analytical mindset: ability to read existing implementations, identify intent versus defect, and make sound re-engineering decisions
· Comfortable operating across both hands-on engineering and technical architecture
· Strong communication skills — able to engage business, governance, and engineering stakeholders with clarity
· Experience working in regulated or enterprise-scale environments (financial services a plus)