Position Overview
We are seeking a
hands-on Data Engineer (5–7 years)
to build and support a modern, governed data platform for a leading US-based
insurance organization. The role focuses on
implementation and delivery
under defined architecture, with strong emphasis on
PySpark, SQL, Azure Fabric Dataflows Gen2, and knowledge of CI/CD
.
Responsibilities
-
Build and maintain
Azure Fabric data pipelines
using
Pipelines, Dataflows Gen2, and Notebooks (PySpark)
.
-
Implement
Lakehouse-first / Medallion (ODS, Bronze–Silver–Gold)
data patterns, including
Type-2 SCD
and golden records.
-
Develop
PySpark transformations
; use
SQL extensively
for validation, reconciliation, and performance tuning.
-
Create
Lakehouse SQL views and materialized views
, enabling
DirectLake
consumption.
-
Ingest data from
on-prem SQL Server DWHs
and
API-based sources
(via Azure API Management/connectors).
-
Support
on-prem → Azure Fabric migrations
, including historical data loads and data quality validation.
-
Implement
event-driven ingestion
where required using
Fabric Event Streams and KQL
.
-
Operate within
governed Fabric environments
(OneLake layout, RBAC, standards).
-
Monitor pipelines, handle failures, and optimize performance using logs
and diagnstics.
-
Contribute to
CI/CD pipelines
using
Azure DevOps or Bitbucket
.
Required Experience
-
5–7 years in
data engineering / ETL / DWH development
.
-
Strong hands-on skills in
PySpark and SQL
.
-
Practical experience with
Azure Fabric
(Lakehouse, OneLake, Dataflows Gen2, Pipelines) for 3 to 5 years
-
Solid understanding of
Medallion architecture, SCD Type-2
, and cloud data migrations.
-
Experience with
API-based ingestion
and
CI/CD for data platforms
.
-
Familiarity with
data governance and security
in regulated environments.