Job Title:
Databricks Engineer
Experience:
4–16 Years Overall | 5+ Years Hands-on Databricks Experience
Employment Type:
Full-Time
Location:
Hybrid
Department:
Data & AI Engineering / Client Delivery
Role Overview
We are looking for an experienced
Databricks Engineer
to design, develop, and optimize enterprise-scale
Data Lakehouse solutions
using Databricks, Delta Lake, PySpark, SQL, and Unity Catalog.
The role will involve building scalable data pipelines and supporting
financial crime platforms covering
Anti-Money Laundering (AML), Know Your Customer (KYC), Customer Risk
Assessment (CRA), Sanctions Screening, Transaction Monitoring, Fraud
Detection, and Regulatory Reporting
.
Key Responsibilities
-
Design, develop, and maintain
Databricks Workspaces, Clusters, Jobs, Workflows, and Compute Pools
across development, testing, and production environments.
-
Implement
Unity Catalog
for data governance, access control, fine-grained permissions, and
data lineage.
-
Design and implement
Delta Lake
tables with partitioning, Z-Ordering, OPTIMIZE, VACUUM, schema
evolution, Time Travel, and Change Data Feed (CDF).
-
Build
Medallion Architecture
using Bronze, Silver, and Gold layers.
-
Develop reliable
Delta Live Tables (DLT)
pipelines with data quality expectations.
-
Develop scalable batch and streaming data pipelines using
PySpark, Spark SQL, and Delta Lake
.
-
Build real-time ingestion pipelines using
Kafka, Event Hubs, or Kinesis
.
-
Optimize Spark workloads using
Broadcast Joins, Adaptive Query Execution (AQE), Dynamic Partition
Pruning, caching, and query plan analysis
.
-
Implement robust error handling, retry mechanisms, and dead-letter
queue patterns.
-
Integrate Databricks with cloud platforms and services such as
Azure ADLS, AWS S3, EMR, Glue, or GCP
.
-
Implement CI/CD pipelines using
Azure DevOps, GitHub Actions, or GitLab CI
.
-
Use
Databricks Asset Bundles (DABs) or Terraform
for Infrastructure as Code (IaC).
-
Manage data ingestion using
Auto Loader, COPY INTO, Fivetran, dbt, or Airbyte
.
-
Monitor pipeline health, cluster utilization, and cost optimization.
-
Support
MLflow, Feature Store, MLOps, GenAI, RAG pipelines, and Vector
Search
workloads as required.
-
Implement security practices including
row-level security, column masking, dynamic views, IAM, managed
identities, and service principals
.
-
Document architecture decisions, operational runbooks, and technical
processes.
Mandatory Skills
-
Strong hands-on experience with
Databricks
-
PySpark
-
Python
-
Advanced SQL
-
Delta Lake
-
Delta Live Tables (DLT)
-
Unity Catalog
-
Databricks Workflows and Jobs
-
Experience with
Azure / AWS / GCP
-
Strong understanding of
Big Data and Distributed Computing
-
Experience in
data pipeline development and performance tuning
Preferred Skills
-
MLflow
-
Databricks Feature Store
-
Mosaic AI / GenAI
-
RAG and Vector Search
-
Apache Kafka / Confluent
-
dbt
-
Terraform / Pulumi
-
Apache Iceberg / Apache Hudi
-
Power BI / Tableau / Looker
-
Experience with
BFSI, Healthcare, Retail, or Manufacturing
domains
Qualifications
-
Bachelor's or Master's degree in
Computer Science, Information Technology, Data Engineering
, or a related field.
-
4+ years
of experience in Data Engineering or Software Engineering.
-
Strong hands-on experience with Databricks in production environments.
-
Excellent problem-solving, troubleshooting, communication, and
documentation skills.
-
Experience working in
Agile/Scrum
delivery environments.
Certifications – Preferred
-
Databricks Certified Data Engineer Associate / Professional
-
Databricks Certified Associate Developer for Apache Spark
-
Databricks Certified Machine Learning Associate / Professional
-
Azure Data Engineer Associate (DP-203)
-
AWS Data Analytics Specialty
-
GCP Professional Data Engineer
-
dbt Analytics Engineer Certification
What We Offer
-
Opportunity to work on large-scale, enterprise Databricks
implementations.
-
Exposure to the complete Databricks ecosystem, including
Lakehouse, Streaming, GenAI, and MLOps
.
-
Continuous learning and certification support.
-
Competitive compensation and comprehensive benefits.
-
Hybrid/remote work flexibility.
-
Collaboration with Databricks account teams and technical partners.