Loading open roles
Loading open roles
Loading role

GreenTree Advisory Services Pvt. Ltd. · posted 2 months ago
JOB DESCRIPTION
Databricks Tech Lead
Data Lakehouse | Delta Lake | PySpark | MLflow | Unity Catalog
Role Overview
We are looking for a skilled and passionate Databricks Engineer to design, build, and optimize enterprise-scale data lakehouse solutions on the Databricks platform. The successful candidate will be responsible for creating Databricks pipeline delivering Financial Crime platforms covering Anti-Money Laundering (AML), Know Your Customer (KYC), Customer Risk Assessment (CRA), Sanctions Screening, Transaction Monitoring, Fraud Detection, and Regulatory Reporting
Position Details
Job Title: Databricks Engineer
Department: Client Delivery / Data & AI Engineering
Experience Required: 8-10 + Years (overall) | 5+ Years Databricks hands-on
Employment Type: Full-Time
Location: Hybrid / Remote (as per business requirement)
Reporting To: Delivery Manager / Technical Project Manager
Key Responsibilities
Databricks Platform Engineering
● Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
● Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage.
● Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance.
● Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
● Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management.
Delta Lake & Lakehouse Architecture
● Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM).
● Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
● Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations.
● Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
● Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).
Data Pipeline Development (PySpark / SQL)
● Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
● Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
● Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning.
● Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
● Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines.
MLflow & AI/ML Workloads
● Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks.
● Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances.
● Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features.
● Enable GenAI workloads — LLM fine-tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI).
● Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints.
Cloud Integration & DevOps
● Integrate Databricks with cloud-native services: Azure Data Lake Storage (ADLS).
● Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI.
● Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure-as-code (IaC) deployment of Databricks resources.
● Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte).
● Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools.
Governance, Security & Optimization
● Implement row-level security, column masking, and dynamic data views using Unity Catalog policies.
● Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations.
● Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement.
● Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog.
● Document architecture decisions, runbooks, and operational guides for Databricks workloads.
Required Qualifications
Education
● Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or related field.
Experience
● 8+ years of total experience in data engineering or software engineering.
● 3+ years of dedicated hands-on experience with the Databricks platform in production environments.
● Strong background in big data engineering, cloud data platforms, and distributed computing.
Databricks Platform
● Deep expertise in Databricks Workspaces, Clusters, Jobs, Workflows, and Repos.
● Proficiency with Unity Catalog — metastore setup, catalog/schema/table management, access controls, and data lineage.
● Hands-on experience with Delta Live Tables (DLT) — pipeline development, expectations, and monitoring.
● Strong command of Delta Lake internals — transaction log, ACID guarantees, file layout, and optimization techniques.
● Experience with Databricks SQL Warehouses, SQL Analytics, and dashboard creation.
● Knowledge of Databricks Photon engine, serverless compute, and cost optimization strategies.
PySpark & SQL
● 4+ years of PySpark development — DataFrames, Datasets, Spark SQL, RDD operations.
● Expert-level SQL — window functions, lateral joins, CTEs, recursive queries, and analytical functions.
● Experience with Spark performance tuning — AQE, query plans (EXPLAIN), partitioning, and caching.
● Proficiency with Python for pipeline development, utilities, and automation.
Cloud Platforms
● Hands-on experience with at least one: Azure (ADLS Gen2, ADF, Azure Databricks), AWS (S3, EMR, Glue, AWS Databricks), or GCP (GCS, BigQuery, Dataproc).
● Experience with cloud networking for Databricks: VNet/VPC injection, private endpoints, and firewall configurations.
● Familiarity with IAM roles, managed identities, and service principal authentication for Databricks.
MLflow & ML Engineering (Nice to Have)
● Working knowledge of MLflow — experiment tracking, model registry, and deployment.
● Experience supporting ML pipelines on Databricks for training, evaluation, and serving.
● Exposure to Databricks Feature Store and Mosaic AI / GenAI capabilities.
Preferred Certifications
● Databricks Certified Associate Developer for Apache Spark (PySpark or Scala).
● Databricks Certified Data Engineer Associate / Professional — strongly preferred.
● Databricks Certified Machine Learning Associate / Professional.
● Azure Data Engineer Associate (DP-203) / AWS Data Analytics Specialty / GCP Professional Data Engineer.
● dbt Analytics Engineer Certification.
Preferred Qualifications
● Experience with dbt (data build tool) for SQL-based transformation on Databricks SQL.
● Knowledge of Apache Kafka / Confluent for real-time streaming into Databricks.
● Familiarity with Terraform or Pulumi for Databricks infrastructure-as-code.
● Exposure to Apache Iceberg or Apache Hudi in addition to Delta Lake.
● Experience with BI tool integration: Power BI, Tableau, or Looker connected to Databricks SQL.
● Knowledge of data mesh principles and federated data governance at scale.
● Domain experience in BFSI, Healthcare, Retail, or Manufacturing data programs.
Core Competencies
● Deep technical depth in Databricks with the ability to architect and troubleshoot complex lakehouse systems.
● Strong problem-solving skills — ability to diagnose pipeline failures, performance bottlenecks, and data quality issues.
● Collaborative team player comfortable working with data engineers, ML engineers, and business stakeholders.
● Excellent documentation and communication skills for technical and non-technical audiences.
● Continuous learner — actively keeps pace with Databricks platform updates and the broader data/AI ecosystem.
● Delivery-oriented mindset with experience in Agile/Scrum delivery environments.
What We Offer
● Work on large-scale, production Databricks implementations for marquee enterprise clients.
● Exposure to the full Databricks ecosystem — lakehouse, streaming, GenAI, and MLOps.
● Databricks certification sponsorship and continuous learning support.
● Competitive compensation, performance bonuses, and comprehensive benefits.
● Hybrid/remote work flexibility and inclusive, high-performance team culture.
● Direct collaboration with Databricks account teams and technical partners.
Confidential | Client Delivery | EXL Service | 2026