Job Summary
We are seeking an experienced Databricks Data Engineer with 9-12 years of expertise in designing, developing, and optimizing scalable data solutions on cloud platforms. The ideal candidate will have strong hands-on experience with Databricks, Apache Spark, PySpark, Delta Lake, and modern data engineering practices. The candidate will be responsible for building enterprise-grade data pipelines, data lakehouse solutions, and supporting analytics and AI/ML initiatives.
Key Responsibilities
· Design, develop, and maintain scalable ETL/ELT pipelines using Databricks and PySpark.
- Build and optimize data ingestion frameworks for batch and real-time data processing.
- Implement and manage Delta Lake architectures and Lakehouse solutions.
- Develop high-performance data transformations and workflows on large-scale datasets.
- Collaborate with Data Architects, Data Scientists, Business Analysts, and DevOps teams to deliver data solutions.
- Design and implement data models for analytics and reporting requirements.
- Optimize Spark jobs, cluster configurations, and query performance.
- Ensure data quality, consistency, governance, and security across data platforms.
- Automate data workflows using orchestration tools such as Airflow, Azure Data Factory, or similar technologies.
- Establish monitoring, alerting, and operational support processes for data pipelines.
- Participate in architecture discussions and contribute to cloud migration initiatives.
- Mentor junior data engineers and drive engineering best practices.
Required Skills Databricks & Big Data
· Databricks Workspace Administration
- Apache Spark
- PySpark
- Spark SQL
- Delta Lake
- Unity Catalog
- Databricks Workflows
- Structured Streaming
Cloud Platforms (Any One Preferred)
· Azure (Preferred)
o Azure Data Lake Storage (ADLS)
- Azure Data Factory (ADF)
- Azure Synapse Analytics
- Azure Key Vault
- AWS
o S3
- Glue
- EMR
- Redshift
- Google Cloud Platform (GCP)
o BigQuery
Data Engineering
· ETL/ELT Frameworks
- Data Warehousing Concepts
- Data Lakehouse Architecture
- Data Modeling
- Real-time Data Processing
- Data Governance & Quality
Programming
· Python
- SQL
- PySpark
- Shell Scripting
DevOps & CI/CD
· Git/GitHub/GitLab
- Azure DevOps
- Jenkins
- Terraform
- Infrastructure as Code (IaC)
Preferred Qualifications
· Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or related field.
- Databricks Certified Data Engineer Associate/Professional certification.
- Cloud certifications such as Azure Data Engineer Associate, AWS Data Analytics, or GCP Professional Data Engineer.
- Experience working in Agile/Scrum environments.
- Strong understanding of data governance, security, and compliance frameworks.
Desired Experience
· 9-12 years of overall IT experience.
- Minimum 4+ years of hands-on experience with Databricks.
- 5+ years of experience in Spark/PySpark development.
- Experience implementing enterprise-scale Lakehouse architectures.
- Exposure to Machine Learning data pipelines and MLOps is desirable.
- Experience handling large-scale structured and semi-structured datasets.
Key Competencies
· Problem Solving & Analytical Thinking
- Stakeholder Management
- Technical Leadership
- Solution Design
- Performance Optimization
- Mentoring and Team Collaboration
- Communication and Presentation Skills
Nice to Have
· AI/ML pipeline integration
- MLOps exposure
- Data Catalog and Governance tools
- Snowflake integration with Databricks
- Kafka/Event Streaming platforms
- Data Quality tools such as Great Expectations or Deequ
Technology Tower: Data & AI → Data Engineering → Databricks / Spark / Lakehouse Engineering