Job Title: Data & AI Engineer
Experience:
5–8 Years
Location:
Pune/Gurugram
Role Overview
We are looking for a
Data & AI Engineer
with 5–8 years of experience to modernize legacy data platforms and build
AI-powered solutions. The role involves migrating Hadoop/Spark workloads to
AWS and Databricks
, developing
LLM-based AI agents
, and optimizing data pipelines for scalable AI applications.
Key Responsibilities
-
Migrate legacy Hadoop, Spark, and MapReduce workloads to
AWS & Databricks
.
-
Design and develop
AI agents
using
LLMs (Anthropic Claude)
and
LangChain/Mosaic AI
.
-
Build AI-ready data pipelines for
RAG, vector search, and unstructured data processing
.
-
Maintain and optimize existing Hadoop/Spark data pipelines.
-
Improve pipeline performance using
Delta Lake
and Databricks optimization features.
-
Support BI dashboards (Tableau/Power BI/Looker) through enhancements and
performance tuning.
-
Implement Databricks AI capabilities such as
Vector Search, Mosaic AI, and Lakeflow
.
Required Skills
-
5–8 years
of Data Engineering or Software Engineering experience.
-
Strong experience with
Python (PySpark)
or
Scala
,
SQL
, and
Apache Spark
.
-
Experience with
Hadoop (HDFS, Hive, MapReduce)
.
-
Hands-on experience with
AWS (S3, Glue, EMR, IAM)
and
Databricks
.
-
Knowledge of
Delta Lake, Unity Catalog, Vector Search, Mosaic AI
.
-
Experience building
LLM/GenAI applications
, AI Agents, and
LangChain
(or similar frameworks).
-
Familiarity with
Anthropic Claude
, Copilot, or other LLM APIs.
-
Experience with BI tools like
Tableau, Power BI, or Looker
.
-
Understanding of distributed computing concepts (partitioning, caching,
shuffling, joins).
Preferred Tech Stack
-
Big Data:
Hadoop, Spark, Hive, HDFS
-
Cloud:
AWS (S3, Glue, EMR, IAM)
-
Data Platform:
Databricks, Delta Lake, Unity Catalog
-
AI/GenAI:
Anthropic Claude, LangChain, Mosaic AI, RAG, Vector Search
-
Orchestration:
Airflow, Control-M
-
BI:
Tableau, Power BI, Looker