Data Engineer
We are seeking a highly skilled and motivated Data Engineer to join our
dynamic data platform team in either our Gurgaon or Chennai office. The
successful candidate will be a critical player in designing, building, and
optimizing our next-generation data architecture and pipelines.
This role requires expert-level proficiency in PySpark and SQL for building
scalable, high-performance ETL/ELT processes that transform vast amounts of
raw data into high-quality, actionable insights for analytics, reporting, and
Machine Learning. We are looking for an immediate joiner who can hit the
ground running and contribute significantly from day one.
Key Responsibilities
Data Pipeline Development & Optimization
-
Design and Build:
Architect, develop, and maintain robust, scalable, and fault-tolerant
ETL/ELT pipelines for ingesting data from diverse sources (e.g., databases,
APIs, streaming sources) into our data lake and data warehouse.
-
PySpark Expertise:
Write and optimize complex data transformation jobs using PySpark and the
Spark DataFrame API to process petabytes of structured and unstructured data
efficiently.
-
SQL Mastery:
Utilize Advanced SQL for complex querying, data manipulation, stored
procedures, performance tuning, and optimizing database schema design in
relational and analytical databases.
-
Data Quality & Governance:
Implement data validation, cleansing, and monitoring routines to ensure high
data quality, integrity, and adherence to security and governance standards.
Architecture and Infrastructure
-
Data Modeling:
Design and implement optimal data models (e.g., Dimensional Modeling, Data
Vault, Snowflake Schema) for our data warehouse to support business
intelligence and analytical needs.
-
Cloud Integration:
Work with cloud-native data services (e.g., AWS S3, Glue, EMR, Redshift, or
Azure Data Lake, Databricks, Synapse, or Google BigQuery, Dataflow) to build
and deploy data solutions.
-
Automation:
Implement orchestration tools like Apache Airflow, Azure Data Factory, or
AWS Step Functions to automate data workflows and manage pipeline
dependencies.
Collaboration and Operational Excellence
-
Cross-Functional Teamwork:
Collaborate closely with Data Scientists, Data Analysts, Product Managers,
and Business Stakeholders to understand data requirements and translate them
into technical specifications.
-
Monitoring & Support:
Monitor, troubleshoot, and resolve issues in production data pipelines,
ensuring maximum uptime and timely data delivery.
-
Best Practices:
Participate in code reviews, enforce coding standards, and contribute to the
continuous improvement of development and deployment practices (CI/CD, Git).
Required Technical Skills (Mandatory)
-
PySpark:
Expert-level, hands-on experience in developing and optimizing large-scale
data processing applications using PySpark (Python for Apache Spark).
-
SQL:
Mastery of Advanced SQL (including window functions, complex joins, stored
procedures, and query performance tuning) across various database systems
(e.g., PostgreSQL, MySQL, Snowflake, Redshift).
-
Programming:
Strong proficiency in Python for scripting, automation, and general data
manipulation libraries (e.g., Pandas).
-
Big Data:
Solid understanding of Big Data concepts, distributed systems architecture,
and data warehousing principles.
-
ETL/ELT:
Proven experience in designing, building, and maintaining robust ETL/ELT
pipelines.
Preferred Qualifications (Good to Have)
-
Experience with Databricks or other managed Spark environments.
-
Hands-on experience with a major Cloud Platform (AWS, Azure, or GCP) and its
data-related services.
-
Familiarity with workflow orchestration tools like Apache Airflow or
comparable technologies.
-
Experience with real-time/streaming data processing (e.g., Spark Structured
Streaming, Kafka, or Kinesis).
-
Knowledge of Data Governance, Data Cataloging, and Data Security best
practices.
Candidate Profile
-
Educational Background:
Bachelor's or Master's degree in Computer Science, Engineering, or a related
quantitative field.
-
Mindset:
Proactive, self-motivated, and a strong sense of ownership and urgency.
-
Communication:
Excellent verbal and written communication skills to articulate complex
technical concepts to non-technical stakeholders.
-
Availability:
Must be an Immediate Joiner or have a short notice period (15 days maximum).