Loading open roles
Loading open roles
Loading role

GirnarSoft · posted 5 months ago
Job Description:
🚀 Job Description: Data Engineer (PySpark, SQL)
Attribute Details
Role Data Engineer
Mandatory Skills PySpark, Advanced SQL
Experience Mid-Level to Senior (Typically 3+ to 8+
years)
Location(s) - Remote
Joiner Status Immediate Joiner (or candidates with a
notice period of 15 days or less)
Job Type Full-Time
🎯 Role Overview
We are seeking a highly skilled and motivated Data Engineer to join our
dynamic data platform team in either our Gurgaon or Chennai office. The
successful candidate will be a critical player in designing, building, and
optimizing our next-generation data architecture and pipelines.
This role requires expert-level proficiency in PySpark and SQL for building scalable, high-performance ETL/ELT processes that transform vast amounts of raw data into high-quality, actionable insights for analytics, reporting, and Machine Learning. We are looking for an immediate joiner who can hit the ground running and contribute significantly from day one.
🔑 Key Responsibilities
Data Pipeline Development & Optimization
Design and Build: Architect, develop, and maintain robust, scalable, and
fault-tolerant ETL/ELT pipelines for ingesting data from diverse sources
(e.g., databases, APIs, streaming sources) into our data lake and data
warehouse.
PySpark Expertise: Write and optimize complex data transformation jobs using PySpark and the Spark DataFrame API to process petabytes of structured and unstructured data efficiently.
SQL Mastery: Utilize Advanced SQL for complex querying, data manipulation, stored procedures, performance tuning, and optimizing database schema design in relational and analytical databases.
Data Quality & Governance: Implement data validation, cleansing, and monitoring routines to ensure high data quality, integrity, and adherence to security and governance standards.
Architecture and Infrastructure
Data Modeling: Design and implement optimal data models (e.g., Dimensional
Modeling, Data Vault, Snowflake Schema) for our data warehouse to support
business intelligence and analytical needs.
Cloud Integration: Work with cloud-native data services (e.g., AWS S3, Glue, EMR, Redshift, or Azure Data Lake, Databricks, Synapse, or Google BigQuery, Dataflow) to build and deploy data solutions.
Automation: Implement orchestration tools like Apache Airflow, Azure Data Factory, or AWS Step Functions to automate data workflows and manage pipeline dependencies.
Collaboration and Operational Excellence
Cross-Functional Teamwork: Collaborate closely with Data Scientists, Data
Analysts, Product Managers, and Business Stakeholders to understand data
requirements and translate them into technical specifications.
Monitoring & Support: Monitor, troubleshoot, and resolve issues in production data pipelines, ensuring maximum uptime and timely data delivery.
Best Practices: Participate in code reviews, enforce coding standards, and contribute to the continuous improvement of development and deployment practices (CI/CD, Git).
🛠️ Required Technical Skills (Mandatory)
PySpark: Expert-level, hands-on experience in developing and optimizing
large-scale data processing applications using PySpark (Python for Apache
Spark).
SQL: Mastery of Advanced SQL (including window functions, complex joins, stored procedures, and query performance tuning) across various database systems (e.g., PostgreSQL, MySQL, Snowflake, Redshift).
Programming: Strong proficiency in Python for scripting, automation, and general data manipulation libraries (e.g., Pandas).
Big Data: Solid understanding of Big Data concepts, distributed systems architecture, and data warehousing principles.
ETL/ELT: Proven experience in designing, building, and maintaining robust ETL/ELT pipelines.
🌟 Preferred Qualifications (Good to Have)
Experience with Databricks or other managed Spark environments.
Hands-on experience with a major Cloud Platform (AWS, Azure, or GCP) and its data-related services.
Familiarity with workflow orchestration tools like Apache Airflow or comparable technologies.
Experience with real-time/streaming data processing (e.g., Spark Structured Streaming, Kafka, or Kinesis).
Knowledge of Data Governance, Data Cataloging, and Data Security best practices.
👥 Candidate Profile
Educational Background: Bachelor’s or Master’s degree in Computer Science,
Engineering, or a related quantitative field.
Mindset: Proactive, self-motivated, and a strong sense of ownership and urgency.
Communication: Excellent verbal and written communication skills to articulate complex technical concepts to non-technical stakeholders.
Availability: Must be an Immediate Joiner or have a short notice period (15 days maximum).