Loading open roles
Loading open roles
Loading role

Vrinda Global · posted 1 month ago
Key Responsibilities
• Design, build, and maintain scalable, reliable ETL/ELT data pipelines across
cloud and on-prem sources,
ensuring data quality, lineage, and auditability.
• Develop and optimize Python/ PySpark and SQL-based data transformations for
large-scale, high-volume
financial datasets.
• Architect and manage data pipeline orchestration (e.g., Airflow, Databricks
Workflows, Step Functions) to
automate ingestion, transformation, and delivery.
• Build and maintain CI/CD pipelines using GitHub/GitHub Actions to support
automated testing,
deployment, and version-controlled infrastructure changes.
• Develop cloud-based solutions on AWS (S3, Glue, EMR, Redshift, Lambda, IAM)
supporting analytics,
reporting, and downstream ML use cases.
• Deploy and manage infrastructure and pipelines as code, following best
practices for environment
promotion, rollback, and monitoring.
• Monitor, troubleshoot, and optimize pipeline performance, query efficiency,
and cost across the data
stack.
• Partner with data scientists, analysts, product, and risk/compliance teams
to translate business
requirements into robust data solutions.
• Enforce data governance, security, and regulatory compliance standards
appropriate for financial data (PII,
SOX, PCI, etc.).
• Document pipeline architecture, data models, and processes; contribute to
engineering standards and
code review practices.
Required Technical Skills
• Advanced proficiency in Python for scripting, automation, and data
engineering workflows.
• Strong hands-on experience with PySpark for distributed data processing at
scale.
• Expert-level SQL and Advanced SQL (window functions, query optimization,
complex joins, performance
tuning).
• Solid experience with AWS cloud services and cloud-based application/data
development (S3, Glue, EMR,
Redshift, Lambda, IAM, CloudWatch).
• Proven expertise building and orchestrating data pipelines (Airflow,
Databricks Workflows, Step Functions,
or equivalent).
• Hands-on CI/CD experience using GitHub / GitHub Actions for automated build,
test, and deployment.
• Deep understanding of ETL/ELT design patterns, data modeling, and data
warehousing concepts.
• Experience deploying infrastructure and pipelines via code (e.g.
version-controlled deployments).
• Demonstrated ability to optimize pipeline performance, query execution, and
cloud resource/cost
efficiency.
Preferred / Desired Skills (Nice to Have)
• Hands-on experience with Databricks (Delta Lake, Unity Catalog, notebooks,
cluster optimization).
• Familiarity with Terraform or CloudFormation for infrastructure as code.
• Experience with streaming data technologies (Kafka, Kinesis, Spark
Structured Streaming).
• Exposure to data quality/testing frameworks (Great Expectations, Dbt tests).
• Knowledge of Dbt for transformation and analytics engineering workflows.
• Understanding of financial data domains — payments, lending, risk, fraud, or
accounting data.
• Relevant certifications (AWS Certified Data Analytics/Solutions Architect,
Databricks Certified Data
Engineer).
Qualifications
• Bachelor’s degree in computer science, Engineering, Data Science, or a
related field (or equivalent practical
experience).
• 5+ years of experience in data engineering, analytics engineering, or a
related technical role.
• Prior experience working within banking, fintech, or financial services,
with awareness of regulatory and
data-security requirements.
• Demonstrated track record delivering production-grade data pipelines in a
cloud environment.