Key Roles & Responsibilities
Job Summary:
We are looking for a Senior Python Data Engineer who can take ownership of our
large-scale data pipelines and make them run faster, leaner and more reliably.
The person in this role will spend most of their time working on real
production code, optimizing it end-to-end and bringing down pipeline runtimes
that today stretch into many hours.
This is primarily a Python role, not a distributed-systems role. Our pipelines
run on high-memory single-node servers (500GB+ RAM, 16+ CPUs), so the focus is
on writing efficient Python code, managing memory carefully, parallelizing
where it makes sense, and getting the most out of modern Python data
libraries. Existing pandas-heavy code needs to be re-engineered into faster
equivalents, and new pipelines need to be built with performance in mind from
day one. You will be working closely with the actuarial, MIS and reporting
teams.
The right person will be hands-on, comfortable reading and rewriting other
people’s code, and curious enough to figure out why one approach takes 27
hours and another takes 90 minutes. Output equivalence is non-negotiable –
whatever we ship has to match the original numbers exactly, every single time.
Key Responsibilities:
-
Own end-to-end optimization of existing Python pipelines – profile them,
identify the slow steps, and rewrite them so the same job that takes 20+
hours today finishes in a fraction of that time.
-
Re-engineer pandas-heavy code into faster, more memory-efficient
equivalents. Most of our existing scripts are written in pandas; we are
moving towards modern Python data libraries (Polars, DuckDB, PyArrow) where
they make sense.
-
Build new data pipelines for various data process and similar workloads –
from source extraction all the way to the final reporting layer.
-
Handle large-volume data extraction from Oracle and PostgreSQL using Python
connectors. Tune fetch parameters (arraysize, prefetchrows, chunking,
parallel reads) to keep ingestion fast and predictable.
-
Work with parquet and other columnar formats – partitioning strategies,
compression, projection pushdown – to keep intermediate storage and read
times under control.
-
Manage memory carefully on high-RAM single-node servers (500GB+). This
includes chunking large joins, releasing memory between steps, avoiding OOMs
on multi-hundred-million-row datasets, and using parallelism
(ThreadPoolExecutor, multiprocessing) where it actually helps.
-
Ensure 100% output equivalence between the old and new versions of every
pipeline. Build comparison and validation scripts so that nothing slips
through during a rewrite.
-
Write advanced SQL for both data validation and PostgreSQL output tables –
window functions, CTEs, joins on tens of millions of rows, indexing, and
basic query tuning.
-
Translate business logic from actuarial, MIS and reporting teams into clean,
testable Python code – reinsurance calculations, earned premium,
member-based reporting, portfolio cuts and so on.
-
Support monthly production runs – reconfiguring parameters, monitoring
runtimes, fixing data issues as they come up, and keeping the reporting
calendar on track.
-
Maintain proper version control, logging and elapsed-time reporting on every
script so that production behaviour is observable and reproducible.
Key Requirements – Education & Certificates
B.E. / B.Tech / M.Tech / MCA in Computer Science, Information Technology, or
any related discipline. Equivalent degrees in Statistics, Mathematics or
Engineering are also fine if backed by strong hands-on data engineering
experience.
Certifications in Python, SQL or any cloud data platform (AWS / Azure / GCP)
are good to have but not mandatory.
Key Requirements - Experience & Skills (Must have)
Must Have
-
5+ years of hands-on Python experience, with a strong portion of that spent
on data engineering / ETL work – not just scripting around dashboards.
-
Strong working knowledge of pandas – merges, group-bys, multi-index
handling, the usual NaN / null gotchas, and a clear sense of where pandas
hits its limits. Experience migrating pandas code to faster equivalents is a
big plus.
-
Solid SQL skills – should be comfortable writing and reading complex queries
(joins, CTEs, window functions, aggregations) on large tables.
-
Hands-on experience pulling data from Oracle and PostgreSQL using Python
connectors (cx_Oracle / oracledb / psycopg2 / asyncpg etc.). Should
understand fetch tuning, chunking and connection management for large
volumes.
-
Comfort with parquet and columnar file formats – reading, writing,
partitioning, compression options, and projection / predicate pushdown.
-
Real-world experience with memory optimization on single-node servers –
chunked processing, releasing memory between steps, debugging OOMs, and
using parallelism (ThreadPoolExecutor / multiprocessing) sensibly.
-
Good debugging instincts – should be able to take an existing, undocumented
Python script and figure out what it’s doing and where it’s slow.
-
Comfort working in a Jupyter notebook based environment for development, and
a habit of writing clean, version-controlled, well-logged code.
Good to Have
-
Exposure to Polars, DuckDB or PyArrow. We don’t expect everyone to know
these – if you do, it’s a strong plus; if you don’t, willingness to pick
them up quickly is enough.
-
Working knowledge of PySpark / Apache Spark. Useful for context and for any
future distributed work, but the day-to-day on this role is single-node
Python.
-
Prior experience in health insurance, BFSI or any other regulated industry.
Understanding of premium, claims, reinsurance or member-level reporting is a
definite advantage.
-
Familiarity with workflow / scheduling tools like Airflow and with Git-based
version control.
-
Some exposure to one of the major cloud platforms (AWS / Azure / GCP) and to
Excel-heavy reporting workflows used by business teams.