Data Scientist (2+ Experience)
Role Overview
We are looking for Data Scientists with strong foundations in mathematics,
statistics, and scalable data engineering, along with hands-on experience in
Python, PySpark, SQL, and ML lifecycle tools like MLflow.
Key Responsibilities
-
Build and evaluate machine learning models
- Perform EDA and data preprocessing
-
Develop scalable pipelines using Python & PySpark
-
Track experiments and manage models using MLflow
-
Write optimized SQL queries for data extraction and transformation
-
Apply statistical methods for insights and validation
Technical Skills (Must-Have)
Programming & Tools
-
Strong Python (Pandas, NumPy, Scikit-learn)
- Working knowledge of PySpark
- Strong SQL skills
- Familiarity with MLflow:
-
Experiment tracking (logging params, metrics)
- Model versioning
- Basic model registry usage
Mathematics & Statistics
-
Hypothesis testing, probability distributions
- Descriptive & inferential statistics
Machine Learning
- Regression, classification, clustering
- Model evaluation metrics
Good to Have
-
Experience with visualization tools (Matplotlib, Seaborn, Tableau)
-
Basic knowledge of cloud platforms (AWS/GCP/Azure)
- Exposure to NLP or Deep Learning
Education
Bachelor’s or Master’s in Data Science or related field
Key Traits
-
Strong analytical thinking with a data-driven approach to problem-solving
-
Curious and eager to learn new tools, techniques, and methodologies
-
Good communication skills to explain insights to non-technical stakeholders
-
Collaborative team player with a proactive attitude
-
Detail-oriented with focus on data quality and accuracy
-
Takes ownership and delivers tasks effectively within timelines