Job Description: Data Engineer - Acquisition
As a core member of our data acquisition operations team, a data engineer is
responsible for maintaining crawlers and parsers and performing required fixes
on demand to ensure product data within the assigned geographical region is
successfully and consistently acquired daily.
An ideal candidate has a minimum of 3 years of hands-on programming experience
in Python, HTML, and JavaScript. The candidate has demonstrable knowledge in
web development or web scraping and is comfortable working with JSON/XML files
to extract data.
Key Responsibilities
-
You will design, develop, and maintain large-scale web scraping pipelines to
extract valuable platform data.
-
You will be responsible for implementing scalable and resilient data
extraction solutions, ensuring seamless data retrieval while working with
proxy management, anti-bot bypass techniques, and data parsing.
-
Optimizing scraping workflows for performance, reliability, and efficiency
will be a key part of your role.
-
Additionally, you will ensure that all extracted data maintains high quality
and integrity.
Our Expectations from the Right Candidate
-
Strong experience in Python and web scraping frameworks such as Scrapy,
Selenium, Playwright, or BeautifulSoup.
-
Knowledge of distributed web crawling architectures and job scheduling.
-
Familiarity with headless browsers, CAPTCHA-solving techniques, and proxy
management to handle dynamic web challenges.
-
Experience with data storage solutions, including SQL, and cloud storage.
- Understanding of big data technologies like Spark and Kafka (a plus).
-
Strong debugging skills to adapt to website structure changes and blockers.
-
A proactive, problem-solving mindset and ability to work effectively in a
team-driven environment.
```