
About Chirag
Chirag co-founded TheDataHQ and has spent over twelve years building web scrapers and data pipelines for businesses that need data off the public web. He works in Python, with Scrapy and Playwright, on the collection side of the problem - the part that has to keep working after a site changes its layout.
He leads the engagements: scoping what a client actually needs, deciding whether it can be collected without getting around a site’s access controls, and saying so early when it cannot. TheDataHQ has delivered more than 260 projects across 10,000+ websites, running upwards of a million pages a day at peak.
What he works on
- Python, Scrapy and Playwright
- Anti-bot and rate-limit handling
- Pipeline design and data quality
- Scoping and compliance review
Selected work
Engagements delivered by TheDataHQ, where Chirag was part of the team.
- Venture Capital
Global startup intelligence pipeline
We built an automated pipeline covering funding rounds, founder backgrounds, growth metrics and technology stacks, collected continuously and delivered through a REST API and a custom dashboard.
2.5M+ Startups tracked120+ Sources200GB+ Data collected
- Human Resources
Job market data at national scale
We built a continuous collection pipeline pulling employment listings from thousands of sources, normalised into one schema so roles, salaries and locations stayed comparable across markets.
50K+ Jobs per day15K+ Sources12 Countries
- E-Commerce
Competitor price monitoring
We automated price and product monitoring across the competitor sites they cared about, with daily pricing reports and real-time alerts when a tracked product changed.
200+ Competitor sitesDaily Price refresh
Writing by Chirag
Posts on web scraping, data extraction and what it costs to run properly.
Read the blog