The client needed to aggregate millions of vehicle data points across global markets in real-time to find pricing arbitrages. However, their legacy scrapers were constantly blocked by enterprise anti-bot systems (like Cloudflare) and IP bans. The Solution: I designed a distributed, highly scalable data extraction pipeline utilizing rotating proxy clusters, headless browser automation, and Kafka event streaming to feed a centralized Elasticsearch data lake. The Impact: Successfully bypassed strict enterprise-grade anti-bot protections, scaling data ingestion to millions of records per day with 99.9% uptime, unlocking massive revenue through real-time market arbitrage.