I need a lightweight, reliable Python scraper to extract contact data from Australian and New Zealand adult classifieds and escort directories. These sites are simpler than us/eu targets, with mobile numbers, whatsapp, and telegram handles usually publicly visible or hidden behind a single javascript click-to-reveal layer. The primary goal is B2B sales outreach, requiring clean, deduplicated, text-only CSV output. I am open to freelancer recommendations on which sites to prioritize based on traffic, data quality, and extraction ROI. Target Websites: The project should prioritize
.com.au and
.co.nz domains. International sites are acceptable only if they have verified, high-volume au/nz sections. Primary regional targets include ScarletBlue, EscortsAndBabes, RealGirls, NZGirls, KiwiEscorts, and Locanto (au/nz sections). International sites with a heavy au/nz presence include tryst, listcrawler (au/nz filters), and slixa (au/nz filters). SkipTheGames is excluded due to zero AU presence. The focus is strictly on active regional directories. Freelancers are encouraged to include other decent-sized au/nz platforms with independent operators in their proposals, with brief justifications. Technical Requirements: The scraper must be Python-based, utilizing frameworks like Playwright, Selenium, or Scrapy. It needs to effectively handle dynamic JavaScript click-to-reveal mechanisms for phone, WhatsApp, and Telegram fields. Implementation of rotating datacenter proxy support is essential; I will provide a proxy list or the freelancer can use their own for testing. Extraction must be text-only, with no images, media, or file downloads. The output should be a clean CSV file with specific columns: phone, whatsapp, telegram, location, ad_title, post_date, profile_url, extracted_at, and Verified. Records must be deduplicated, and phone/Telegram formats validated. The solution should be lightweight and efficient, designed to run smoothly on resource-constrained environments.
Delivery term: Not specified