Python Data Extraction · Project
Website Classification Pipeline
Built a lightweight Python pipeline that scrapes 5,000 domain homepages, extracts key page signals, and classifies each website with ChatGPT into CSV-ready output.
The implementation.
Delivered a clean Python script for a one-time local run on Windows to process a large list of domains at speed. The script visits each homepage, extracts core content signals including page title, meta description, first H1, and a short text snippet, then sends the structured payload to the ChatGPT API using client-provided classification instructions. Final classifications and scraped fields are exported to CSV for immediate analysis. The implementation uses requests, BeautifulSoup, and lightweight concurrency to improve throughput while keeping setup simple.
My contribution
Local Windows script, website signal extraction, API classification, and CSV output.

Start with one workflow.
Tell me the repetitive task and the result you need. We’ll define a small pilot, check it against real examples, and agree on the full scope.