Classification · Practical guide
How to plan website classification and article scoring automation
Define a rubric, extract useful page signals, and evaluate a sample before classifying domain lists or scoring research articles at scale.
Opening a long list of websites to categorize each business is a different task from collecting a fresh price feed. The output is a judgment under a rubric, so extraction quality and classification quality need separate checks.
I built a Website Classification Pipeline using a local Windows Python script. It extracted homepage titles, meta descriptions, the first H1, and short text snippets, then passed those signals to the ChatGPT API under client-provided instructions and exported results to CSV.
I also built an Article Scoring workflow in n8n that extracted URLs from Gmail, retrieved article content, scored it against a custom rubric, and logged qualified items to Google Sheets. These are related patterns with different source content and evaluation needs.
Choose categories or scores according to the decision
Categorization puts an item into a defined group. Scoring measures its fit against agreed criteria. Decide which output helps the next person act.
For a domain list, your team may need a business category and a short reason. For a research inbox, it may need relevance scores and the criteria that caused an article to be selected. Asking for both without a downstream use adds complexity to review.
Write category definitions or scoring conditions in plain language. Include positive examples, negative examples, and ambiguous cases. Decide whether more than one category is permitted and when “unknown” is the correct result.
Retain the evidence used for the judgment
The pipeline can classify only the signals it receives. A homepage with an empty description or a loading screen does not supply the same evidence as a complete business description.
Retain the source URL, extraction status, and useful page signals beside the result. An article workflow also needs to distinguish the article body from menus, recommendations, and surrounding page text.
An illustrative domain export could contain:
| Field | What it lets you review |
|---|---|
| input_id | The corresponding item in the original list |
| final_url | The source after any redirect |
| page_title | A captured business signal |
| extracted_text | The content supplied to classification |
| category | The result under the agreed definitions |
| reason | An explanation tied to the retrieved content |
| review_status | Unreachable, ambiguous, or ready for use |
Do not label an inaccessible page as an ordinary negative result. Failed extraction and a valid non-match should remain distinct.
Evaluate against a reviewed sample
Select a representative set your team labels manually. Include clear examples, sparse homepages, overlapping categories, and URLs that cannot be retrieved.
Compare workflow results with those labels. Review false positives and false negatives for each important category, then decide whether the problem is missing content, an unclear definition, or an interpretation error.
Keep the evaluation set separate from the examples used to tune the rubric where possible. Repeatedly optimizing for the same examples can give a misleading impression of quality on the full batch.
Plan a resumable batch
Record processing status per input item so a stopped run can resume. Keep the extraction and classification results associated with the correct input ID even if items finish in a different order.
Define retry behavior for inaccessible pages and failed model calls. Agree on limits and costs after the pilot reveals how much content and processing each item actually needs.
For a recurring workflow, record which rubric version produced a result. A later category definition change should be visible when your team compares old and new outputs.
Deliver a list with a review path
The final output should help the team act: accepted classifications, items needing review, and failures with an explanation. Setup instructions should show how to rerun selected items and update the rubric.
Website & Content Classification covers these batch and workflow patterns. If you need broader company research and qualification, see Lead Enrichment & Scoring. Bring a sample URL list and your category definitions to scope an evaluation batch.
Start with one workflow.
Tell me the repetitive task and the result you need. We’ll define a small pilot, check it against real examples, and agree on the full scope.
Keep reading.
Lead research
How to scope lead enrichment and scoring in Google Sheets
Turn an existing company list into source-backed research and explainable qualification scores with clear field rules and a review process.
Read the guideEmail workflows
Designing an n8n email classification workflow for Gmail
Define message categories, duplicate handling, review queues, and structured logging before automating a shared inbox with n8n.
Read the guideWebsite data
What to define before scheduling website data collection
Choose the sources, fields, update schedule, duplicate rules, and delivery format that make a recurring dataset useful to your team.
Read the guide