One of the biggest problems in outbound list building is selection bias.
Attempt 1: Scraping "Best of" lists. Companies on these lists are there because they are good at getting listed. Award pages surface the marketing-loud tier: massive teams, enterprise clients, and heavy PR budgets. But the technical founder who is heads-down delivering client work is invisible there. And that was exactly who I was looking for.
Attempt 2: Scraping a tool partner directory. This is where practitioners list themselves to get technical client referrals. I ran the exact same ICP parameters, the exact same AI classifier, and the exact same grading rubric across 153 companies.
The Result: 33 qualified keeps. A 22% hit rate.
Same ICP. Same classifier. Eight times the density.
Why? Because the second source is a place my ICP self-identifies, instead of a place they campaign to appear.
The lesson I keep re-learning: list quality is decided before enrichment, before copy, and long before an AI agent touches a single record. It is not determined by how you filter. It is determined by where you fish.
Where does your ICP self-identify? Because that is your list.
What sourcing pools have outperformed for you? Directories, communities, certification pages, or something weirder?