Lead scraping tools differ less in their features than in their data source, and everything else hangs off that source: what you are contractually allowed to do with the data, how good it is, and who answers for it. Three groups cover the field. First, tools built on official interfaces and open registers: access is regulated, and in exchange the provider dictates what you may store. Second, bought contact databases: you get coverage and speed, but you take on a provenance you have not checked yourself. Third, tools that harvest platforms automatically: the terms of use of LinkedIn and XING forbid exactly that, and what gets restricted, in case of doubt, is your account, not the tool. Running across all three is a duty that appears in almost no tool list: anyone who collects personal data from somewhere other than the person themselves has to inform them under Article 14 GDPR about the purpose, the legal basis and the source, unless one of the exemptions in paragraph 5 applies. So the choice is not decided by feature count but by one question: can you say, for every row, where it came from?
I find my own clients with an AI employee who works squarely in the first group and deliberately leaves the rest alone. Which B2B sources exist at all is covered in Finding B2B Leads: LinkedIn and Alternatives; this post is about the tools in front of them. I am not a lawyer, and this is not legal advice, it is a sorting of rules you can look up yourself.
The three groups at a glance
The order is not a ranking. It sorts by what you have in your hand if things go wrong.
Group 1: official interfaces and open registers
The advantage is obvious: you use the route the provider built for you. The price is their rules, and those are tighter than most people expect. The Google Maps Platform Terms of Service, last modified 26 August 2026, say under "No Scraping" that the customer must not "export, extract, or otherwise scrape … for use outside the Services", and give as an express example not to "copy and save business names, addresses, or user reviews". What you may store permanently is essentially the identifier of a place, the place_id; everything else falls under the caching prohibition with the narrow exemptions in the service-specific terms. So anyone who builds their own company database through a maps interface is in breach of the terms of the service, even though the calls themselves were paid for and authorized.
Registers are the other half of this group. The German commercial register is open to anyone for information purposes, but expressly "durch einzelne Abrufe", through individual retrievals, which is precisely not a bulk download; the notices of public contracting authorities are freely available and come with an open-data interface. Neither costs a licence fee, and in both cases the permitted route is spelled out. What you do not get: a contact person, an email address, a reason to call. The register sources are set out at length in the B2B post.
Rule of thumb for this group: the interface supplies the candidate, you do the assessment yourself, and you take the contact details from where the business published them itself.
Group 2: bought contact databases
Here you are not buying a tool, you are buying a stock of data. Providers in this segment advertise precisely that: that they have handled the sourcing, legal questions included, for you. One example of the tone in which that happens: on its own compliance page, Cognism writes that it notifies business contacts of their inclusion in its database "within GDPR timeframes" and regularly scrubs its data against "15 major Do Not Call and TPS lists". That is a vendor statement about the vendor's own holdings, not a verified assurance about your later use of them, and that is exactly how it should be read.
Because buying does not shift one thing: the moment you put a list into your own system and work with it, you are processing the data yourself. The legal basis for your processing, typically legitimate interests under Article 6 (1) (f) GDPR, is something you have to assess for yourself. What a provider has settled for its own holdings is an argument for that, not a substitute.
Substantively you get coverage, and that is the real advantage when your target group is large and scattered. What is hard to verify is the individual row: when it was last correct, which source it came from, and whether the email address was found or derived. Ask the provider about exactly that, field by field rather than in general. If the answer to the provenance question is nothing but marketing, that tells you something too.
Group 3: tools that harvest platforms automatically
This is the group that coined the search term, and the one the tool lists go quietest about. LinkedIn's User Agreement, in effect since 3 November 2025, lists among the prohibited acts that you must not
"Develop, support or use software, devices, scripts, robots or any other means or processes (such as crawlers, browser plugins and add-ons or any other technology) to scrape or copy the Services, including profiles and other data from the Services"
and equally must not
"Use bots or other unauthorized automated methods to access the Services, add or download contacts, send or redirect messages"
Three things about that decide your tool choice. First, browser extensions are named word for word, so the most convenient build is not the more harmless one. Second, LinkedIn's own help pages name the consequence: people using such tools "risk having their accounts restricted or shut down", and the tools themselves may "become non-operational without notice". The risk sits with your account and your pipeline, not with the tool's vendor. Third, this is not only LinkedIn: XING likewise prohibits the use of mechanisms, software or scripts in connection with using its websites, and both wordings are quoted in the B2B post.
On top of that comes a right that is often overlooked because it has nothing to do with data protection. Under section 87b (1) of the German Copyright Act (UrhG) the maker of a database has the exclusive right to reproduce it in whole or in a part that is "nach Art oder Umfang wesentlichen", substantial in nature or extent, and the "wiederholte und systematische" repeated and systematic extraction of insubstantial parts is treated the same way where it conflicts with normal exploitation. A database in that sense is any systematically arranged collection whose procurement required a substantial investment (section 87a UrhG). Trade directories and portals can fall under it. Whether a given collection reaches that threshold, and whether a given extraction is substantial, is decided case by case, and that is a job for a lawyer, not for a tool list.
So there is no how-to here. The practical consequence is banal anyway: reading and researching by hand is the use these platforms were built for. Harvesting the surface automatically is not. And whether you may then write to the person you found that way is a second, separate question: for the contact route section 7 UWG applies on top, which prohibits advertising where it is evident that the market participant addressed does not want it. The detail is in the B2B post.
What "AI" usually means in these tools
For about two years now nearly every vendor has carried "AI" in the name of its features. Behind the word there are usually three things: a search that turns free text into filters, an enrichment that pulls further fields in from a company name, and a scoring that ranks candidates by fit. The first is the genuinely useful one, because it saves time.
One point deserves suspicion, whoever the vendor is: derived email addresses. When a tool builds an address out of first name, last name and domain, that is a guess with a hit rate, not a find. For the information duty it is awkward, because there is nothing you can say about the provenance except "calculated". And for a first approach it is the worst possible footing. My own rule on this is hard and has served me well: without a documented address from the imprint or the contact page, there is no lead.
The duty you buy along with the tool
Once a natural person behind a row is identifiable and the detail was not collected from that person, Article 14 GDPR applies. A plain company address with no link to a person does not automatically fall under it; a name in the contact-person column does, and in all three groups that column is exactly what the tools are paid for. Among the mandatory items, paragraph 2 (f) includes "aus welcher Quelle die personenbezogenen Daten stammen und gegebenenfalls ob sie aus öffentlich zugänglichen Quellen stammen", from which source the personal data originate and, where applicable, whether they came from publicly accessible sources. Paragraph 3 names three points in time, of which the earliest counts: at the latest within one month of obtaining the data, at the latest at the time of the first communication where the data is used to communicate with the person, and at the latest at the first disclosure where disclosure is envisaged. Paragraph 5 sets out exemptions, among them where informing the person would involve disproportionate effort.
For the tool choice that means: a tool that does not hand you the provenance per row makes this duty harder. A tool that carries it along takes work off you that otherwise piles up at the end. Three columns per row are enough: where the detail came from, when you recorded it, when you informed the person.
Free: what that means in each group
"Free" means something different in each of the three groups. In group 1 it really does mean free: registers and tender notices cost no licence fee, the effort sits in the analysis. In group 2 it usually means a limited starter allowance out of the same bought holdings, with the same questions about provenance. In group 3 you pay with your account. A free route is most likely where the data is public to begin with.
How Olaf cuts this differently
Olaf is my AI employee for the groundwork, and he deliberately sits in group 1. He researches businesses through the Google Places API, checks every website he finds with a quick technical review covering mobile usability, HTTPS, the age of the site builder and whether the contact page is reachable, and takes screenshots of the site, plus mobile ones in the deep check, because he only judges a design with the picture in front of him. The contact address comes from the business's own imprint or contact page, and it is never guessed. Without it there is no A or B, there is a C. Companies with no website at all are not leads, because the material to judge them by is missing. And he contacts nobody: no emails, no calls. He finds, I approach.
That is not a better tool, it is a different cut: the group 3 tools are built for volume, Olaf for a short list with a reason per entry. From my lead runs in July 2026: 198 candidates checked across eight assignments, and in the seven location assignments with 20 candidates each that typically turned into 3 to 5 real leads; the eighth was a differently built community list. The frame for rebuilding that cut on your own criteria is in AI Employees for Consultants and Agencies, further roles on the same pattern in AI Agents: 10 Real-World Examples, and the line for client data in AI and Privacy: What the AI Gets to See.
Frequently asked questions
Which lead scraping tool is the best?
That depends on the data source, not on the feature list. If you need to be able to evidence every row, you work with official interfaces and registers. If you need coverage in a large market, the route runs through a bought database. Tools that harvest professional networks automatically bring the densest data and the highest risk to your own account.
Is lead scraping legal?
There is no blanket answer, because three levels have to be checked separately. The terms of use of the source govern whether the retrieval is contractually permitted. Database rights under section 87b UrhG govern the extraction of substantial parts of someone else's collection. The GDPR governs your processing. With professional networks the first level is already unambiguous: automated harvesting is prohibited.
Are there free lead tools worth using?
Yes, if you stay with public sources. The commercial register and the federal notice service for public contracts cost no licence and deliver dated triggers, and for both the permitted route is spelled out: the individual retrieval for the register, the open-data interface for the notice service. Inside that route you are on safe ground; a bulk download is a different matter. What they do not deliver is the contact person. Free allowances from commercial databases, by contrast, come out of the same holdings as the paid ones, with the same open questions about provenance.
What does AI actually contribute in a lead tool?
Most in the search, because it turns a description into filters, and in the pre-sorting by fit. Least in contact details: an email address derived from a name and a domain is a guess, not a source. And for the information duty under Article 14 GDPR it is precisely the source that you need.
Who is liable if bought data is not clean?
You are responsible for your own processing the moment you use the list. The provider answers for its holdings and for what it assures you contractually. A marketing claim on a product page is not such an assurance. So the question about provenance and legal basis belongs before the purchase and in the contract, not in the hoping.
How to continue
Before you compare tools, write down two sentences: which source makes your ideal client visible, and what trigger you can spot there. After that the choice of group almost makes itself, and the comparison shrinks to two or three candidates. Check three things on each: provenance per row, email addresses found or derived, and whose account carries the risk.
How I handed that cut to an AI employee is on the page about Olaf, my lead scout. The template to rebuild it on your own criteria, and the people currently building their own, are in my community Claude Practitioners.