What AI SEO Agents Still Get Wrong in Link Building
Author: Justin Davis, Link Builders
Originally published on LinkedIn.
I build these agents. I run them every week, against real client work, on real budgets. Over the past several months I’ve built three of them: one that mines competitor backlink profiles through the Ahrefs API, one that runs boolean-operator prospecting across Google SERPs, and one that finds a named contact and a verified email for every site on a sheet. That work is the continuation of a decade of doing all of it by hand, and it’s made my team roughly ten times more efficient than the manual process it replaced.
So when I tell you what they get wrong, understand where it’s coming from. This isn’t a skeptic’s takedown written by someone who tried ChatGPT once and didn’t like the output. This is a field report from someone whose margins depend on these things working, and who has spent a lot of hours watching them fail in ways that were expensive to discover.
Six failure modes. Every one of them cost me something before I learned to design around it.

1. The site that passes every filter and is still worthless
An agent evaluates a link prospect on measurable signals. Domain Rating, estimated organic traffic, referring domain count, whether the site has a resources page, whether it’s published recently. Those are the things you can query, so those are the things a filter can act on.
The problem is that every one of those signals is also the specification sheet for building a fake site. Private blog networks, rebuilt expired domains, and link farms are engineered to pass exactly the checks an automated system runs. They have DR. They have traffic numbers. They have a links page. That’s the whole point of them. The distinction between a backlink profile that accumulated naturally and one that was assembled is obvious to a practiced eye and nearly invisible to a query.
Early in my first build, I set the qualification band at DR 10-70 with a minimum of 100 estimated monthly organic visits. Reasonable thresholds. What came back included a lot of sites that scored perfectly well and that I could tell were junk in about four seconds of looking at them. The tell isn’t in any field the API returns. It’s the way the homepage is laid out. It’s an “About” page that describes no actual person. It’s twelve unrelated verticals in the same sidebar. It’s the specific visual flatness of a site built to hold links rather than to be read.
I don’t have a clean technical fix for this. I’ve gotten a long way by training the system on the quality signals and spam indicators I learned to recognize manually: consistent authorship, a real publication schedule, thin content, footer widget links. But “a long way” isn’t “all the way,” and so the rule stands: no site enters a client campaign without a human having looked at it. The agent’s job is to take tens of thousands of candidates down to a few hundred worth looking at. The looking is still mine.

2. Relevance that’s real on paper and fake in practice
This one cost me an entire first deliverable.
On an early run I pointed the agent at five competitors in a geographically specific niche and told it to surface relevant prospects. Mining a competitor’s backlink profile is the single highest-yield prospecting method I know, and the agent did it fast. It came back with a list sorted by Domain Rating, and near the top were Google, the New York Times, and Reddit.
Those sites are, strictly speaking, relevant. They have pages on the topic. They link to competitors. Every criterion I’d written was satisfied. They’re also completely non-actionable as link prospects, and any human who has done this work for a week knows that instantly.
The deeper version of the same problem is subtler and more dangerous. An agent judging relevance by semantic similarity will happily surface a site that covers the right subject for entirely the wrong audience. Topically adjacent, practically useless. A general travel publication and a site about one specific city are semantically close and commercially unrelated.
What fixed it wasn’t a better model. It was me writing down the geographic and topical keyword set explicitly (the specific city names, region names, and landmarks that had to appear in a page title or URL for the page to count) and having the agent hard-filter on that before any judgment call happened. This is the same discipline that makes advanced boolean operators work: you are not asking a system to understand your niche, you are describing your niche in terms it can execute against. Relevance stopped being something the agent decided and became something I defined and the agent applied.
That’s the general shape of the lesson. Agents are good at applying a rule you wrote precisely. They are unreliable at inventing the rule.

3. Outreach that reads like outreach
I’ll be blunt about this one because it’s the section most people in my industry don’t want to hear.
I read something like a thousand backlink pitches a week, and I’ve written up what that inbox actually looks like. Roughly 2% of what arrives is worth a reply. The reason that number keeps falling isn’t that editors got busier. It’s that a large share of the outreach hitting their inbox is now generated by the same handful of models, trained on the same corpus of “here’s how to write a great outreach email” blog posts, producing the same structure every time. The warm opener referencing a recent post. The compliment. The pivot. The soft ask. The “happy to send over a few ideas.”
When every agent writes that email, the format itself becomes a spam signal. Editors don’t consciously think “this is AI.” They pattern-match on the shape and delete. And this is happening in a search environment where Google is visibly sharpening its treatment of machine-generated, self-promotional content. So the downside isn’t only a dead campaign.
So in my own pipeline, the agents don’t write outreach copy from scratch. I supply fixed templates per link type, and the agent fills in a personalization layer and nothing more. That’s a deliberate downgrade in what I let the automation do, and I made it on purpose. The volume advantage of machine-written outreach is worthless if the volume converts at zero.
If you take one thing from this article: automating prospecting is a force multiplier. Automating outreach copy at scale is a way to burn your sender reputation faster than you used to be able to. That distinction is most of what building links in an AI-saturated environment now comes down to.

4. No read on the person at the other end
Related, but distinct. Agents have no model of who they’re talking to.
A blogger who launched eight months ago and has never been pitched is a completely different conversation from a managing editor at an established publication who receives two hundred pitches a week. The first one needs reassurance that you’re legitimate. The second needs you to get to the point in one sentence and prove you’ve read the site. Same email to both is wrong for both.
A human prospector reads that from the site in a few seconds: the publishing cadence, the tone, whether there’s an advertising page, whether the contact page has the weary specificity of someone who’s been burned. My contact-finding agent can tell you the person’s name, their role, and their email with good accuracy. It cannot tell you which of those two people you’re about to email.
It also can’t tell you which link acquisition tactic fits the site in front of it. Resource page, guest post, unlinked mention, and sponsored placement each call for a different opening line, and choosing between them is a judgment about the publisher, not a property of the domain. This is why my contact priority ladder is a human-authored rule rather than an agent judgment: editor, then content manager, then marketing, then owner, then generic inbox. The agent executes the ladder. It didn’t derive it and it couldn’t.

5. Confident wrongness
This is the behavioral problem underneath all the others, and it’s the one that makes unsupervised operation genuinely dangerous.
An agent produces a justification for every decision it makes. That justification is equally fluent, equally reasonable-sounding, and equally well-structured whether the decision was correct or catastrophic. There’s no tremor in the voice. Nothing in the output tells you which rows to check.
Two concrete cases from my own builds.
First: on an early run the agent delivered a clean, well-formatted, confident-looking list of qualified domains. It looked finished. It wasn’t. It had worked at the domain level when the actual opportunity lives at the page level, and it had sampled rather than swept a backlink profile that ran past 18,000 links in the target DR band for a single competitor. The output looked exactly as authoritative as a complete run would have. I only caught it because the result felt thin against what I know that niche contains. Anyone running the same competitor analysis through a chat window rather than a controlled pipeline is exposed to the identical failure and has even less ability to detect it.
Second, and worse in principle: an agent asked to find email addresses will, unless you forbid it, construct them. It knows the person’s name, it knows the domain, it knows that firstname@domain is a common pattern, and it will produce a plausible address and present it in the same column as the ones it actually found on a page. Nothing in the spreadsheet distinguishes them.
My standing rule now is absolute: no constructed emails, ever. Only addresses actually found on a page or returned by a lookup API go in the sheet. If nothing was found, the cell is empty and the row gets a contact form URL instead. An empty cell is honest. A guessed address damages your deliverability and your client’s reputation, and you won’t know which ones were guesses.
The design principle I’d give anyone building this: assume every output is confidently wrong until you’ve built a way to check it. My prospects hit roughly a 75% placement rate, which matches what I was achieving manually. But that number exists because verification is a budgeted cost inside the system, not a nice-to-have bolted on at the end.

6. No institutional memory
Link building compounds. You placed something on a site eight months ago, the editor knows you, the next placement takes one email instead of six. Across a decade and a hundred and fifty clients, that accumulated relationship graph is most of what an established agency is actually selling. And it’s the part of the scaling story nobody puts in the pitch deck.
An agent starts every run from zero. It doesn’t know you’ve already worked with a site. It doesn’t know an editor asked you not to pitch again. It doesn’t know a domain went bad last year. Without explicit suppression lists it will cheerfully surface the same prospects run after run, and it will pitch a site your colleague pitched last Tuesday.
Dedupe is not a nice feature. It’s load-bearing infrastructure, and it has to compare against every prior batch, not just the current one. The first time an agent re-pitches a contact who already told you no, you’ve spent credibility that took a year to build. Re-establishing it is exactly the kind of slow, unglamorous work that clients don’t picture when they imagine what a campaign looks like.
There’s a second-order version of this that I think is underrated. The relationship knowledge in an agency’s head is exactly the asset automation can’t copy. Which is, if you run an agency, reasonably good news. It’s also why I’m comfortable writing publicly about what my own tooling does. The software was never the moat.

What this means if you’re buying
Most people reading this aren’t building agents. They’re deciding whether to hire someone who claims to have them. So here’s the checklist: six questions, and the answers tell you more than any case study.
- Where do your prospects actually come from? “AI-powered” is not an answer. Competitor backlink mining, boolean SERP prospecting, and a purchased list are three completely different things with three completely different quality profiles. Any vendor worth hiring can walk you through their process step by step without hedging.
- Who looks at a site before it enters my campaign, and what are they looking for? If the answer is “our system scores it,” ask what the score is made of. If it’s made of DR and traffic, see failure mode one. Ask them to explain what kind of link they’re actually buying you on each placement.
- Is the outreach copy generated per-recipient by a model? If yes, ask to see five real examples from the last month. You’ll know within thirty seconds whether an editor would delete them. Most of the industry will not pass this test.
- How do you know an email address is real? The correct answer involves finding it on a page or verifying it through a lookup service. If any part of the answer involves patterns or formats, walk.
- What happens to placements that fail verification? Links get removed, go nofollow, or quietly never go live. Ask whether anyone checks after the invoice, whether you get credited, and what documented results they can show you from a comparable client.
- What’s the ratio of human oversight to output volume? Not whether there’s a human in the loop. Everyone says yes. How many hours of review per hundred prospects. A vendor who’s thought about this has a number. A vendor who hasn’t will change the subject.
One more thing worth knowing before you get quoted a price: automation lowered the cost of finding prospects. It did not lower the cost of a good placement, and anyone telling you otherwise is selling you something cheaper than what you think you’re buying. If you want to sanity-check a proposal, run the numbers yourself and read why the $500 link and the $100 link are not the same product.
There’s also a path that sits between hiring an agency and building agents yourself: buying placements from a vetted catalog. If you already know which pages you want to push and you want control over the sites, a backlink marketplace lets you browse niche sites by DR, traffic, and price, pay once with the article included, and track each link until it goes live. That does not replace judgment about which pages deserve links. It removes the prospecting and outreach stack when you want placements on sites that have already been reviewed by hand.
The honest position
The agents made me dramatically faster. Prospecting that used to consume four or five hours a day now runs in about one, and it sweeps parts of a backlink profile I’d never have had the patience to reach manually. That’s real and I’m not walking it back.
What they didn’t do is make me omniscient. Every one of the six failures above is a place where the system produces output that looks exactly like good output and isn’t. The oversight didn’t disappear when I automated. It moved: out of execution and into judgment, quality control, and designing the checks that catch the confident mistakes.
Anyone selling you automated link building without that second half is selling you volume. Volume was never the hard part.
If you run an agency and want this capacity without building it yourself, that’s what our white label program and agency outsourcing are for. Tell us about the campaign and we’ll tell you honestly whether we’re a fit.
Recent Posts
- What AI SEO Agents Still Get Wrong in Link Building
- How AI Outreach QA Gates Protect Link Building Quality
- I Rebuilt My Prospecting Process as an AI Agent — Here’s the Full Workflow
- How AI Personalization Layers Improve Link Building Outreach
- Is Google Targeting LLM-Focused Self-Promotional Content?
