I run an SEO agent on my own site. Not a sandbox, not a demo account. It reads a query report, watches a fixed prompt set across four assistants every week, and hands me a list of what changed since the last run. It has done seven jobs well enough that I would not go back to doing them by hand, and it has three limits I have stopped expecting it to cross.
An SEO agent is software that reads your search and AI visibility data, decides what needs doing, and runs that work on a schedule instead of waiting for you to ask. It is strong at repetitive measurement work and weak at judgement. Most of the disappointment people report comes from asking it for the second thing.
The honest scoreboard: an SEO agent is excellent at watching and reporting, and unreliable at deciding. It cannot see which question triggered an AI answer, cannot count every AI visit, and cannot tell a person from a machine inside your own data. Buy or build it for the first half, and keep the second half human.
That split is the whole argument of this post. Here is what each half looks like in practice, with the numbers from my own runs.
What an SEO agent actually is
The term gets used for three different things, and only one of them is worth paying for.
A rank tracker is a dashboard. It shows you positions and waits. A visibility tool measures one surface, usually whether your brand appears in AI answers, and it also waits. A chat assistant answers a question and forgets it existed the moment you close the tab.
An agent holds state. It remembers what it found last time, it runs on a schedule you set once, and when it comes back it checks its own earlier findings against fresh data. That memory is the difference between a report and a process. A report is a snapshot you have to interpret. A process tells you what changed and what it wants to do about it.
The practical test I use: does it come back on its own, and does it remember? If both are no, you have a dashboard with a new name on the box.
Quick question: does an SEO agent need access to my Search Console data? Yes, and that access is most of its value. An agent planning from generic keyword lists produces generic pages. An agent reading your own queries finds the pages that already get impressions and no clicks, which is the cheapest work available.
The 7 jobs an SEO agent does better than a person
These are the ones I would hand over without hesitation.
1. It audits on a schedule you do not have to remember
Manual audits happen when someone thinks of them. Agent audits happen on the calendar. On my own site the quarterly audit turned into a weekly one purely because nobody had to remember it.
People search for this as a geo audit (about 140 monthly searches in the US, keyword difficulty 0), which tells you how thin the competition is for a job that most teams are still doing by hand.
2. It watches a fixed prompt set across every assistant
Here is the number that made me take this seriously. My prompt set is 31 questions, and every run puts each one to four assistants. That is 124 checks per cycle, and the cycle runs weekly.
A person can hold four assistants in their head for one afternoon. Nobody does it 124 times a week, and nobody records the answers in a comparable format. That is the entire case for an agent in a single sentence, and it is why prompt monitoring (roughly 90 searches a month, difficulty 0) is one of the cheapest positions in this category to own.

3. It tracks AI Mode and AI Overviews presence
Google now reports the impressions your pages earn inside its AI features. It is impressions only, with no clicks attached, and the agent is a better reader of that report than I am because it sees the same page list every cycle and notices when one drops out.
The search phrase here, ai mode tracking, sits near 70 monthly searches at difficulty 11. Small numbers, and the reason I care is competitive rather than commercial: almost nobody has built the habit of watching it yet.
4. It finds pages that earn impressions and no clicks
This is where my own data got uncomfortable. Over 28 days, my query report listed 375 queries and 2,311 impressions. 363 of those 375 queries, or 96.8%, earned exactly zero clicks.
One question alone pulled 120 impressions at an average position of 4.7 and produced nothing. Ranking fourth and being chosen zero times is not a ranking problem, and a person scanning a chart page by page will not find that pattern. An agent sorting by impressions and filtering for zero clicks finds it in one pass.
5. It compares your citation share per prompt
Per prompt, not per site. The question I want answered is narrow: when someone asks this specific question, does the answer name me, or does it name someone else? Aggregate visibility scores hide the split. Prompt-level counts expose it.
6. It writes the fix list, not the fix
The agent drafts what should change and why, ranked by what the data says is worth doing first. It does not publish, and I would not let it. Writing the list is mechanical. Choosing which item to act on this week is the part where I still want a person in the loop.
7. It re-runs and proves whether the change worked
Anyone can change a title tag. Very few teams go back 30 days later, re-run the same checks, and compare like for like. An agent does it because the schedule already exists, and that comparison is the only thing that separates a fix from a guess.
Where these jobs have real demand, if you want to see what people actually type:
| Job | Search phrase | US monthly | Difficulty |
|---|---|---|---|
| Audits | geo audit | 140 | 0 |
| Monitoring | prompt monitoring | 90 | 0 |
| Tool selection | best llm optimization tools for ai visibility | 210 | 5 |
| AI Mode | ai mode tracking | 70 | 11 |
| Category | aeo tools | 720 | 18 |
| Comparison | generative engine optimization tool | 480 | 24 |
Volumes are US monthly figures from Google Ads keyword data reported in a July 2026 category scan. For scale, the head term this post targets, seo agent, carries roughly 27,100 monthly searches combined with its reversed form.
The 3 hard limits of an SEO agent
I have watched all three of these closely. None of them are prompt-engineering problems you can talk your way out of.
1. It cannot tell you which question triggered the AI answer
The report that shows your pages inside Google’s AI features gives you impressions and nothing else. No queries, no click-through rate, no history before the feature launched.
So when the agent says this page lost 400 AI impressions, the honest next sentence is that nobody knows which questions produced them. You can infer. You cannot know. An agent that presents an inference as a fact is worse than one that admits the gap, and telling the difference is your job, not its job.
2. It cannot count all the AI traffic it is measuring
AI referrals arrive without a referrer header more often than they arrive with one, which means a large share of them are filed as direct traffic. On top of that, clicks from Google’s own AI answers are counted as ordinary organic search rather than separated out.
The result is a number that only ever moves in one direction: whatever your agent reports for AI traffic is a floor. Every estimate it produces is on the low side, and no amount of instrumentation inside your own site fixes that, because the information is missing before the request ever reaches you.
3. It cannot tell a person from a machine inside your own data
This is the limit that cost me the most time, and the one I would want a buyer to understand before signing anything.
My query report looked like it was telling me something about my audience. A handful of rows suggested demand in areas I had not written about, with position data and everything. So I read the raw rows instead of the summary. Three of them were not searches at all.
One was a ten-clause exclusion string, the kind of thing a person types into a tool rather than a search box. One was a test someone ran against our own tracking setup, with a date stamp still attached. One was an instruction to look up a named individual’s employer and return nothing else, an email address still stuck on the end of it.
Those rows sat in the same table as real demand, formatted identically, waiting to be treated as evidence.
Then the shape of the file confirmed it. 139 of the 375 queries, or 37.1%, ran to eight words or longer. A further 66 were one or two word fragments. Real searches cluster in the middle, and a file that is one-third long-form instruction is not a record of people looking for things.

Research on generative engine visibility found that adding quotations and statistics improved how often content was surfaced, by up to 40%, while keyword stuffing performed worse than doing nothing at all (Aggarwal et al., Generative Engine Optimization, KDD 2024).
Read that finding next to this limit and the shape of the problem is clear. The tactics that work are the ones a machine has to read carefully. The noise that misleads you is also machine generated. Your agent reads both with equal confidence, because nothing in the data marks which is which.
How to check an SEO agent’s work
Four habits cover most of the risk, and none of them are technical.
Fix the prompt set. If the questions change every cycle, the comparison is meaningless. My set has been stable long enough that a change in the answers means something changed in the world.
Re-run on a schedule and compare like for like. Same questions, same assistants, same window length. A weekly run compared against last week is signal. A run compared against an ad-hoc test from three months ago is noise wearing a chart.
Read the raw rows before you trust a trend. This is the lesson from the third limit, and it is the one I would put first. Any time a number surprises you, open the file and look at the individual entries. The summary hides the machine.
Track what a run costs, in units. Checks consumed per run, checks per fix, cost per cycle. Agents do not notice waste, so it has to be somebody’s job, and the only available candidate is you.
Quick question: how do I know a re-run is comparing like with like? Same prompt set, same assistants, same window length, same date boundaries. Change any one of those and you are comparing two different measurements and calling the difference a result.
Two tools I lean on for the outside picture: Google’s own documentation for AI features in Search, which is more conservative than most vendor posts about what any of us can actually measure.
What it costs to run one
I will not quote you a monthly figure, because the arithmetic depends entirely on your scale and I would be guessing.
What I can give you is the shape. Cost scales with checks per cycle, not with seats, and a check is one question put to one assistant. My own run is 124 checks a week. A larger site watching 100 prompts across four assistants is running 400 checks per cycle, which is roughly three times the work and roughly three times the cost.
That per-check model is also why prompt-set hygiene matters more than it sounds. A duplicated or malformed entry is not a one-off error, it is a recurring charge on every cycle from then on, and it will not announce itself. More on that in the AI agent budget control breakdown, which covers the three levels of spend and why pre-flight checks matter more than dashboards.
Set against an agency retainer, the comparison is straightforward. A retainer buys you someone’s hours. An agent buys you a fixed number of checks, and the checks happen whether it is a holiday or not.
Where visibility.so fits
We built the agent that runs these checks into visibility.so, because researching this by hand is what made us build it in the first place.
The product watches a prompt set across four assistants, tracks whether the brand shows up in each answer, and keeps the ranking and traffic data next to it so a change in one can be read against the others. It runs on a schedule, it remembers the last run, and it reports what changed. If you want to see the checks we apply to a single page before any of that, the answer readiness checks are the same standard applied in isolation.
The product is one part of the answer. The other part is the person deciding which fix to ship, which is why we describe it as an operating system for hybrid teams rather than a tool that replaces a person.
Challenges and what we do about them
The most common doubts I hear, and my honest position on each.
“My data is too messy for an agent to work with.” Then the agent surfaces the mess, which is worth knowing regardless. My own query file was one-third machine noise and I had no idea until I read the rows. Start with one project, read the first two runs by hand, and you will learn more about your data quality than a year of dashboards taught you.
“What if it makes a recommendation that costs me traffic?” Keep it out of the publishing path until it has earned trust. Mine writes lists and does not touch the site. Every change goes through a person, and the agent’s job is to arrive at that review with a ranked list instead of a 200-row export nobody reads.
“How do I know the recommendations are any good?” Re-run and compare, which is job seven on the list. A recommendation with no follow-up measurement is an opinion.
“Is this not just SEO with a new label?” For the measurement half, no. Traditional SEO tools did not have to answer whether a language model named your brand when someone asked it a question. That is a new surface with its own data, its own blind spots, and its own reporting quirks. The underlying craft, deciding what to publish and how to structure it, carried over almost intact.
Frequently Asked Questions
What is an SEO agent?
An SEO agent is software that reads your search and AI visibility data, decides what needs attention, and runs that work on a schedule rather than waiting for you to ask. The defining features are memory and cadence: it remembers previous findings and comes back on its own.
Do AI SEO agents actually work?
They work on measurement and repetition, and they struggle with judgement. If the work is running the same 124 checks every week and reporting what changed, an agent wins outright. If the work is deciding whether to rewrite a title or restructure a page, keep that with a person.
How much does an SEO agent cost to run?
It scales by checks per cycle rather than by user seats. One check is a single prompt put to a single assistant, so 100 prompts across four assistants is 400 checks per run. Price your expected check volume before comparing vendors, because two products at the same seat price can differ tenfold in work per cycle.
Can an SEO agent replace my SEO team?
No, and any vendor telling you otherwise is selling a subscription rather than a system. It replaces the remembering, the re-checking and the compiling. Your team still decides what to do and is still accountable for whether it worked.
What is the difference between an SEO agent and a GEO tool?
A GEO tool measures whether your brand appears in AI answers. An SEO agent does that and then keeps going: it holds a task list across runs, tracks rankings and traffic alongside visibility, and re-checks its own earlier findings. Most GEO tools stop at the report.
Why does an SEO agent sometimes report traffic that never happened?
Because agency-run crawlers, testing tools and automated agents generate entries that look exactly like user queries. In my own export three rows were machine generated, including a ten-clause exclusion string and a test run against our tracking setup. Long-form entries are the tell, and a file where more than a third of queries run to eight words or more deserves a closer read.
Is an SEO agent worth building instead of buying?
Build it if your requirements are narrow and you already have engineering time. Buy it if you want the schedule, the memory and the reporting without maintaining them. One project, not a multi-year roadmap.
What this data cannot tell you
The query file behind this post covers 28 days ending 28 September 2026 on one site, which is small and recent. The 96.8% zero-click figure is specific to a site with modest traffic, where most queries inevitably collect nothing. A larger site would show a friendlier number, and the pattern behind it would not change.
What holds regardless of scale is the shape of the problem. Long machine-generated entries sit alongside real demand. AI impressions arrive without the queries that produced them. Referred AI visits arrive without referrers. An agent will read all of it with perfect confidence unless somebody builds the checks that force it to hesitate.
The question about results over time, how quickly AI visibility work shows up, is the one I get most often after this, and the honest answer there is also a range rather than a promise: what determines how fast AI visibility work shows results covers the variables, and how AI search engines work explains the retrieval mechanics underneath it.
If you are earlier in the decision, the complete guide to agentic SEO covers how the operating model differs from hiring, and the agentic SEO services breakdown lists what to ask any vendor before you sign. I would read this post first, then those, in that order.