Claude can do a FAIR amount of SEO work. That makes it worth using. It also makes it very easy to give it more responsibility than the evidence justifies. However, having ran my own tests using AI for SEO, I can say with absolute certainty that it's not ready to replace SEO consultants, agencies or specialists just yet.
AI is great for specific tasks, even more complex tasks now that agents can effectively sub-delegate tasks, but the issue with AI overall is the unification between the work, analysing the outcomes whilst building brand and adjusting strategies in response to organic growth.
I know this FIRST HAND because I test AI SEO agents, from Openclaw to Hermes, from chatGPT6 Astra through to other agentic solutions, I test and test and test.
Give AI such as Claude a crawl export, Search Console data and a clear brief, and it can help you work through problems, organise findings and produce useful recommendations (to a degree). Connect it to the right tools via API/MCP and it can retrieve information, write code and support implementation - again, to a degree and not without limitations.
I have no interest in pretending those capabilities are trivial, If an agency’s entire contribution is a generic report and a handful of recycled recommendations, businesses are right to question what they are paying for.
A lot of what made SEO work in the past could be done with relatively "low effort" - so it wasn't hard to produce a crawl report, highlight broken or poorly/non optimised page elements and generate recommendations for implementation, the same for content, a lot of the SEO agency model was built on pushing out blogs and articles, and, because search was simpler, CTRs were higher it was relatively easy to get traffic.
But, SEO has changed considerably, not just from an effort perspective but from an opportunity perspective too - mainly because Google's incessant push on AI overviews and expanded paid ads amongst other things have made it considerably harder to get clicks.
Replacing an expensive process (SEO) is different from establishing that the replacement makes sound decisions. The part I want businesses to examine is what happens between an AI recommendation and a change to a website that generates revenue.
What evidence supports the recommendation? What information is missing? What could the change affect elsewhere? Who checks whether it worked?
Those questions matter considerably more than how impressive the answer looks in a chat window.
Can AI / AI agents really replace SEO agencies or SEO consultants? IMO no, no they can't and won't be able to for some time yet.

Where do I think that AI platforms such as Claude / chatGPT will struggle the most with taking the labour from skilled SEO specialists?

Crawling websites that are large, or even medium sized, where there are firewalls, crawl restrictions, complex setups
Being able to properly document and organise SEO issues in such a way that they are actionable and measurable
Being able to understand context of domain / sub domain / sub folder configurations for say multi-language sites
Being able to properly assess rendering
Being able to output recommendations that are actionable or can be distributed to the right people in the right way
Being able to properly assess all google search console data, especially larger sites
Being able to process larger data sets with 50k API request limits from GSC API
Being able to understand the relationship between page groups / page types / page clusters
Being able to properly audit content and again, understand content context, quality, coverage, NLP
Being able to measure the true quality of a link / domain / brand profile
Being able to measure the impact of implemented work
There's more, but these are the key ones.
Some of the other things to note and what I would consider:
Crawl issues and challenges can be partially negated with MCP providers such as AHREFS who provide crawl data over MCP - however then we are still relying on another 3rd party data set, and as anyone in SEO will tell you, third party tools aren't a catch all for SEO issues
When I used chatGPT6 Astra with AHREFS it did an audit

But as expected, truly underwhelming and missing a lot of key data required for on site SEO:


So, as I expected, a mediocre crawl / output that still relied on a third party tool anyway.
What the evidence actually shows about businesses using Claude/chatGPT for SEO
There are real commercial applications here, there are also promotional claims that deserve a much deeper look.
The examples below establish that Claude is being used for SEO and related marketing work, but they do not establish that an unattended Claude setup reliably outperforms a competent consultant or agency across an entire SEO programme - this is really key to understand, you can't just set it and forget it, it still needs someone to oversee the work, and, if you aren't an SEO well, good luck.

Seeing the example above, Cox Automotive is particularly useful because the operational detail matters more than the headline, this is a business combining AI with domain knowledge, connected data and review. Its case study says generated content is presented to users for validation and acceptance.
That is a credible way to improve production, it does not support the claim that expertise has become less useful or totally redundant.

Snapshot 1. Anthropic’s Cox Automotive case study explicitly retains validation and acceptance, captured 3rd October 2026.
I also found public posts claiming large savings from replacing SEO retainers with Claude (funnily enough none of these have any visible case studies) personally I would not build a business case around those figures without seeing the underlying data (I mean GSC data and not AHREFS snapshots that are vague or where the domain hasn't been disclosed.
1 search-indexed LinkedIn post included the qualification that the strategist stayed, while its full page could not be retrieved during this research. Its performance claims are therefore excluded from the evidence above.
This distinction is important. A founder doing the strategy, a writer creating the content and a developer implementing the changes is still a team doing SEO, even if Claude coordinates a lot of the work.
The first fallacy is treating a convincing answer as a verified finding
Anthropic’s own documentation acknowledges that Claude can generate factually incorrect information or information inconsistent with the context provided. It recommends grounding answers in sources, allowing uncertainty and checking claims. It also states that these techniques do not eliminate hallucinations entirely.
That should be enough to rule out treating every polished recommendation as an established fact.
An SEO hallucination does not have to be a ridiculous claim that anybody would spot. It could be an invented search-volume figure, an unsupported statement about a competitor, or an assertion that Google has changed a policy when the cited page says something much narrower.
Hallucinations can often lead to false or wrong information - now, consider that Google recently updated its policies on AI generated content we see >

The interesting thing here is the fact that Google is expecting people to fact check their content, and that AI content is okay if it's gone through a process of manual review and where the content output is not low effort (not just churning out endless content) but actually making the content something people can use.
So, imagine if you have Claude or chatGPT running your SEO completely unchecked, you'll end up with content that's exposed and ultimately
risk losing rankings/traffic.
The following are illustrative failure modes, not transcripts of Claude responses produced for this article:

A link at the end of an answer is not sufficient, open it, check that it exists, that it is current enough for the question, and that it supports the claim being made.
A human in the loop is needed.
I want recommendations that can survive inspection, a confident explanation is useful only when the evidence holds up.

Snapshot 2. Anthropic’s own guidance recommends uncertainty, factual grounding and citation checks, captured 3rd October 2026.
The second fallacy is assuming connected data means complete context
Connecting Claude to Search Console is useful, It does not give it a complete understanding of your business - plus the key thing here is that a lot of websites are going to have MORE data than can be pulled in a single request, even paginated requests, this can expose AI to a lack of full data meaning decisions COULD be made that are wrong because the data is sampled or incomplete.
Google’s Search Analytics API documentation says it does not guarantee every data row and prioritises top rows, the documented maximum is 25,000 rows per request, with pagination available. Google separately explains privacy filtering and daily data limits, these are source constraints, not evidence that Claude/chatGPT has failed.
The failure happens when a workflow ignores those constraints and presents a partial extraction as a comprehensive analysis - again, if you aren't an SEO, getting an AI platform to do SEO without knowing all the surrounding contexts is an accident waiting to happen.
A larger context window cannot restore data that was never retrieved (again, limitations), an unlimited allowance for tool requests cannot make a source expose information it withholds i.e. sampled data, anonymised data etc..

Snapshot 3. Google explicitly qualifies the completeness of Search Analytics API results.
Even an accurately retrieved dataset leaves business questions unanswered, Google Search Console does not tell you which service or product has the strongest margin, which enquiries sales rejects, which locations you can realistically serve or whether a product will be discontinued next month - this is where a human in the loop is required, whilst you can mitigate this to a degree with GA4 and automations, it gets messier the more data connections are required as part of your "SEO setup" that's not agency based - you are sacraficing the knowledge and expertise for something that can DO SEO to a degree but not without someone who understands what's going on.
Consider an illustrative example, a business attracts substantial traffic to a low-margin product category but wants to grow a specialist service with fewer searches and a much higher order value. A traffic-led recommendation could prioritise the category whilst a commercially informed plan might prioritise the service. This can happen far too easily because your "AI platform" might determine low impression queries as not of value when ultimately a keyword with only 10 searches a month could be one that generates a £1m order.
AI cannot make a decision to run with unless it is given FULL CONTEXT all the way around, from product to service, to keyword search data to target SERPS, audiences, competitors etc - and it's not just replacing an SEO agency with an AI that makes things worse, it's that using AI can ultimately send things in the wrong direction if not used properly.
Neither decision can be evaluated properly from clicks alone, or many other metrics for that reason, at least not without context.
Before asking Claude to prioritise SEO, I would provide a current business brief covering products, target customers, markets, profitability, capacity, competitive positioning and measurement limitations. I would also record previous migrations, releases, redirects and major content changes.
There is a further distinction between storing context and using it reliably, this is complex in itself because you need to ensure all the data you have is made available to AI in a way that it can be used and without losing contex. Anthropic’s context-engineering guidance discusses declining retrieval performance as context grows and the need to curate useful information. It does not establish that every current model fails at a particular document size. It does explain why uploading everything is not a substitute for organising the information.
I wrote an article about using AI for keyword research - you may find this particularly interesting because each AI model/platform showed different data, so imagine what would happen at SEO campaign level? See "can you use AI for keyword research"
The third fallacy is accepting a plausible explanation as a diagnosis
Organic traffic has dropped, Claude offers an explanation involving weak content, competitors and missing internal links and you know, It can sound reasonable.
You can give AI snapshots and data -

But unless you provide context, you may be disappointed with the answer, for example AI could easily go through search console data to compare page performance over time to ascertain where losses occured, but beyond that, why did those losses occur? well AI won't be able to answer that.
And we have to ask ourselves but has it established which pages declined, in which countries, on which devices, for which queries? Has it checked whether impressions fell, positions changed or click-through rates shifted? Has it reviewed releases and changes to tracking? One mistake I have seen AI make is when it sees pages have LOST CLICKS, it will highlight the URLS but it won't automatically check to see if the URLS have been removed (thus the click gap).
Google’s own diagnostic guidance considers several possible causes, including technical issues, algorithmic changes, changing search interest and seasonality, It recommends investigating the shape and scope of the decline - which is what you would usually pay your seo agency/consultant for.
My practical interpretation is straightforward: a useful AI workflow should compare explanations before prescribing work.
For example, a seasonal fall in demand, a broken template and weaker search-result click-through rates can all reduce clicks, they require different responses. Rewriting 50 pages does not repair an accidental noindex instruction, reworking internal links does not prove there was anything wrong with demand in the first place.
I would ask for an observed finding, a proposed explanation, the evidence supporting it, evidence against it and the next check needed. If the cause remains uncertain, the report should say so.
That is a higher standard than asking for ten reasons traffic might have fallen and accepting whichever sounds most convincing.
The fourth fallacy is assuming technically correct advice is appropriate everywhere
SEO is full of conditional decisions, advice can describe a real technique and still be wrong for the page being changed.
The examples below are validation scenarios, they are not a claim that Claude/chatGPT necessarily recommends these mistakes.
Blocking crawling is not the same as removing a page from search
Google explains that robots.txt controls crawling and is not a reliable mechanism for excluding a webpage from search results. A URL can still appear without its content being crawled, blocking access can also stop a crawler seeing an instruction on the page.
If the objective is removal from the index, the recommendation needs to address indexing. If the objective is managing crawl activity, it needs to address crawling, those are related but different requirements that would typically require a professional opinion as opposed to a prediction engines output.
A redirect needs a relevant destination
Redirecting every retired page to the homepage is not a sensible universal cleanup policy, google warns that redirecting many old URLs to an irrelevant destination can confuse users and may be treated as a soft 404.
My review would consider the original content, backlinks, user intent and whether a genuine replacement exists. The fact that a redirect rule runs successfully does not establish that its destination makes sense. Can AI perform a redirect map? it can, and it can generally identify pages based on queries, titles and other hook data before comparing old URL to new.
Canonicalisation depends on the relationship between pages
Google documents redirects and canonical annotations as signals for consolidating duplicate or very similar pages, they are not a universal treatment for pages sharing a keyword.
2 pages can discuss the same subject while serving different needs, before consolidating them, I want evidence that their overlap is actually a problem, a comparison page, category page and buying guide may all belong in the journey.
JavaScript needs inspection rather than slogans
Google can render JavaScript. Its documentation also describes rendering constraints, blocked resources and the importance of crawlable links. It recommends keeping canonical signals consistent and warns about initial noindex instructions that may prevent rendering.
So neither “Google cannot see JavaScript” nor “it works in my browser, therefore Google sees everything” is an adequate audit conclusion.
I would inspect the server response, rendered HTML, relevant resources and Google’s own inspection output. Critical content missing from the initial response deserves investigation; it is not automatically proof that Google cannot index it.
Valid schema does not guarantee a search enhancement
Google explicitly says that correct structured data does not guarantee a rich result. Markup also needs to represent the page and meet the relevant eligibility rules.
Generating syntactically valid JSON-LD is a useful task. Deciding which markup is appropriate, checking the facts and evaluating the result are separate tasks.
The fifth fallacy is treating more content as a complete growth strategy
Claude makes content production easier. It does not make every additional page worth publishing.
Google’s scaled content abuse policy focuses on generating many pages primarily to manipulate rankings while offering little value. It applies regardless of whether the work is produced by humans, automation or a mixture of both.
The accurate argument is not that Google automatically penalises AI content. Google’s guidance allows useful applications of generative AI while stressing accuracy, relevance and value.

Snapshot 4. Google’s policy concerns purpose and value, rather than a blanket prohibition on AI-written content.
The business question is what a page adds. Does it answer a real customer question? Does it contain evidence, experience, useful functionality or information unavailable elsewhere? Does it help somebody make a better decision?
If the only distinction is that the same advice has been rewritten around another keyword, I would question the investment.
There is also a maintenance cost. Prices change. Services change. Product details become inaccurate. Five hundred pages create five hundred things that may need reviewing. Publication speed should not conceal that obligation.
The sixth fallacy is confusing brand language with brand credibility
Claude can help express a brand position. It can organise interview notes, draft a case study from supplied evidence and identify inconsistencies in messaging. Those are useful contributions.
It cannot make an invented customer result true or turn a fabricated author biography into genuine expertise.
Google’s people-first content guidance discusses evidence of experience and expertise, trustworthy sourcing and accurate authorship. It also clarifies that E-E-A-T is not itself a single ranking factor.
I would therefore avoid any recommendation that promises an E-E-A-T score increase because a generated biography or author box has been added.
Brand development needs something real behind the wording: a service customers value, evidence of results, an identifiable point of view, competent people and relationships with relevant audiences. AI can help organise and communicate that work. The business still has to substantiate it.
For an SEO programme, this means joining content planning to customer research, product knowledge and public relations. If customers repeatedly question implementation times, give the content team accurate implementation evidence. If sales keeps losing on trust, investigate the missing proof rather than automatically commissioning another introductory article.
Those are strategic choices. They require information from outside the search dashboard.
The seventh fallacy is optimising visibility without the customer journey
Getting somebody to the site is only part of the job - as is the old adage in SEO, you can bring a horse to water but you can't make it drink.
Can the website visitor understand the offer? Compare the options? Find the delivery information? Use the form on a phone? Establish whether the business can solve their problem?
A recommendation can be good for one local objective and poor for the wider experience, adding more keyword-focused navigation links might make the menu harder to use. Expanding a service page might bury the information needed to enquire. Removing low-traffic content might eliminate useful reassurance for a small but valuable audience.
These are possible trade-offs to investigate, not universal reasons to avoid those changes.
Google says Core Web Vitals are used by its ranking systems, but good scores do not guarantee top rankings. Its guidance also distinguishes wider improvements to page experience from direct ranking benefits.
I would make the same distinction with engagement. An improved GA4 engagement rate can be useful evidence about behaviour. It should not be presented as proof that a specific Google ranking signal has improved.
Measure the customer outcome directly: completed enquiries, qualified leads, sales, revenue or another appropriate business objective. Then investigate how search visibility and on-site behaviour contribute to it.
The eighth fallacy is assuming a collection of AI tasks creates a unified SEO programme
You can have a content prompt, a technical audit prompt, a link prospecting prompt and a reporting prompt. Each can produce useful work. They can also produce recommendations that conflict / don't go together properly.
Imagine an illustrative situation where a content workflow proposes new location pages, a technical workflow proposes reducing similar URLs and a navigation workflow removes links to those pages. All three may sound sensible in isolation. Someone must resolve the underlying objective and agree what should exist.
This is the coordination problem businesses need to solve before handing over more responsibility.

Claude can help coordinate these workstreams when the data, instructions and workflows support it. My concern is treating coordination as something that appears automatically because the same model generated all the documents.
A maintained roadmap should explain what matters most, why it matters now, what depends on it and how success will be assessed. Recommendations should be checked against that roadmap before they become tickets or published changes.
The ninth fallacy is believing the right prompt removes the need for independent judgement
Good instructions improve the quality of AI work, they cannot guarantee it - the outcomes are CONSIDERABLY poorer when AI is delegated the responsibility of being an AI SEO agency as opposed to using an actual agency where you have specialists, often with 10+ years SEO experience manning the strategy.
Anthropic’s 2023 research examined sycophancy: models giving answers aligned with a user’s beliefs rather than the most truthful answer. That research concerned the models and tasks tested at the time. It is not a measured error rate for current Claude SEO advice.
The practical risk is still worth considering. If a prompt assumes that an agency is wasting money, that Google has penalised the site or that all low-traffic pages should be deleted, it can frame the analysis around an unproven conclusion.
I would ask Claude to challenge the premise and look for evidence that would change the recommendation. I would also review the answer independently. Asking the same system whether it is confident does not create an external check.
A recent Local Falcon exercise illustrates why the distinction matters. The company asked ChatGPT, Claude and Gemini for checklists, then used Claude to merge them. It reported no outright errors in the combined checklist, but identified missing detail and oversimplification.
That is a small vendor-led exercise, with merged outputs rather than a controlled model comparison. It supports using AI for a starting checklist. It cannot tell us how often Claude independently makes SEO mistakes, or prove that every claim in the vendor’s commentary is established ranking science.
There is no defensible universal percentage in this research for how often Claude gets SEO wrong. Inventing one would undermine the entire argument.
What I would require before acting on an AI recommendation
I would apply this standard to an agency’s recommendation as well. Human delivery is not automatically accurate.
1. Evidence. Identify the affected URLs, metrics, dates, source and reproducible finding.
2. Coverage. State what was retrieved and what was unavailable, filtered, sampled or outside scope.
3. Reasoning. Separate the observation from the proposed explanation and consider alternatives.
4. Commercial relevance. Explain which customer or business outcome the change supports.
5. Dependencies. Check content, technical behaviour, brand, navigation and conversion implications.
6. Validation. Test the change on staging or a limited set where appropriate, and inspect the actual output.
7. Ownership. Name the person approving implementation and define how the change can be reversed.
8. Measurement. Record the baseline and deployment date, then assess outcomes with an appropriate comparison.
For calculations, I would use reproducible code or spreadsheet formulas and reconcile the totals. For technical claims, I would inspect the affected URLs. For a content claim, I would check the source and the subject expertise behind it.
The cost comparison should also include data tools, integration, staff time, review, implementation and maintenance. A subscription price is not the full cost of an SEO operation.
When Claude can reduce your need for external SEO support
A smaller business with a straightforward site, someone willing to learn and access to occasional specialist review may be able to bring substantial work in-house. An experienced SEO team can also use Claude to remove repetitive production and analysis work.
That is a reasonable conclusion. It does not need to be exaggerated into a claim that expertise no longer matters.
For a complex ecommerce site, a migration, multiple countries or a business heavily dependent on organic leads, I would place more weight on competent oversight. The consequences of a wrong decision are larger, and the interactions are harder to assess from one report.
An agency should still have to earn its place. Ask what it has implemented, what improved, what remains uncertain and how it is using AI to make delivery better. A retainer is not evidence of value.
My position is that businesses should use Claude aggressively for work they can validate, while retaining clear responsibility for the decisions that matter. It can help reduce the cost of research, analysis and production. It can also help experienced people see and do more.
If you replace your agency, make sure you have replaced its useful responsibilities as well as its invoice.
























