BUYER'S GUIDE PUBLISHED SEPTEMBER 1, 2026 · 15 MIN READ

How To Hire An AI Search Agency.

The monitoring platform an agency uses to track your AI visibility costs around a hundred dollars a month and covers two clients. The retainer built on top of it commonly runs $2,000 to $10,000. That gap is not automatically unreasonable — but you should know it exists before you evaluate a proposal.

$2–10KTypical mid-market monthly retainer
20–30%Common surcharge on an SEO retainer
14Questions before you sign
1That does most of the work
Quick Answer

AI search agency pricing in 2026 clusters into four bands: one-time audits at $1,500 to $5,000, entry-level programmes around $1,500 to $2,500 monthly for basic monitoring and foundational work, mid-market retainers at $2,000 to $10,000 monthly, and enterprise engagements from $10,000 to $30,000-plus. Experienced freelancers charge $75 to $150 hourly or $1,500 to $4,000 on retainer. Many agencies sell the service as a 20 to 30% surcharge on an existing SEO retainer. The single most useful filter is asking to see a prompt-testing report showing which AI platforms cite which of your pages for which queries — an agency charging at the top of these ranges without producing citation evidence has not justified the price, regardless of how good the deliverables list looks. The second most useful is asking them to separate on-site work from off-site work, because the on-site half is comparatively cheap and commoditised while the off-site half is where results actually come from and where most proposals are thinnest.

Every SEO agency on earth added an AI search page to their site in the last eighteen months. Some of them did the work first.

Disclosure, and it matters more here than usual

Evolve Media Agency is an AI search agency. You are reading a guide to evaluating vendors written by a vendor, which is a genuine conflict and you should weigh it accordingly. I have written the version I would want if I were buying, including the tool costs agencies do not usually volunteer, the sections where the honest answer is to do it yourself, and a fair account of why custom pricing is sometimes legitimate rather than evasive. If any of it reads as self-serving, that is a reasonable thing to test against the other guides in this list.

The category has a specific structural problem. AI search optimization is new enough that there is no established credential, no standard measurement, and no agreed vocabulary — so a competent practitioner and a confident one look identical in a first meeting. Meanwhile the underlying work overlaps heavily with SEO, which means an SEO agency can add the service line honestly or dishonestly and the deck looks the same either way.

The good news is that the gap between real and repackaged shows up quickly under specific questions. Not clever ones — just concrete ones about measurement and evidence.

Definition

AI search optimization — the practice of making a brand accurately represented and reliably cited inside AI-generated answers. It splits into an on-site half, commonly labelled AEO, covering answer-first structure, schema and extractability, and an off-site half, commonly labelled GEO, covering entity consistency and third-party corroboration across the sources engines actually retrieve.

01/12SECTION

The Repackaging Problem

To be fair to the industry: real AI search work genuinely does overlap with SEO. Crawlability, structure, page speed, schema and content quality all matter to both. An SEO agency extending into this space is not automatically bluffing, and some of the best practitioners came from exactly that background.

The problem is that the overlap makes bluffing easy. A deck can list "technical foundation, content optimization, entity building" and describe either a serious programme or an SEO retainer with the labels changed.

Three tells that appear early

  • Rank language. Talking about "ranking in ChatGPT" or "position one in AI" suggests they are still modelling a ranked list. Citation is probabilistic and frequency-based, not positional.
  • Keyword framing. If the proposal centres on keyword research rather than prompt panels and entity corroboration, it is an SEO proposal.
  • No mention of off-site work. The on-site half is the cheap, commoditised half. A proposal that is entirely on-site is selling you the easy part.

The honest counterweight

A vendor who is straightforwardly an SEO agency with an AI monitoring add-on is not necessarily the wrong choice — if your technical foundation is genuinely weak, that work has to happen first anyway and they may be excellent at it. The failure is not the SEO heritage, it is being sold that as a complete AI programme at AI programme pricing.

02/12SECTION

What the Work Actually Splits Into

Insisting a proposal separate these two halves is the fastest way to see what you are buying.

DimensionOn-site (AEO)Off-site (GEO)
What it coversAnswer blocks, FAQ schema, Speakable markup, content restructuringEntity consistency, editorial mentions, Reddit, review platforms
Who controls itYou, entirelyThird parties, mostly
DifficultyModerate, well-documentedHard, relationship-dependent
Time to effectWeeksQuarters
CommoditisationHigh, and risingLow
Where results come fromNecessary, not sufficientWhere the ceiling is set

Why the split matters commercially

Adding schema and an llms.txt to your own site is the cheap, table-stakes half. The expensive and effective half is off-site — getting mentioned in the places models actually pull answers from. Research into citation behaviour consistently finds owned domains account for a small share of total citations, which means a proposal weighted entirely toward on-site work is structurally capped no matter how well executed.

Ask for the proposal split by hours or by budget across the two. If off-site is under a third, ask why. There may be a good reason — a genuinely broken technical foundation, for instance — but it should be a stated reason rather than an omission.

03/12SECTION

What It Costs in 2026

EngagementTypical rangeUsually includes
One-time audit$1,500–5,000Baseline panel, technical review, roadmap
Entry retainer$1,500–2,500/moMonitoring plus foundational optimization
Mid-market retainer$2,000–10,000/moContent, schema, prompt testing, some off-site
Enterprise$10,000–30,000+/moMulti-market, large-scale schema, weekly audits
Freelance$75–150/hr or $1,500–4,000/moDepth in one area, not full coverage
SEO add-on+20–30% on existing retainerMonitoring bolted onto current work

The number agencies do not volunteer

An AI visibility monitoring platform — the tool that produces the prompt-tracking reports in your monthly deck — costs on the order of a hundred dollars a month for capacity covering roughly 150 prompts weekly across five models, enough for about two client projects. So the tooling underlying a $6,000 retainer costs the agency around fifty dollars.

That is not an accusation. Tool cost is a poor proxy for value in any professional service, and the same logic would condemn every SEO, legal or accounting engagement. What it does mean is that you are paying for judgment, execution and relationships, not for software — so the proposal should be able to describe those things concretely. If the deliverable list is mostly reports, you are paying agency rates for a subscription.

The fair test

Ask what proportion of the retainer is monitoring and reporting versus active work. A programme that is 70% reporting is a dashboard with a consultant attached. A programme that is 70% doing things — content restructuring, schema deployment, outreach, review generation — is what you meant to buy.

04/12SECTION

Questions 1–4: Measurement

Start here. If measurement is vague, nothing downstream can be verified.

  1. Show me a prompt-testing report from a real client. Anonymized is fine. You are looking for named platforms, specific queries, and which URLs were cited. This single request separates most of the field.
  2. What is my baseline, and how will you establish it? A programme that begins optimizing before measuring cannot demonstrate causation later. The baseline is the deliverable of week one.
  3. How many prompts will you track, on which engines, how often? Specific numbers. Fifty to a hundred prompts across four or five engines, sampled quarterly or monthly, is a reasonable shape.
  4. What is the metric, and what is the denominator? Citation rate, share of answer, appearance frequency — any is defensible. Not defining one is not.

Why question one does most of the work

Producing a prompt-testing report requires having actually run the panels, which requires having built them, which requires understanding what to measure. It is very difficult to fake and trivial to provide if you do it. An agency charging at the higher end of the ranges above should be able to hand you one within a day.

Our AI visibility audit guide covers what a proper baseline diagnostic contains, which is useful context for evaluating whether theirs is thorough.

05/12SECTION

Questions 5–9: Methodology

  1. Split the proposal into on-site and off-site. What is the ratio? Covered above. If off-site is under a third, ask for the reasoning.
  2. How do you approach third-party corroboration specifically? Listen for concrete surfaces — review platforms, community threads, editorial outreach, industry directories — rather than "we build authority."
  3. What is your position on llms.txt? A useful honesty test. The evidence for it moving citation rates is thin. An agency presenting it as a primary lever is either behind or overselling; one that says it is cheap and worth having but not the mechanism is being straight with you.
  4. How do you handle the differences between engines? Perplexity weights freshness far more heavily than others; ChatGPT behaves differently again. A vendor treating all engines as one surface has not measured them separately.
  5. What would you do first, and why that? The answer should be diagnostic — find out whether crawlers can reach you, establish the baseline — rather than a list of deliverables they sell to everyone.
Ask what they think about llms.txt. It is the cheapest available honesty test, because the evidence is thin and everyone in the field knows it.
Ian Smith · Evolve Media Agency

For context on what the off-site half actually involves, see our Reddit strategy for AI citations and the 30-signal citation audit.

06/12SECTION

Questions 10–12: Deliverables

  1. What proportion of the retainer is reporting versus doing? The single best value question. Aim for a programme weighted toward action.
  2. Who does the work, and are they the people in this meeting? The senior-sells, junior-delivers pattern is as common here as anywhere. Ask for names and ask who you will actually speak to monthly.
  3. What do I own at the end? Content, schema implementations, the prompt panel itself, the historical measurement data, any accounts created on your behalf. All of it should be yours.

The ownership question in more detail

The prompt panel is the one people forget. It is the instrument that measures everything, it takes real thought to construct, and an agency that keeps it means you restart measurement from zero when you leave — and lose the ability to compare before and after their engagement. Get it in writing that the panel and its historical data transfer to you.

FREE 30-MINUTE CALL

Comparing proposals?

Send us what you have been quoted. We will tell you what is standard scope, what is thin, and which questions to push on — including if the answer is that you do not need an agency.

Book a Strategy Call →
FREE RESOURCE

The Ecom Profit Box

Eleven playbooks on listings, conversion, images, and email. Built for operators, no fluff, no email sequence.

Grab It Free →
07/12SECTION

Questions 13–14: Accountability

  1. What result at six months would you consider a failure? A vendor unwilling to name a failure threshold has committed to nothing, and you will have no basis for a difficult conversation later.
  2. What is the exit? Notice period, and what transfers? Thirty days is reasonable. Ninety with auto-renewal is a trap. Confirm what you walk away with.

On guarantees

Some vendors now offer refund guarantees — money back if you are not being recommended within a defined window. That is a genuinely interesting development because it puts something at risk, and it is more meaningful than any case study.

Read the condition carefully, though. "Being recommended" for which prompts, on which engines, measured by whom? A guarantee against a panel the vendor constructs and grades is weaker than it sounds. A guarantee against a panel you construct is a real commitment.

The timeline honesty test

Ask how long before meaningful movement. An honest answer is three to six months for first signal, longer for the off-site half, because third-party corroboration depends on other people publishing. Anyone promising results in thirty days is describing a timeline the mechanism does not support — and our analysis of how citations compound over twelve months covers why.

08/12SECTION

Good Answers vs Worrying Answers

QuestionGood answerWorrying answer
Prompt-testing reportSends one within a dayExplains why they cannot share client data
BaselineWeek one deliverable, named method"We'll track improvements as we go"
On-site vs off-site splitA number, with reasoningTreats the question as odd
llms.txtCheap, worth having, not the leverPresented as a primary deliverable
Engine differencesNames specific behavioural differences"We optimize for all AI platforms"
Reporting vs doingA ratio, weighted toward doingDeliverables list is mostly documents
Failure thresholdA specific metric and number"It depends on many factors"
What I ownEverything, including the panelVague, or panel retained

The meta-signal

Notice how many of the good answers are simply specific. That is the actual filter. This field is new enough that nobody has all the answers, and a vendor saying "we do not know yet, here is how we would find out" is more credible than one with a confident answer to everything. Certainty is the tell.

09/12SECTION

Why Nobody Publishes Pricing

Worth addressing fairly, because the absence of published rates reads as evasive and is sometimes legitimate.

The legitimate reasons

  • Scope genuinely varies enormously. A ten-page site and a two-thousand-page catalog are different projects.
  • Measurement standards are not uniform. Without agreed metrics, a published price attaches to an undefined deliverable.
  • Platforms change frequently. A rate card written against last quarter's engine behaviour ages badly.
  • Discovery is usually necessary before a meaningful scope exists. Custom proposals are the norm in the category rather than automatically an evasion tactic.

The less legitimate reason

Vague pricing is easier to inflate. When nobody publishes, every buyer negotiates without a reference point, and the same scope can be quoted at $3,000 to one client and $9,000 to another based on perceived willingness to pay.

How to protect yourself either way

  • Get two or three proposals with the same brief. The variance itself is informative.
  • Ask for the price broken into components — audit, monthly monitoring, content production, off-site work — so you can compare like with like.
  • Ask what a smaller version costs. A vendor who cannot scope down usually has one package rather than a practice.
  • Consider a fixed-scope build first. One-time foundation engagements exist in the $1,500 to $5,000 range and let you assess quality before committing to an open-ended retainer.

Our AI search agency pricing guide goes deeper on the ranges, and the in-house versus agency calculator covers the build-or-buy arithmetic.

10/12SECTION

What You Should See at 1, 3, 6, 12 Months

REASONABLE EXPECTATIONS BY STAGEAGREE THESE UPFRONT
MONTH 01
Baseline And Diagnosis

Prompt panel built and run, crawler access verified and fixed, technical gaps documented. You should know exactly where you stand.

MONTH 03
Foundation Complete

Schema deployed and validated, priority pages restructured answer-first, entity definition consistent everywhere. First re-measure, modest movement.

MONTH 06
Measurable Movement

Citation rate up against baseline on the same panel. Off-site work producing its first mentions. This is the honest first judgment point.

MONTH 12
Compounding

Corroboration accumulating, citation rate rising non-linearly, competitors appearing less often alongside you in your own category.

The point where most engagements fail

Month three. The foundation work is done, the re-measure shows little, and the temptation to cancel is strongest — precisely when the slow off-site levers have not had time to mature. A good agency tells you this in month one so it is expected rather than alarming. A poor one lets you discover it and then explains why it is not their fault.

Agree in advance that month six is the judgment point and that month three is a checkpoint, not a verdict.

11/12SECTION

When You Do Not Need an Agency

Written by an agency, so weigh it accordingly — but these cases are real and common.

  • You have never measured. Run a prompt panel yourself first. Ninety minutes and a spreadsheet tells you whether you have a problem worth paying to solve, and you will brief vendors far better afterwards.
  • Your crawlers are blocked. If AI bots cannot reach your site, that is a configuration fix, not a retainer. Check robots.txt, your CDN and your render path before buying anything.
  • Your schema is missing entirely. A one-time implementation is a project, not an ongoing programme. Buy the build, not the subscription.
  • You already have a strong content team. The on-site half is well documented and your team can execute it. Consider buying only the off-site work, or only the measurement.
  • You are under roughly $500K in revenue. A $3,000 monthly retainer against that base is a large bet on a channel you have not yet quantified.
  • You want one specific thing. A freelancer at $75 to $150 an hour is often the better instrument for a defined problem than a retainer covering everything.
The sequence I would actually recommend

Measure yourself. Fix crawler access. Deploy schema. Re-measure. Then decide whether to hire, and hire specifically for the gap the second measurement revealed. Most brands who do that discover they need less than they thought, and the ones who genuinely need an agency arrive with a far better brief.

The 90-minute AI audit is the version of this you can run without help.

12/12SECTION

The Decision Framework

Your situationBuy thisBudget
Never measuredDo it yourself, or a one-time audit$0–5,000
Technical foundation weakFixed-scope build, not a retainer$1,500–5,000 once
Foundation solid, no citationsOff-site focused retainer$2,000–6,000/mo
Competitive category, scalingFull mid-market retainer$5,000–10,000/mo
Multi-market or enterpriseEnterprise engagement$10,000–30,000/mo
One defined problemFreelance specialist$75–150/hr
Strong internal content teamMeasurement plus off-site only$1,500–4,000/mo
Under $500K revenueDo it yourself for nowTime, not money

The three things to settle before signing

  1. The baseline — measured, documented, and owned by you.
  2. The split — on-site versus off-site, as a number, with reasoning.
  3. The failure threshold — a specific metric at a specific month, agreed by both sides.

Settle those and you can compare proposals meaningfully. Leave any of them vague and you will be having a difficult conversation in month seven without an agreed basis for it.

For the vendor landscape itself, see our review of AI search optimization agencies, and for the measurement tooling, our AI visibility tracking tools comparison.

Key Takeaways

The Short Version

  • Pricing bands: one-time audits $1,500 to $5,000, entry retainers $1,500 to $2,500 monthly, mid-market $2,000 to $10,000, enterprise $10,000 to $30,000-plus. Many agencies sell it as a 20 to 30% surcharge on an SEO retainer.
  • The monitoring platform behind your reports costs around a hundred dollars monthly and covers roughly two clients. You are paying for judgment and execution, not software — so the proposal should describe those concretely.
  • Ask for a prompt-testing report from a real client. It is very hard to fake, trivial to provide if you do the work, and separates most of the field in one request.
  • Insist the proposal split on-site from off-site work. The on-site half is cheap and commoditised; the off-site half sets the ceiling and is where thin proposals are thinnest.
  • Ask what they think about llms.txt. Evidence for it moving citations is thin, so presenting it as a primary lever reveals either overselling or being behind.
  • Month three is where most engagements fail, because foundation work is done and slow off-site levers have not matured. Agree upfront that month six is the judgment point.
  • Before hiring: measure yourself, fix crawler access, deploy schema, re-measure. Most brands then discover they need less than they thought.

Common Questions

Hiring an AI Search Agency
FAQ

How much should an AI search agency cost?

One-time audits run $1,500 to $5,000. Entry-level retainers covering monitoring and foundational work sit around $1,500 to $2,500 monthly. Mid-market retainers cluster at $2,000 to $10,000, and enterprise engagements run $10,000 to $30,000-plus. Experienced freelancers charge $75 to $150 hourly or $1,500 to $4,000 on retainer. Many agencies also sell it as a 20 to 30% surcharge on an existing SEO retainer.

What is the single best question to ask?

Ask to see a prompt-testing report from a real client, anonymized if necessary, showing which AI platforms cited which pages for which queries. Producing one requires having actually built and run measurement panels, which is very difficult to fake and trivial to provide if you do the work. An agency charging at the higher end of market rates should be able to hand you one within a day.

How do I tell a real AI search agency from a rebranded SEO shop?

Three tells appear early. Talking about ranking or position in AI answers suggests they are still modelling a ranked list rather than probabilistic citation. Centring the proposal on keyword research rather than prompt panels and entity corroboration means it is an SEO proposal. And a proposal with no off-site component is selling the cheap, commoditised half. That said, genuine SEO heritage is not disqualifying, since the disciplines really do overlap.

Why do so few agencies publish their prices?

Partly for legitimate reasons: scope varies enormously between a ten-page site and a large catalog, measurement standards are not uniform across the industry, platforms change frequently enough that rate cards age badly, and a discovery phase is usually needed before a meaningful scope exists. Custom proposals are the category norm rather than automatically evasion. The less legitimate reason is that vague pricing is easier to inflate when buyers have no reference point.

What is the difference between AEO and GEO in a proposal?

AEO generally covers on-site work: direct answer blocks, FAQ schema, Speakable markup and restructuring content so models can extract factual answers. GEO generally covers off-site work: building entity consistency and corroboration across editorial press, community platforms, review sites and reference sources. Most agencies handle both as one discipline, but insisting the proposal separates them tells you where the budget is actually going.

How much of my retainer should be off-site work?

At least a third, and often more once your technical foundation is solid. Research into citation behaviour consistently finds owned domains account for a small share of total citations, which means a programme weighted entirely toward on-site optimization is structurally capped regardless of execution quality. If a proposal is under a third off-site, ask for the reasoning, since there may be a good one such as a genuinely broken technical base.

Is llms.txt a good sign or a bad sign in a proposal?

It depends entirely on how it is framed. Evidence for llms.txt moving citation rates remains thin, so an agency presenting it as a primary deliverable is either behind the research or overselling an easy item. An agency describing it as cheap, worth implementing and not the mechanism is being straight with you. It is one of the most useful honesty tests available because everyone working in the field knows the evidence is weak.

How long before I should expect results?

Three to six months for first meaningful movement, longer for the off-site half because third-party corroboration depends on other people publishing. Month one should deliver a baseline and diagnosis, month three a completed foundation with modest movement, month six the honest first judgment point, and month twelve compounding. Anyone promising results in thirty days is describing a timeline the underlying mechanism does not support.

What should I own when the engagement ends?

Everything: content produced, schema implementations, any accounts created on your behalf, the historical measurement data, and critically the prompt panel itself. The panel is the one people forget. It is the instrument that measures everything, takes real thought to construct, and if the agency keeps it you restart measurement from zero and lose the ability to compare before and after their work. Get the transfer in writing.

Are refund guarantees meaningful?

Potentially, because they put something at risk, which is more than most case studies do. Read the condition carefully though. Being recommended for which prompts, on which engines, measured by whom? A guarantee assessed against a panel the vendor builds and grades is considerably weaker than one assessed against a panel you construct. The guarantee is only as strong as the independence of the measurement behind it.

When should I not hire an AI search agency?

If you have never measured, run a prompt panel yourself first, since ninety minutes and a spreadsheet tells you whether you have a problem worth paying to solve. If AI crawlers are blocked, that is a configuration fix rather than a retainer. If your schema is missing, buy a one-time build rather than a subscription. And under roughly $500K in revenue, a $3,000 monthly retainer is a large bet on a channel you have not yet quantified.

Why do most engagements fail at month three?

Because the foundation work is complete, the re-measure shows little, and the slow off-site levers have not had time to mature. That combination makes month three the point of maximum temptation to cancel and minimum evidence either way. A good agency tells you this in month one so it is expected rather than alarming. Agree upfront that month three is a checkpoint and month six is the judgment point.

Ian Smith, Founder of Evolve Media Agency
Ian Smith
Founder, Evolve Media Agency · AI Search & Ecommerce Specialist

Ian co-founded Evolve Media Agency in 2017 with his wife Megan. Over 9 years he has worked with $1M-$10M ecommerce brands on AI search visibility, schema infrastructure, content production, and channel diversification. Based in Colorado. Read Ian’s full bio →

Work With Ian

Including if the answer is no

Send Us The Proposal.

We will tell you what is standard scope, what is thin, and which questions to push on — and if the honest answer is that you should do it yourself for now, we will say that instead.