Every SEO agency on earth added an AI search page to their site in the last eighteen months. Some of them did the work first.
Evolve Media Agency is an AI search agency. You are reading a guide to evaluating vendors written by a vendor, which is a genuine conflict and you should weigh it accordingly. I have written the version I would want if I were buying, including the tool costs agencies do not usually volunteer, the sections where the honest answer is to do it yourself, and a fair account of why custom pricing is sometimes legitimate rather than evasive. If any of it reads as self-serving, that is a reasonable thing to test against the other guides in this list.
The category has a specific structural problem. AI search optimization is new enough that there is no established credential, no standard measurement, and no agreed vocabulary — so a competent practitioner and a confident one look identical in a first meeting. Meanwhile the underlying work overlaps heavily with SEO, which means an SEO agency can add the service line honestly or dishonestly and the deck looks the same either way.
The good news is that the gap between real and repackaged shows up quickly under specific questions. Not clever ones — just concrete ones about measurement and evidence.
AI search optimization — the practice of making a brand accurately represented and reliably cited inside AI-generated answers. It splits into an on-site half, commonly labelled AEO, covering answer-first structure, schema and extractability, and an off-site half, commonly labelled GEO, covering entity consistency and third-party corroboration across the sources engines actually retrieve.
The Repackaging Problem
To be fair to the industry: real AI search work genuinely does overlap with SEO. Crawlability, structure, page speed, schema and content quality all matter to both. An SEO agency extending into this space is not automatically bluffing, and some of the best practitioners came from exactly that background.
The problem is that the overlap makes bluffing easy. A deck can list "technical foundation, content optimization, entity building" and describe either a serious programme or an SEO retainer with the labels changed.
Three tells that appear early
- Rank language. Talking about "ranking in ChatGPT" or "position one in AI" suggests they are still modelling a ranked list. Citation is probabilistic and frequency-based, not positional.
- Keyword framing. If the proposal centres on keyword research rather than prompt panels and entity corroboration, it is an SEO proposal.
- No mention of off-site work. The on-site half is the cheap, commoditised half. A proposal that is entirely on-site is selling you the easy part.
The honest counterweight
A vendor who is straightforwardly an SEO agency with an AI monitoring add-on is not necessarily the wrong choice — if your technical foundation is genuinely weak, that work has to happen first anyway and they may be excellent at it. The failure is not the SEO heritage, it is being sold that as a complete AI programme at AI programme pricing.
What the Work Actually Splits Into
Insisting a proposal separate these two halves is the fastest way to see what you are buying.
| Dimension | On-site (AEO) | Off-site (GEO) |
|---|---|---|
| What it covers | Answer blocks, FAQ schema, Speakable markup, content restructuring | Entity consistency, editorial mentions, Reddit, review platforms |
| Who controls it | You, entirely | Third parties, mostly |
| Difficulty | Moderate, well-documented | Hard, relationship-dependent |
| Time to effect | Weeks | Quarters |
| Commoditisation | High, and rising | Low |
| Where results come from | Necessary, not sufficient | Where the ceiling is set |
Why the split matters commercially
Adding schema and an llms.txt to your own site is the cheap, table-stakes half. The expensive and effective half is off-site — getting mentioned in the places models actually pull answers from. Research into citation behaviour consistently finds owned domains account for a small share of total citations, which means a proposal weighted entirely toward on-site work is structurally capped no matter how well executed.
Ask for the proposal split by hours or by budget across the two. If off-site is under a third, ask why. There may be a good reason — a genuinely broken technical foundation, for instance — but it should be a stated reason rather than an omission.
What It Costs in 2026
| Engagement | Typical range | Usually includes |
|---|---|---|
| One-time audit | $1,500–5,000 | Baseline panel, technical review, roadmap |
| Entry retainer | $1,500–2,500/mo | Monitoring plus foundational optimization |
| Mid-market retainer | $2,000–10,000/mo | Content, schema, prompt testing, some off-site |
| Enterprise | $10,000–30,000+/mo | Multi-market, large-scale schema, weekly audits |
| Freelance | $75–150/hr or $1,500–4,000/mo | Depth in one area, not full coverage |
| SEO add-on | +20–30% on existing retainer | Monitoring bolted onto current work |
The number agencies do not volunteer
An AI visibility monitoring platform — the tool that produces the prompt-tracking reports in your monthly deck — costs on the order of a hundred dollars a month for capacity covering roughly 150 prompts weekly across five models, enough for about two client projects. So the tooling underlying a $6,000 retainer costs the agency around fifty dollars.
That is not an accusation. Tool cost is a poor proxy for value in any professional service, and the same logic would condemn every SEO, legal or accounting engagement. What it does mean is that you are paying for judgment, execution and relationships, not for software — so the proposal should be able to describe those things concretely. If the deliverable list is mostly reports, you are paying agency rates for a subscription.
Ask what proportion of the retainer is monitoring and reporting versus active work. A programme that is 70% reporting is a dashboard with a consultant attached. A programme that is 70% doing things — content restructuring, schema deployment, outreach, review generation — is what you meant to buy.
Questions 1–4: Measurement
Start here. If measurement is vague, nothing downstream can be verified.
- Show me a prompt-testing report from a real client. Anonymized is fine. You are looking for named platforms, specific queries, and which URLs were cited. This single request separates most of the field.
- What is my baseline, and how will you establish it? A programme that begins optimizing before measuring cannot demonstrate causation later. The baseline is the deliverable of week one.
- How many prompts will you track, on which engines, how often? Specific numbers. Fifty to a hundred prompts across four or five engines, sampled quarterly or monthly, is a reasonable shape.
- What is the metric, and what is the denominator? Citation rate, share of answer, appearance frequency — any is defensible. Not defining one is not.
Why question one does most of the work
Producing a prompt-testing report requires having actually run the panels, which requires having built them, which requires understanding what to measure. It is very difficult to fake and trivial to provide if you do it. An agency charging at the higher end of the ranges above should be able to hand you one within a day.
Our AI visibility audit guide covers what a proper baseline diagnostic contains, which is useful context for evaluating whether theirs is thorough.
Questions 5–9: Methodology
- Split the proposal into on-site and off-site. What is the ratio? Covered above. If off-site is under a third, ask for the reasoning.
- How do you approach third-party corroboration specifically? Listen for concrete surfaces — review platforms, community threads, editorial outreach, industry directories — rather than "we build authority."
- What is your position on llms.txt? A useful honesty test. The evidence for it moving citation rates is thin. An agency presenting it as a primary lever is either behind or overselling; one that says it is cheap and worth having but not the mechanism is being straight with you.
- How do you handle the differences between engines? Perplexity weights freshness far more heavily than others; ChatGPT behaves differently again. A vendor treating all engines as one surface has not measured them separately.
- What would you do first, and why that? The answer should be diagnostic — find out whether crawlers can reach you, establish the baseline — rather than a list of deliverables they sell to everyone.
Ask what they think about llms.txt. It is the cheapest available honesty test, because the evidence is thin and everyone in the field knows it.
For context on what the off-site half actually involves, see our Reddit strategy for AI citations and the 30-signal citation audit.
Questions 10–12: Deliverables
- What proportion of the retainer is reporting versus doing? The single best value question. Aim for a programme weighted toward action.
- Who does the work, and are they the people in this meeting? The senior-sells, junior-delivers pattern is as common here as anywhere. Ask for names and ask who you will actually speak to monthly.
- What do I own at the end? Content, schema implementations, the prompt panel itself, the historical measurement data, any accounts created on your behalf. All of it should be yours.
The ownership question in more detail
The prompt panel is the one people forget. It is the instrument that measures everything, it takes real thought to construct, and an agency that keeps it means you restart measurement from zero when you leave — and lose the ability to compare before and after their engagement. Get it in writing that the panel and its historical data transfer to you.
Comparing proposals?
Send us what you have been quoted. We will tell you what is standard scope, what is thin, and which questions to push on — including if the answer is that you do not need an agency.
Book a Strategy Call →The Ecom Profit Box
Eleven playbooks on listings, conversion, images, and email. Built for operators, no fluff, no email sequence.
Grab It Free →Questions 13–14: Accountability
- What result at six months would you consider a failure? A vendor unwilling to name a failure threshold has committed to nothing, and you will have no basis for a difficult conversation later.
- What is the exit? Notice period, and what transfers? Thirty days is reasonable. Ninety with auto-renewal is a trap. Confirm what you walk away with.
On guarantees
Some vendors now offer refund guarantees — money back if you are not being recommended within a defined window. That is a genuinely interesting development because it puts something at risk, and it is more meaningful than any case study.
Read the condition carefully, though. "Being recommended" for which prompts, on which engines, measured by whom? A guarantee against a panel the vendor constructs and grades is weaker than it sounds. A guarantee against a panel you construct is a real commitment.
The timeline honesty test
Ask how long before meaningful movement. An honest answer is three to six months for first signal, longer for the off-site half, because third-party corroboration depends on other people publishing. Anyone promising results in thirty days is describing a timeline the mechanism does not support — and our analysis of how citations compound over twelve months covers why.
Good Answers vs Worrying Answers
| Question | Good answer | Worrying answer |
|---|---|---|
| Prompt-testing report | Sends one within a day | Explains why they cannot share client data |
| Baseline | Week one deliverable, named method | "We'll track improvements as we go" |
| On-site vs off-site split | A number, with reasoning | Treats the question as odd |
| llms.txt | Cheap, worth having, not the lever | Presented as a primary deliverable |
| Engine differences | Names specific behavioural differences | "We optimize for all AI platforms" |
| Reporting vs doing | A ratio, weighted toward doing | Deliverables list is mostly documents |
| Failure threshold | A specific metric and number | "It depends on many factors" |
| What I own | Everything, including the panel | Vague, or panel retained |
The meta-signal
Notice how many of the good answers are simply specific. That is the actual filter. This field is new enough that nobody has all the answers, and a vendor saying "we do not know yet, here is how we would find out" is more credible than one with a confident answer to everything. Certainty is the tell.
Why Nobody Publishes Pricing
Worth addressing fairly, because the absence of published rates reads as evasive and is sometimes legitimate.
The legitimate reasons
- Scope genuinely varies enormously. A ten-page site and a two-thousand-page catalog are different projects.
- Measurement standards are not uniform. Without agreed metrics, a published price attaches to an undefined deliverable.
- Platforms change frequently. A rate card written against last quarter's engine behaviour ages badly.
- Discovery is usually necessary before a meaningful scope exists. Custom proposals are the norm in the category rather than automatically an evasion tactic.
The less legitimate reason
Vague pricing is easier to inflate. When nobody publishes, every buyer negotiates without a reference point, and the same scope can be quoted at $3,000 to one client and $9,000 to another based on perceived willingness to pay.
How to protect yourself either way
- Get two or three proposals with the same brief. The variance itself is informative.
- Ask for the price broken into components — audit, monthly monitoring, content production, off-site work — so you can compare like with like.
- Ask what a smaller version costs. A vendor who cannot scope down usually has one package rather than a practice.
- Consider a fixed-scope build first. One-time foundation engagements exist in the $1,500 to $5,000 range and let you assess quality before committing to an open-ended retainer.
Our AI search agency pricing guide goes deeper on the ranges, and the in-house versus agency calculator covers the build-or-buy arithmetic.
What You Should See at 1, 3, 6, 12 Months
Prompt panel built and run, crawler access verified and fixed, technical gaps documented. You should know exactly where you stand.
Schema deployed and validated, priority pages restructured answer-first, entity definition consistent everywhere. First re-measure, modest movement.
Citation rate up against baseline on the same panel. Off-site work producing its first mentions. This is the honest first judgment point.
Corroboration accumulating, citation rate rising non-linearly, competitors appearing less often alongside you in your own category.
The point where most engagements fail
Month three. The foundation work is done, the re-measure shows little, and the temptation to cancel is strongest — precisely when the slow off-site levers have not had time to mature. A good agency tells you this in month one so it is expected rather than alarming. A poor one lets you discover it and then explains why it is not their fault.
Agree in advance that month six is the judgment point and that month three is a checkpoint, not a verdict.
When You Do Not Need an Agency
Written by an agency, so weigh it accordingly — but these cases are real and common.
- You have never measured. Run a prompt panel yourself first. Ninety minutes and a spreadsheet tells you whether you have a problem worth paying to solve, and you will brief vendors far better afterwards.
- Your crawlers are blocked. If AI bots cannot reach your site, that is a configuration fix, not a retainer. Check robots.txt, your CDN and your render path before buying anything.
- Your schema is missing entirely. A one-time implementation is a project, not an ongoing programme. Buy the build, not the subscription.
- You already have a strong content team. The on-site half is well documented and your team can execute it. Consider buying only the off-site work, or only the measurement.
- You are under roughly $500K in revenue. A $3,000 monthly retainer against that base is a large bet on a channel you have not yet quantified.
- You want one specific thing. A freelancer at $75 to $150 an hour is often the better instrument for a defined problem than a retainer covering everything.
Measure yourself. Fix crawler access. Deploy schema. Re-measure. Then decide whether to hire, and hire specifically for the gap the second measurement revealed. Most brands who do that discover they need less than they thought, and the ones who genuinely need an agency arrive with a far better brief.
The 90-minute AI audit is the version of this you can run without help.
The Decision Framework
| Your situation | Buy this | Budget |
|---|---|---|
| Never measured | Do it yourself, or a one-time audit | $0–5,000 |
| Technical foundation weak | Fixed-scope build, not a retainer | $1,500–5,000 once |
| Foundation solid, no citations | Off-site focused retainer | $2,000–6,000/mo |
| Competitive category, scaling | Full mid-market retainer | $5,000–10,000/mo |
| Multi-market or enterprise | Enterprise engagement | $10,000–30,000/mo |
| One defined problem | Freelance specialist | $75–150/hr |
| Strong internal content team | Measurement plus off-site only | $1,500–4,000/mo |
| Under $500K revenue | Do it yourself for now | Time, not money |
The three things to settle before signing
- The baseline — measured, documented, and owned by you.
- The split — on-site versus off-site, as a number, with reasoning.
- The failure threshold — a specific metric at a specific month, agreed by both sides.
Settle those and you can compare proposals meaningfully. Leave any of them vague and you will be having a difficult conversation in month seven without an agreed basis for it.
For the vendor landscape itself, see our review of AI search optimization agencies, and for the measurement tooling, our AI visibility tracking tools comparison.
The Short Version
- Pricing bands: one-time audits $1,500 to $5,000, entry retainers $1,500 to $2,500 monthly, mid-market $2,000 to $10,000, enterprise $10,000 to $30,000-plus. Many agencies sell it as a 20 to 30% surcharge on an SEO retainer.
- The monitoring platform behind your reports costs around a hundred dollars monthly and covers roughly two clients. You are paying for judgment and execution, not software — so the proposal should describe those concretely.
- Ask for a prompt-testing report from a real client. It is very hard to fake, trivial to provide if you do the work, and separates most of the field in one request.
- Insist the proposal split on-site from off-site work. The on-site half is cheap and commoditised; the off-site half sets the ceiling and is where thin proposals are thinnest.
- Ask what they think about llms.txt. Evidence for it moving citations is thin, so presenting it as a primary lever reveals either overselling or being behind.
- Month three is where most engagements fail, because foundation work is done and slow off-site levers have not matured. Agree upfront that month six is the judgment point.
- Before hiring: measure yourself, fix crawler access, deploy schema, re-measure. Most brands then discover they need less than they thought.

