How to Evaluate a GEO Agency: A 9-Point Checklist
Hiring an agency to improve your generative engine optimisation is a real business decision. It is also a decision being made in a market where the discipline is new enough that bad actors can hide be
Hiring an agency to improve your generative engine optimisation is a real business decision. It is also a decision being made in a market where the discipline is new enough that bad actors can hide behind confusing terminology, where good SEO agencies are legitimately uncertain how much of their expertise transfers, and where almost no one has strong proof of results yet.
This checklist gives you nine criteria to evaluate any GEO agency before signing a contract. Work through each one and you will quickly see which agencies have genuine capability and which are wearing the language without the substance.
1. Can They Define the Specific Outcomes They Are Optimising For?
What to look for: The agency can clearly articulate what GEO success looks like: citation rate improvements in specific AI systems for specific query categories.
Green flag: "We will improve your brand's citation rate in ChatGPT, Perplexity, and Claude for your top 30 target queries from X% to Y% over six months."
Red flag: Vague framing like "improving your AI presence" or "ensuring your brand is visible in the AI era." These are brand promises, not measurable outcomes. Ask what the number is that they expect to move.
Why it matters: An agency that cannot name the specific metric they are optimising for cannot be accountable to results. This is the single most clarifying question you can ask.
2. How Do They Measure AI Visibility?
What to look for: A clear, specific methodology for tracking citation rates across multiple AI systems on a regular cadence.
Green flag: "We track your visibility weekly across ChatGPT, Perplexity, Claude, and Gemini using a defined query set. You get a dashboard showing citation rate by model, prominence score, and competitive position."
Red flag: Manual spot-checks presented as a measurement methodology. "We run some queries each month and report what we see" is not systematic tracking - it is anecdotal observation.
Why it matters: Without systematic measurement, you cannot know whether the agency's work is actually improving your AI visibility or whether they are just producing content activity. Measurement is the foundation of accountability.
See AI Visibility: How It Works for what systematic tracking should look like.
3. Which AI Systems Do They Track?
What to look for: Coverage of at minimum ChatGPT, Perplexity, Claude, and Gemini.
Green flag: A tracking methodology that covers all four major AI systems and explains why each matters (different user populations, different retrieval mechanisms, different citation behaviour).
Red flag: "We track your visibility in ChatGPT" presented as comprehensive AI visibility coverage. ChatGPT is one model. The research-heavy buyers who most influence B2B purchasing are disproportionately on Perplexity and Claude.
4. How Do They Determine Which Content to Create?
What to look for: A gap analysis methodology that identifies specific queries where competitors are cited and the client is not, and uses this to drive content priorities.
Green flag: "We run your target query set across all major AI models, identify the highest-intent queries where your competitors appear and you do not, and build your content roadmap around closing those gaps."
Red flag: A content strategy driven primarily by keyword volume or SEO considerations, with AI visibility as a secondary consideration. This is SEO with GEO branding.
Why it matters: The most valuable GEO content is not the content with the highest keyword volume - it is the content that addresses the queries where you have the biggest citation gap relative to competitors.
5. What Content Formats Do They Recommend and Why?
What to look for: Specific reasoning for content format recommendations based on how AI systems retrieve and cite content.
Green flag: "FAQ pages structured to directly answer the questions AI models receive, comparison guides that explicitly address competitive positioning, and use-case pages with clear, machine-parseable positioning - because these are the formats AI systems pull from most readily."
Red flag: Generic content recommendations - "long-form blog content," "thought leadership articles" - without specific reasoning grounded in AI retrieval behaviour. These formats are fine for SEO; they are not the most effective formats for GEO.
6. Do They Include Structured Data and Technical Implementation?
What to look for: Schema markup implementation (Organisation, Product, FAQ, HowTo) as a standard component of the engagement.
Green flag: Technical deliverables including schema markup on key pages, validation and testing, and a framework for keeping structured data current as content changes.
Red flag: An entirely content-focused proposal with no mention of structured data. GEO without structured data is leaving one of the most direct AI-legibility signals untouched.
7. How Do They Address the Third-Party Reference Network?
What to look for: A programme for building your brand's presence in the external sources AI systems learn from - review platforms, editorial publications, relevant community discussions.
Green flag: A plan for improving your G2/Capterra profile completeness, placing accurate editorial mentions in industry publications, and building presence in the community discussions (Reddit, forums, Slack communities) where your buyers congregate.
Red flag: A proposal that focuses entirely on your own website content without addressing the external reference network. AI systems characterise brands based on what the whole web says about them, not just what your own site says.
See Community Research Guide for how community presence feeds the GEO reference network.
8. What Is Their Track Record, and How Do They Handle the Absence of One?
What to look for: Honesty about what they have and have not proved, combined with a clear case for why their methodology is sound even without an extensive track record.
Green flag (for established agencies): Case studies showing specific citation rate improvements with named clients, defined timelines, and measurement methodology.
Green flag (for newer agencies or consultants): Honest acknowledgment that GEO is a young discipline, combined with a clear articulation of the methodology and why it should work, with a measurement plan that makes the engagement accountable from the start.
Red flag: Vague references to results without specifics, or defensiveness when asked for evidence. The absence of case studies is understandable in a new discipline; the refusal to be specific is not.
9. Do They Define the Feedback Loop from Data to Content?
What to look for: A clear process for how visibility data drives content decisions - how gaps identified in tracking become content briefs and then published content.
Green flag: "Our monthly review process takes the prior month's citation gap data, identifies the three highest-priority unaddressed queries, and generates content briefs for your team or ours to execute against. We track whether each piece of content improves citation rates for its target queries."
Red flag: Tracking and content as separate workstreams that are reported on but not systematically connected. Data without a response protocol is reporting, not a GEO programme.
How to Use This Checklist
Score each agency you are evaluating on a simple green/amber/red system per criterion. Any agency that receives red on criteria 1, 2, 3, or 4 is not delivering a GEO service in any meaningful sense - regardless of what they call it.
Amber ratings on criteria 5-9 are acceptable if the agency is transparent about the limitation and has a plan for addressing it. Red on any of these five warrants a direct conversation before proceeding.
The agencies worth working with are the ones that can engage with this checklist directly, explain their methodology clearly, and make their engagement accountable through systematic measurement.
Monitor your brand in AI answers with Bingly
Track your AI visibility with bing.ly
See how ChatGPT, Perplexity, Claude, and Gemini answer questions about your brand, and monitor community signals across Reddit, Hacker News, and more.
Get started free