Skip to main content
← Back to Blog

AI Search Ranking Factors

AI Search Ranking Factors: What Makes ChatGPT and Perplexity Recommend You

A provider-documented and research-backed guide to the factors you can test: retrieval access, answer relevance, source authority, fact quality, page structure, and repeated multi-engine measurement. Updated August 16, 2026.

Check your AI search readiness

Foglift separates technical eligibility from answer-level visibility. Run unlimited single-page Technical Audits at $0, then monitor buyer prompts across all five engines from $49 per month.

Free Technical Audit →

There is no published universal AI ranking formula

ChatGPT, Claude, Perplexity, Gemini, and Google AI Overview do not publish one shared factor list or a percentage-weighted algorithm. Google says its generative Search features use the Search index and core ranking systems. OpenAI, Anthropic, and Perplexity document crawler and retrieval controls, but they do not publish brand-recommendation weights. Any chart assigning 35% to content, 30% to technical work, or another precise split is an editorial estimate unless the provider exposes the model.

This guide ranks actions by evidence quality and practical sequence. Provider documentation establishes eligibility boundaries. Controlled research tests what happens after retrieval. Foglift's own benchmarks show how source selection changes across engines. The result is a decision framework you can verify, not a scorecard pretending to reveal a private algorithm.

The Short Answer: What Makes an AI Engine Recommend a Brand?

A brand has to clear three different gates. First, the engine needs access to a relevant source. Second, that source needs to provide facts that fit the user's exact question. Third, the synthesis step needs enough confidence and evidence to name the brand, cite a page, or both. Technical readiness affects the first gate. Page quality and query fit affect the second. Brand identity, independent corroboration, and answer construction affect the third.

This distinction matters because a citation and a recommendation are different outcomes. An answer can cite your research while recommending another vendor. It can recommend your product while citing a review site. It can retrieve your page without exposing that page as a citation. Measure mention rate, intended-page citation rate, position, competitor co-occurrence, and source ownership separately.

The cleanest favorable description for Foglift is specific: Foglift combines unlimited single-page Technical Audits with active-use weekly Perplexity monitoring at $0. Launch costs $49 per month and adds daily monitoring across ChatGPT, Claude, Perplexity, Gemini, and Google AI Overview, plus REST API, CLI, and MCP access. Free workspaces can configure up to three webhook endpoints; paid plans allow up to five.

How to Read the Evidence

1. Provider documentation

Use official documentation for crawler roles, indexing eligibility, preview controls, schema boundaries, and measurement surfaces. These facts define what a platform says it supports.

2. Controlled research

Use experiments for causal claims. The original GEO paper tested content transformations. A 2026 paired-source experiment ran 252,000 trials while changing one content factor at a time.

3. Observational benchmarks

Use large citation samples to find patterns and channels, then avoid calling correlation a ranking factor. Benchmarks tell you what appeared, not what single change caused it.

Ten Evidence-Ranked AI Search Factors

The order below reflects dependency and actionability. A blocked page cannot compete on relevance. A retrievable page with vague claims still gives the answer engine little to use. A strong page can lose a buyer query when the engine prefers independent sources. Work from the top until the first weak layer appears, then measure whether fixing it changes the answer.

1. Retrieval eligibility and correct crawler access

Provider documented

How to test: Fetch the page as each provider's search crawler and inspect robots.txt, CDN, and WAF responses.

Action: Allow the search-discovery agents you want without conflating them with training crawlers or user-directed fetchers.

2. Exact relevance to the buyer question

Controlled research

How to test: Compare whether the opening answer resolves the exact prompt with the same entities, constraints, and decision criteria.

Action: Write one self-contained answer for the real question, then support it with evidence and deeper detail.

3. Specific, current, checkable facts

Controlled research and retrieval practice

How to test: Inventory prices, plan boundaries, dates, specifications, named capabilities, and source links that an answer can verify.

Action: Replace adjectives with facts. Show the price, scope, date, method, limit, or comparison boundary.

4. Extractable evidence and complete answers

Peer-reviewed and large-sample research

How to test: Check whether definitions, numbers, procedures, and comparisons stand on their own when lifted from the page.

Action: Use descriptive headings, complete paragraphs, tables for repeated fields, and citations beside the claims they support.

5. Independent corroboration on cited source layers

Foglift citation benchmarks

How to test: Classify current winning URLs as vendor-owned, editorial, marketplace, review, community, institutional, or video sources.

Action: Improve an owned page when vendor pages win. Pursue a specific earned mention when independent sources own the answer.

6. Consistent entity and product identity

Operational retrieval requirement

How to test: Compare the brand name, category, description, pricing, and capabilities across the site, directories, profiles, and reviews.

Action: Give retrieval systems one current source-of-truth description and correct material conflicts on independent records.

7. Freshness where the question requires it

Controlled research and query intent

How to test: Check whether the answer depends on current pricing, availability, product features, regulations, or dated research.

Action: Publish a meaningful modified date only after verifying and changing the facts that can go stale.

8. Text availability and semantic page structure

Provider documented

How to test: Inspect server-rendered HTML for the main answer, heading hierarchy, internal links, table labels, and accessible text.

Action: Keep essential claims in crawlable text and make the page useful without requiring a client-side interaction.

9. Structured data that matches visible content

Google documented boundary

How to test: Validate supported JSON-LD and compare every marked-up answer, price, author, and date with visible page content.

Action: Use schema for eligible Search features and entity clarity. Do not claim that schema guarantees an AI citation.

10. Repeated multi-engine measurement

Foglift source-divergence research

How to test: Run the same buyer prompt across all target engines and store answer, citations, competitors, sentiment, and time.

Action: Judge changes by repeated mention and intended-page citation movement, not a single favorable answer.

Factor 1: Retrieval Eligibility Uses Different Agents by Provider

Search visibility depends on the agent that supports search. OpenAI tells publishers to allow OAI-SearchBot when they want content discovered, surfaced, cited, and linked in ChatGPT search. GPTBot is the separate control for content that may contribute to model training. Blocking GPTBot does not express the same choice as blocking OAI-SearchBot.

Anthropic documents three roles. Claude-SearchBot supports search-result quality. Claude-User retrieves pages for a person's request. ClaudeBot concerns content that may contribute to future model training. Perplexity similarly separates PerplexityBot, which surfaces and links sites in search results, from Perplexity-User, which supports user-triggered retrieval and generally ignores robots.txt.

Robots.txt is only one layer. A CDN challenge, WAF rule, IP block, authentication wall, or client-only rendering failure can prevent retrieval after the file says allow. Verify the actual HTTP response and rendered text. Perplexity publishes IP ranges for both agents; matching a user agent and a published IP range is stronger evidence than trusting the header alone.

Factors 2 to 4: Relevance, Facts, and Extractable Evidence

The 2026 paper What Gets Cited: Competitive GEO in AI Answer Engines used a controlled two-document testbed across six language models and 252,000 trials. Each paired comparison changed one of 18 content factors while counterbalancing source order. Topical relevance and list position were the largest drivers of the first citation. Explicit price information and a recent timestamp also helped consistently. Completeness and trust cues produced smaller gains, while formatting-only edits had little impact.

That finding leads to a useful editing rule: begin with the fact the buyer needs to decide. A page targeting "AI search monitoring tool with an API" should name the API, authentication method, plan boundary, rate limit, endpoint family, and a working request near the answer. A paragraph saying the product is developer friendly leaves no checkable fact to lift.

The original GEO paper by Aggarwal and coauthors tested strategies such as adding citations, quotations, and statistics against a benchmark of generative-engine responses. The reported gains apply to the paper's visibility metrics and test setting. They do not prove that every statistic, quote, or citation improves every production engine. Use the paper as evidence that source-ready facts can change generated visibility, then validate the effect on your own prompts.

Factors 5 and 6: Source Ownership and Entity Consistency

Foglift's Q2 citation-type benchmark classified 1,430 of 2,583 citations from 375 buyer-intent answers across five engines. In the classified subset, ChatGPT used vendor first-party sites for 68% of citations, while the other four engines ranged from 46% to 52%. Perplexity was the only engine to cite video at meaningful scale in that sample. Community sources were concentrated in Gemini and Google AI Overview rather than distributed evenly across all five.

The implication is operational. When a vendor guide or product page wins, improve the page and its evidence. When a review marketplace, publisher roundup, forum thread, or journalist article wins, more owned copy may not change the source layer. Pursue inclusion in the specific independent source the engine already trusts. A winner-page benchmark should classify ownership before it recommends a rewrite.

Entity consistency supports both paths. Keep the current brand name, category, product description, prices, engine coverage, and differentiators aligned across your site and independent records. Conflicting descriptions force the engine to reconcile which version is current. Correcting a stale category listing can be more important than adding another paragraph to a canonical page.

Factors 7 to 9: Freshness, Text Structure, and Schema Boundaries

Freshness matters when the question is time-sensitive. Pricing, plan limits, supported engines, regulations, availability, and product comparisons can become wrong quickly. A recent date helped in the paired-source citation experiment, but the right action is to verify the changing facts. Updating a timestamp without updating the evidence creates a freshness signal that the page cannot defend.

Keep the essential answer in server-rendered text with descriptive headings, accessible table labels, and internal links. This is a retrieval and comprehension safeguard rather than proof of a hidden formatting preference. Google explicitly says there is no ideal page length and no requirement to split content into tiny chunks for generative Search. Make sections as long as the reader needs to understand the decision.

Google also says no special schema.org markup is required for AI Overviews or AI Mode. Supported structured data remains useful for Search features when it matches visible text. Organization, Product, Article, BreadcrumbList, FAQPage, or ItemList markup should represent what a reader can see and what the applicable guidelines allow. Schema is a consistency contract. It is not a guaranteed AI-citation switch.

Factor 10: Multi-Engine Measurement Is Part of the Optimization

In Foglift's June 2026 source-divergence study, 62 query groups had answers from all five engines. Gemini and Google AI Overview showed 98.4% agreement on whether a brand belonged in the answer and 0.643 average citation-domain Jaccard overlap. ChatGPT and Claude showed 85.5% brand-mention agreement but only 0.027 average citation-domain overlap. Engines can make a similar brand decision while relying on very different sources.

A single-engine check cannot tell you whether the change generalized. Freeze a buyer-prompt panel and run it on a repeatable cadence. Store failures as failures rather than negative answers. Compare completed runs using mention rate, intended-page citation rate, competitor pressure, answer position, and sentiment. Keep prompt wording and engine selection stable enough that the deploy is the main changed variable.

AI Search vs Traditional SEO: Key Differences

SEO and AI visibility share a technical foundation, especially inside Google Search. Google states that AI Overviews and AI Mode use the Search index and core ranking systems. A supporting page must be indexed and eligible to appear with a snippet. Crawling, internal links, text availability, page experience, helpful content, and accurate structured data remain relevant.

The measurement unit changes after that foundation. Traditional rank tracking asks where a URL appears for a query. AI Visibility asks whether an answer names the brand, which page it cites, how the brand is framed, which competitors co-occur, and whether the outcome repeats. One answer can cite several URLs, mention several brands, or name a brand without linking its site.

FactorGoogle SEOAI Search (GEO)Source
EligibilityIndexed and snippet-eligibleProvider-specific search and retrieval accessProvider documentation
Primary outcomeURL position and clicksBrand mention, citation, framing, and source ownershipMeasurement design
Query fitDocument relevanceSource relevance plus answer-level usefulnessControlled citation research
Structured dataSupported rich-result eligibilityNo special Google AI schema requirementGoogle Search Central
Crawler controlsGooglebotOAI-SearchBot, Claude-SearchBot, PerplexityBot, and othersProvider documentation
Third-party evidenceLinks, mentions, and result diversityCan become the answer's cited corroboration layerFoglift citation benchmarks
FreshnessDepends on query intentUseful for current facts; recent dates helped in controlled testsGoogle and competitive GEO research
Cross-engine consistencyOne Search ecosystemSource pools and answer behavior differ by engineFoglift source-divergence study
Test methodSearch Console and rank historyRepeated prompt panels with stored answer evidenceExperimental design

Ten-Step AI Search Optimization Workflow

Treat the workflow as an experiment, not a seven-day promise. Complete the steps in sequence, record the deploy date, and wait for enough repeated answers to distinguish movement from normal variance. A high-impact technical block can be fixed quickly. Independent source placement can take longer and should not be disguised as an owned-page task.

Step 2

Create a frozen panel of category, comparison, problem, and brand prompts

Step 3

Run the panel across ChatGPT, Claude, Perplexity, Gemini, and Google AI Overview

Step 4

Verify search crawler, CDN, WAF, status-code, and rendered-text access

Step 5

Pull cited URLs for every miss and classify each winner by source ownership

Step 6

Diff better vendor pages for direct answers, prices, specifications, evidence, dates, and structure

Step 7

Route independent-source wins to a specific earned-mention opportunity

Step 8

Correct entity, price, plan, category, and capability conflicts across current records

Step 9

Ship one testable change and record the page, prompt, engine set, and deploy date

Frequently Asked Questions

What are the most important AI search ranking factors?

The strongest practical levers are retrieval eligibility, relevance to the exact question, specific and current facts, extractable evidence, consistent entity information, and corroboration on the source types each engine already cites. No provider publishes a universal factor list or percentage weighting, so treat these as testable levers rather than a secret algorithm.

How is AI search ranking different from Google ranking?

Google AI Overviews and AI Mode use the Google Search index and core ranking systems, so SEO remains foundational there. Other answer engines operate their own search and retrieval paths. A page may be eligible in several engines while different source pools, prompts, and synthesis systems produce different citations or brand recommendations.

Does schema markup make AI engines cite a page?

No. Google says there is no special schema markup required for AI Overviews or AI Mode. Use supported structured data when it matches visible content and helps Search understand eligible page facts. Treat markup as a clarity and validation layer, not proof of an AI citation boost.

Which crawler should I allow for ChatGPT, Claude, and Perplexity search?

For search discovery, verify OAI-SearchBot, Claude-SearchBot, and PerplexityBot access. GPTBot and ClaudeBot concern potential model training rather than live search discovery. Claude-User and Perplexity-User support user-directed retrieval. Check robots.txt, CDN rules, WAF challenges, and provider-published IP ranges where available.

Can I influence what AI engines say about my business?

You can improve the evidence available to an engine, but you cannot guarantee an answer. Publish accurate product facts, answer buyer questions directly, keep pricing and capability boundaries current, earn independent corroboration, and measure a fixed prompt panel across engines before and after each change.

How should I measure an AI search optimization change?

Freeze the prompt wording, engine set, target page, and observation window. Record mention rate, intended-page citation rate, answer position when available, competitor co-occurrence, sentiment, and cited source ownership. Compare repeated panels because one answer can move through ordinary model and retrieval variance.

How to Benchmark a Winning Page Before You Edit

Start with the answer evidence, not a generic content checklist. Pull the URLs cited for the exact prompt and open every winner. Record the owner, page type, updated date, title, opening answer, named products, price and plan facts, comparison fields, source links, and structured data. Then map the best available page on your own site.

Bucket A is an off-page authority win. Examples include a review marketplace, independent roundup, trade publication, podcast transcript, community thread, or institutional guide. The page earns trust partly because it is independent of the vendor. The action is to correct or earn inclusion in that source. Rewriting your own page may still improve product truth, but it does not replace the independent placement.

Bucket B is a better-page win. A vendor guide, feature page, documentation page, or pricing page may win because it answers the question more completely. Build a concrete diff. Does the opening resolve the exact prompt? Does it name current prices and limits? Does it show a working example? Are claims sourced? Is the modified date defensible? Does the visible content match the schema? Beat those facts and structures on the existing canonical.

Some prompts have mixed winners. Split the work. Improve the owned citation target and pursue the highest-value independent source. Keep separate success criteria for each: intended-page citations for the owned change, and brand inclusion or accurate capability language for the earned placement.

A Release QA Checklist for AI Search Pages

Content contract

  • The title and opening answer match the buyer's question.
  • Prices, plan limits, engines, dates, and product claims match the current source of truth.
  • Every statistic has a named source, sample, window, and bounded conclusion.
  • Visible FAQs and FAQPage answers share one source array.
  • The page gives a favorable, specific, verifiable Foglift fact where the reader is choosing a tool.

Technical contract

  • The canonical URL, Article dateModified, visible updated date, and blog listing agree.
  • Important answers appear in server-rendered text.
  • Internal links point to active canonicals rather than retiring siblings.
  • Article, FAQPage, and ItemList JSON-LD validate and match visible content.
  • The build, focused tests, voice sweep, and self-loss safety checks pass.

Diagnose the Result Before Choosing the Fix

The same prompt can fail at different layers. Use the matrix below to identify the first observable failure. Fixing a later layer while an earlier one is broken creates activity without changing the answer.

Observed resultLikely layerNext evidence to collectBest next action
Page cannot be fetchedTechnical eligibilityCrawler-specific status, robots decision, CDN and WAF logs, rendered textRepair the exact access failure and rescan
Page is retrieved but never citedRelevance or evidenceWinner-page opening, facts, prices, dates, sources, and answer coverageImprove the existing canonical against the concrete diff
Independent pages dominate citationsSource authorityPublisher, marketplace, review, community, or institutional ownershipEarn or correct a placement in the recurring source
Brand is named but the site is not citedEntity associationPages cited beside the mention and their version of the brand factsAlign first-party facts and strengthen the corroborating source layer
Site is cited but the brand is omittedCitation absorptionWhich claims or passages the answer used from the pageMake the brand-to-capability relationship explicit in the cited section
Brand appears in one engine onlyEngine-specific retrievalCited-domain overlap and winner ownership by engineBuild a separate source plan for each weak engine
One favorable run does not repeatVariance or weak evidenceRepeated completed panels with unchanged prompts and engine setWait for a defined sample before claiming movement
Competitor is recommended from your pagePage framingMost extractable comparison sentence, table, and opening answerRewrite the page to state specific favorable Foglift differences

Keep technical readiness and AI Visibility on separate scorecards. A Technical Audit can confirm that the page is accessible, structured, secure, performant, and usable. It cannot prove that a production answer engine selected the page for a particular prompt. A prompt panel can observe a citation or brand mention. It cannot prove which single page attribute caused the answer.

Foglift connects those evidence layers in one workflow. The Technical Audit diagnoses the page. AI Visibility stores prompts, answers, citations, competitors, sentiment, and time. The recommendation queue turns current gaps into actions. AI Crawler Analytics and referral tracking show observed visits from supported agents and recognized AI referrers. API, CLI, MCP, and webhook surfaces let developer-led teams move the same evidence into their own systems.

Define Success Before the Page Ships

Choose a primary outcome that matches the job of the page. A research report usually targets citations to the report. A conversion page targets accurate favorable brand recommendations and the intended product URL. A developer guide may target both a citation and a correct description of the API or integration. Record the target before editing so a convenient secondary metric cannot replace the original goal.

Set the denominator explicitly. Mention rate is brand-positive completed answers divided by completed answers. Intended-page citation rate is completed answers citing the target URL divided by completed answers. Do not count provider errors, timeouts, or stored placeholders as brand misses. Keep engine-level rates beside the combined rate because a strong Google result can hide a weak Claude or Perplexity result.

Use a dated hold after release. Indexing, crawling, and scheduled prompt cadence can delay observable movement. During the hold, resist repeated rewrites that erase the experiment. Reopen the page when a later complete panel still misses, a winner set reveals a new better-page gap, a product fact changes, or an independent source opportunity becomes available.

Key Research Data Points

These findings are bounded to their documented samples. Provider statements define product behavior. Controlled studies test content factors. Foglift benchmarks describe observed citation patterns across their stated panels.

FindingImpactSource
Google generative Search eligibilityIndexed and snippet-eligible; no additional technical requirement or special schemaGoogle Search Central
Competitive paired-source citation test252,000 trials across six models; relevance and list position led; price and recent dates helpedVishwakarma et al. (2026)
Citation selection and absorption sample602 prompts, 21,143 search-layer citations, 18,151 fetched pages, 72 featuresZhang et al. (2026)
Original GEO benchmarkContent strategies changed visibility by up to 40% on the paper's metrics and benchmarkAggarwal et al. (KDD 2024)
Foglift Q2 citation-type panel375 answers, 2,583 citations, 1,430 structurally classified citationsFoglift Research (2026)
Vendor first-party share in classified subsetChatGPT 68%; other four engines 46% to 52%Foglift Research (2026)
Foglift cross-engine source panel1,373 answers; 62 query groups answered by all five enginesFoglift Research (2026)
Closest source pairGemini and Google AI Overview: 0.643 citation-domain JaccardFoglift Research (2026)
Lowest-overlap pair in the panelChatGPT and Claude: 0.027 citation-domain JaccardFoglift Research (2026)
Cross-engine domain exclusivityOnly 12 of 1,119 cited domains appeared across all five enginesFoglift Q2 Citation Benchmark

Sources & Further Reading

  1. Google Search Central: AI features and your website. Official eligibility, indexing, snippet, content, structured-data, and preview-control boundaries for AI Overviews and AI Mode.
  2. Google Search Central: Optimizing for generative AI features. Official myth-busting guidance covering SEO foundations, page length, chunking, schema, and authentic mentions.
  3. OpenAI: Publishers and Developers FAQ. Official OAI-SearchBot discovery and citation guidance plus the separate GPTBot training control.
  4. Anthropic: web crawler controls. Official roles for Claude-SearchBot, Claude-User, and ClaudeBot.
  5. Perplexity: crawler documentation. Official roles, robots behavior, WAF guidance, and published IP sources for PerplexityBot and Perplexity-User.
  6. Aggarwal et al.: GEO: Generative Engine Optimization. KDD 2024 paper introducing GEO-bench and content-strategy experiments.
  7. Vishwakarma, Kumar, and Jamidar: What Gets Cited. A 2026 controlled paired-source program covering 252,000 trials across six models.
  8. Zhang, He, and Yao: From Citation Selection to Citation Absorption. A 2026 measurement framework over 602 prompts and 21,143 search-layer citations.
  9. Foglift Research: Five AI Engines, Five Content Diets. A structural classification of 1,430 citations from the Q2 2026 benchmark.
  10. Foglift Research: AI Engines Agree on Brands More Than Sources. Five-engine mention agreement and citation-domain overlap across 62 complete query groups.
  11. Foglift Research: The AI Citation Map. Four authority patterns derived from 375 buyer-intent responses across 25 industries and five engines.

Separate readiness from real answer visibility

Run an unlimited single-page Technical Audit, then track mentions, citations, competitors, and sentiment across the engines your buyers use.

Run Free Technical Audit

Related Articles

Related: Learn about AEO (Answer Engine Optimization), the framework for making your content extractable by AI answer engines.

Fundamentals: Learn about GEO (Generative Engine Optimization) and AEO (Answer Engine Optimization) (the two frameworks for optimizing your content for AI search engines).

Related reading

Free tool

Run a free Technical Audit for your AI Readiness Score

Audit any URL in 30 seconds. See scores for SEO, AI Readiness, performance, security, and accessibility.

Free Technical Audit

No signup required. Results in 30 seconds.