Skip to main content
Guides geo research papers

GEO Research, Explained: What the Papers Behind AI-Visibility Scoring Actually Found

A plain-language walkthrough of the peer-reviewed research behind Generative Engine Optimization — what was tested, found, and what wasn't.

·

Two academic papers currently anchor most of what's testable about Generative Engine Optimization: a Princeton-affiliated study that measured which content changes make a passage more likely to be used in an AI-generated answer (KDD 2024), and a newer automated-optimization system that builds on it (AutoGEO, arXiv 2025). This guide reads both directly — not secondhand summaries — and reports exactly what they tested, what they found, and where the evidence stops.

Why bother reading the research in a niche full of hype

"AI SEO" and "GEO" attract a lot of confident-sounding advice with no citation behind it: percentages, rankings, and "AI optimization" claims that trace back to nothing. The honest response isn't to add more unsourced claims — it's to go read the two papers that most of this category's tooling, including GeoReady's own scoring engine, actually points to, and say plainly what's in them and what isn't. That's what this page does. If you want the practical difference between "getting placed in the sources a model reads" and "making your own site retrievable," see the companion guide, LLM seeding and GEO , published the same day as this one.

Paper 1: GEO — Generative Engine Optimization (Aggarwal et al., KDD 2024)

What was tested. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande built GEO-Bench, a benchmark of real-world-style queries, and tested a set of concrete content-level rewrite methods against it — adding quotations, adding statistics, citing sources, optimizing for fluency, using more authoritative language, using more technical terms, using easier-to-understand language, increasing unique-word variety, and, as a deliberately old-fashioned comparison point, keyword stuffing. Each method's effect was measured as a relative change in how much a page's content was drawn on in the generated answer, compared to the unmodified baseline.

What was found. The paper's own results table shows the methods split clearly into two groups:

  • Quotation Addition — Relative lift over baseline: ≈ +41%
  • Statistics Addition — Relative lift over baseline: ≈ +33%
  • Fluency Optimization — Relative lift over baseline: ≈ +29%
  • Cite Sources — Relative lift over baseline: ≈ +28%
  • Technical Terms / Authoritative / Easy-to-Understand language — Relative lift over baseline: mid-range, smaller lift
  • Unique Words — Relative lift over baseline: smaller lift
  • Keyword Stuffing — Relative lift over baseline: underperformed the baseline

The paper describes its three strongest methods — Cite Sources, Quotation Addition, and Statistics Addition — as achieving a relative improvement in the 30–40% range on its visibility metric.

What it means for practitioners. The result that matters most isn't the top of the list — it's the bottom. Keyword stuffing, one of the oldest tricks in classic SEO, actively hurt performance in this test. That's a real, testable finding, not an assumption carried over from Google-era habits: for these generative-answer systems, adding real quotations, real statistics, and real citations to sources outperformed volume-based, low-substance tactics. This is directly relevant to what a GEO audit checks — GeoReady's content-quality and factual-density signal categories draw on exactly this kind of finding: does a page contain quotable statements, sourced statistics, and citations, rather than just more words.

Paper 2: AutoGEO (Wu, Zhong, Kim, and Xiong, arXiv 2025)

What was tested. This paper — formally titled "What Generative Search Engines Like and How to Optimize Web Content Cooperatively," and referred to by its system name, AutoGEO — takes the Princeton work a step further: instead of hand-picking a rewrite method, AutoGEO automatically learns a generative engine's content preferences by prompting frontier LLMs to explain those preferences, extracts reusable rules from the explanations, and applies them two ways — as prompt-based guidance for a larger system (AutoGEO-API) and as training rewards for a smaller, cheaper model (AutoGEO-Mini). Both were tested on the original GEO-Bench plus two newly built benchmarks using real user queries.

What was found. AutoGEO-API's automated rewriting beat the strongest single hand-designed method from the Princeton paper, Fluency Optimization, by up to 50.99% on one benchmark/metric combination (Researchy-GEO, "Overall"), with an average improvement across the paper's GEO metrics of about 36%. The smaller, cost-efficient AutoGEO-Mini model achieved an average improvement of about 21% while running at roughly 1/140th (≈0.0071×) the computational cost of the API-based version.

One caution when you see this paper cited elsewhere: AutoGEO is sometimes described as "ICLR 2026" with a flat "+50.99% improvement." As of September 2026 the paper is an arXiv preprint with no conference acceptance listed, and 50.99% is its best single-benchmark result — the representative figures are the ~36% and ~21% average improvements above.

What it means for practitioners. This is a research prototype, not a product you can buy or run today — there's no indication it's been independently replicated or peer-reviewed at a venue yet. The useful takeaway isn't "use AutoGEO," it's what the result implies: a systematically, automatically learned set of content rules outperformed even the best fixed, hand-picked method from the earlier study. That supports treating GEO as an evolving practice you periodically re-check against new evidence, rather than a fixed checklist you complete once.

Limits of the research — stated plainly

Both papers are genuinely useful, and both have real limits worth naming before anyone over-relies on them:

  • Benchmark, not the open web. Both papers evaluate against purpose-built benchmarks (GEO-Bench and, for AutoGEO, two newly constructed query sets) — not a random sample of the whole web, and not a measurement of any specific live production system's day-to-day behavior.
  • Lab conditions, not your site in production. The lift percentages describe how the tested setups responded to content changes under benchmark conditions. Production answer engines add retrieval ranking, safety filtering, freshness weighting, and frequent model updates that an academic benchmark can't fully reproduce — a technique that lifted a benchmark score may not transfer one-to-one to what ChatGPT, Perplexity, or Google's AI Overviews do with your specific page today.
  • Aggregate lift, not a guarantee for any one page. These are average or maximum effects across a benchmark's queries, not a promise that any individual page or query will move by a specific amount.
  • AutoGEO is very recent and not yet independently confirmed. Submitted to arXiv in October 2025, it hasn't (as far as could be verified for this guide) been independently replicated or confirmed as peer-reviewed at a conference. Treat its numbers as a promising early result, not settled science.

None of this means the findings are wrong — it means they're evidence, not certainty, and should be treated that way.

How GeoReady applies this research

GeoReady's scoring engine is grounded in this research, not a reproduction of either paper's exact experiment. The content-quality and factual-density categories in the 8-category GEO score reward the same kind of signal the Princeton paper found effective — quotable statements, sourced statistics, and cited claims — and the open-source scoring weights are published so you can check that connection yourself rather than take it on faith. For the full scoring breakdown and category weights, read the research explained in the methodology .

For the companion question — how "LLM seeding" (getting mentioned in third-party sources) differs from this readiness-focused research — see LLM seeding and GEO .

See how your own site measures up

Reading the research is the first step. Checking where your own site stands against it is the next one.

Run the free GEO audit

Apply this guide

Run an AI SEO audit before you change pages.

Use the audit to find which signal is holding the site back: crawler access, schema, llms.txt, content clarity, AI discovery, or entity strength.

  • Best for pages that need a technical and content baseline.
  • Next metric: AI readiness score plus the weakest signal category.

Frequently asked questions

Is GEO backed by real research, or is it marketing?

Both papers covered here are real, checkable academic work — one peer-reviewed at KDD 2024, one a 2025 preprint not yet confirmed at a conference venue. That's a genuine, if young, research base — stronger than most of the unsourced "AI SEO" advice circulating, but still early and narrower in scope than a mature field like classic information retrieval.

Did the research really find that keyword stuffing hurts AI visibility?

Yes — that's one of the more useful, counter-intuitive findings in the Princeton KDD 2024 paper: keyword stuffing underperformed the unmodified baseline in their test, while adding real quotations, statistics, and cited sources outperformed it.

Is AutoGEO a product I can use?

No — it's a research prototype described in an academic paper, not a commercial tool. Its result (automated content-rule learning beating a hand-picked method) is informative for how to think about GEO, not something to install.

Does GeoReady's score come directly from these papers' formulas?

No. GeoReady's 8-category scoring is its own engineered rubric, grounded in the same kind of findings these papers report (quotability, sourced statistics, structural signals), but it is not a line-by-line reproduction of either paper's experiment. See the methodology (/methodology/) page for exactly what GeoReady measures and how it's weighted.

Where can I read the actual papers?

Both are public: the Princeton GEO paper is on arXiv at arxiv.org/abs/2311.09735 (published at KDD 2024), and AutoGEO is at arxiv.org/abs/2510.11438 (arXiv preprint, October 2025).

Get the monthly State of GEO report

AI search readiness benchmarks, adoption stats, and the actions that move the needle — delivered monthly. No spam.

By submitting, you agree to receive the monthly GeoReady newsletter: benchmark data from the State of GEO dataset, practical GEO guidance, and product updates. You can unsubscribe anytime. See our Privacy Policy.