What Does a GEO Score Measure? Prompts, Models, and Judges Change the Number (2026)
Same questions, same day, only the model changed: our brand mention rate was 43% versus 5%. A September 2026 methods paper explains why GEO scores depend on prompt sets, execution conditions, weights, and judges, and what to check in an AI visibility report.
AI-assisted drafting · reviewed by the RanketAI Editorial Team — Editorial policy ›
Key takeaway: Before comparing GEO scores, check the conditions that produced them. A methods paper published on September 6, 2026 (Martinez, arXiv:2609.06811, preprint) lays out how a GEO visibility score is built from prompt sets, prompt wording, execution conditions, weights, and scoring rules, and notes that changing only the judge instruction can change the score of an identical answer. When RanketAI asked the same set of questions on the same day and changed only the model, our own brand mention rate came out at 43% versus 5% in ChatGPT and 74% versus 40% in Gemini. With GPT-6 just released, here is how to check the conditions behind a score.
Three-line summary
- A GEO score is not a single fact about ; it is the output of a measurement design. Which questions were collected, which model was asked and when, and what counted as a mention or citation all feed into the number.
- In our own test, changing nothing but the model moved the mention rate several-fold. So the measurement model should be judged not by how new it is but by whether it is the model the users you want to represent actually see. GPT-6 Sol and Luna launched on September 22, 2026, yet on September 24 the default model in free users' regular chat was GPT-5.6 Luna.
- The paper is a methods survey with no new experiments, so its numbers are not evidence. What it offers is a framework for what to ask when you evaluate an AI visibility tool or report.
Same questions, same day — only the model changed, and the mention rate moved
We checked how much measurement conditions move a score in our own data first. On September 4, 2026, while reviewing our measurement conditions, RanketAI ran a set of category questions that did not include any brand name on the same day, changing only the model, over repeated rounds. The target was our own brand, in the AI search visibility diagnostics category.
| AI service | Model condition | Our brand mention rate |
|---|---|---|
| ChatGPT | Model A | 43% |
| ChatGPT | Model B (lighter model) | 5% |
| Gemini | Version A | 74% |
| Gemini | Version B (newer version) | 40% |
In ChatGPT, the gap moved together with search depth. Model A averaged 4.3 web searches per answer and model B 1.2, and every answer that mentioned our brand also cited one of our pages. One reasonable reading is that deeper search checks more sources and widens the candidate list accordingly.
The Gemini result makes a different point. Moving to a newer version does not guarantee a higher score either. In this test, the earlier version produced the higher mention rate. Both numbers are real AI answers, so the table alone cannot say which one is "right" — what the measurement is meant to represent has to be settled first.
The limits are clear. One brand (ours), one category, a small sample, one day. The magnitudes should not be generalized to other brands or categories; what the table shows is direction, not size: the choice of measurement model alone can move the same brand's score several-fold.
The five components of a GEO score, according to the paper
The methods paper published on September 6, 2026 turns this phenomenon into a framework. Its author, Olivier Martinez, is the researcher who in July published a critical survey grading 45 GEO studies by strength of evidence. This paper steps back from "do GEO techniques work?" to ask "what does a GEO score measure?"
The paper splits the instrument that produces a GEO visibility score into five components.
| Component | What it sets | What changes when it changes |
|---|---|---|
| Situation annotation | The usage situations the questions represent, and the rules for classifying them | The same score means something different depending on which situations were measured |
| Prompt formulations | How each situation is phrased as a question | Wording alters retrieval, the competing sources, and the answer itself |
| Execution conditions | Model and version, date and time, interface, prior conversation, parameters | The same question produces a different answer when the model changes (see our test above) |
| Weights | The relative importance of each situation | Weights can flip the ranking of two brands |
| Scoring rules | How mentions and citations are identified in the answer; for an LLM judge, its model, rubric, and instruction | The same answer can score differently when the judge instruction changes |
The point of the table is simple: a different score does not by itself mean different visibility. If even one of the five components differs, the two scores may be measuring different things.
The "answer market" — your question set defines what you measure
The paper's central concept is the "answer market." The questions used for measurement, together with their weights, make up the market being measured.
"We call this weighted set of answer opportunities an “answer market.” This constructed market does not necessarily represent actual user demand." — Martinez, Measuring GEO Visibility, arXiv:2609.06811 (2026-09-06)
In the same passage, the author argues that a sentence like "a source is cited in 40% of answers" is not enough on its own. To interpret it, you also need to know whether those answers responded to open-ended requests, to prompts naming the source, or to comparison requests — and which engine and period they came from.
Two problems covered earlier on this blog fit under this one concept. The brand-name prompt post showed that prompts containing the brand name almost always produce a mention in web-search-enabled AI, so they carry no information as a performance metric. The demand-side data post showed that there is no data to check how close measurement questions are to real purchase-evaluation questions. Both are cases where the answer market drifts away from real demand.
Execution conditions — a model's launch date is not its default date
The model is the execution condition that changes most often. OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026. The announcement makes clear, however, that a launch and a change of default model are not the same event.
"GPT‑6 Sol and Luna roll out today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users." — OpenAI Developer Community announcement (2026-09-22)
| Surface (checked 2026-09-24) | Model provided |
|---|---|
| Default model in ChatGPT free users' regular chat | GPT-5.6 Luna |
| Free and Go users, desktop app trial | GPT-6 Luna |
| Paid users (Plus, Pro, Business, Enterprise, Edu) in ChatGPT Work and Codex | GPT-6 Sol, GPT-6 Luna |
GPT-5.6 Luna's availability to free users comes from OpenAI's announcement; the regular-chat default as of September 24 was checked directly in a free account. We will update this table as models change.
When choosing a measurement model, decide first whose answers you want to represent. If the measurement is meant to represent what free users see, the GPT-6 launch itself is no reason to change conditions; the point to change is the day the default model in free users' regular chat changes. If it is meant to represent users in paid work environments, the reference point is different. Either way, what to check is whether the tool has defined its reference point and whether it records when that reference point changed.
The paper is specific about how to read scores after a model change, too.
"A stable identifier does not guarantee the absence of silent updates: interleave conditions, use sentinel controls, and segment collection if a break is suspected." — Martinez, arXiv:2609.06811, measurement protocol step 4
In practice, scores from before and after a default-model change are better not read as one continuous trend line. Record when the change happened and treat the next few rounds as a new baseline; only then can before and after be compared.
Date is an execution condition as well, because answers differ from run to run even on the same model. In our re-measurement benchmark, measuring 12 Korean B2B SaaS companies twice on the same day produced a median individual score change of 6 points, while the median change over a 20-day interval was 5.5 points — the month's change sat inside same-day re-measurement noise. How many repetitions it takes before a result can be trusted is covered in our St. Gallen study analysis.
Scoring — the same answer can score differently under a different judge instruction
Receiving the answer does not settle the score yet. A scoring step still has to identify brand mentions and citations in the answer, and if that step is handed to a language model, another instruction enters the measurement.
"This instruction can change the scoring even when the evaluated answers remain identical." — Martinez, arXiv:2609.06811
The paper proposes three safeguards: if an LLM judge is used, declare its model, rubric, and instruction; keep the raw answer text and the coding rule; and cross-check by applying a different judge to the same saved answers. From a user's side, this reduces to one question: when a score changes, can you tell whether the answer changed or the judging changed? That requires being able to see the raw answers alongside the results.
Weights — reporting a range instead of a single score
A composite score hides situation weights. The paper gives an example where the choice of weight alone flips a ranking. The author labels it a fictional example; discovery situations are questions where users look for candidates for the first time, and verification situations are questions checking something they already know.
| Discovery success rate | Verification success rate | Discovery weight 0.2 | Discovery weight 0.8 | |
|---|---|---|---|---|
| System A | 0.30 | 0.90 | 0.78 | 0.42 |
| System B | 0.50 | 0.60 | 0.58 | 0.52 |
A leads when the discovery weight is below 0.6, and B leads above it. If the weight cannot be pinned down anywhere between 0.2 and 0.8, the paper recommends reporting the score as a range: 0.42–0.78 for A and 0.52–0.58 for B. A single score hides the fact that A swings widely with the weight while B stays stable.
The practical question is whether the composite score's weights match your goal. If the goal is winning new customers, discovery situations — where users who do not know your brand look for candidates — should carry more weight; if the goal is retaining existing customers, verification situations can reasonably weigh more.
Being cited is not the same as contributing
"Our citation count went up" and "that source contributed to the answer" are different claims.
"The presence of a citation is insufficient to establish a source’s contribution to an answer." — Martinez, arXiv:2609.06811
To examine contribution, the paper proposes fixing the question, the conditions, and the list of retrieved sources, removing one source, regenerating the answer, and comparing the two. In that design, "was it cited?" cannot be the outcome measure, because removing a source mechanically removes its citation. What has to be compared is whether the information in the answer changed.
This distinction points the same way as two earlier observations on this blog. The ghost citations post covered cases where a page is cited but the brand is never mentioned, and the critical survey analysis covered how being selected as answer material and being absorbed into the answer body are separate problems.
Seven things to check when evaluating a measurement tool or report
Turning the paper's seven-step measurement protocol into questions a user can ask gives the list below. Question-design criteria were covered in the brand-name prompt post, so this list focuses on the other conditions.
| Question to ask | Why it matters |
|---|---|
| Which situations does the question set represent? Does it include prompts containing the brand name? | The question set is the market being measured |
| Does the measurement model match the users you want to represent (free or paid)? | Changing only the model moved the mention rate several-fold |
| Is the date of a default-model change recorded? | Scores before and after a change cannot be read as one trend line |
| How many times, and over how long, was each question repeated? | Same-day re-measurement noise was about as large as a month's change |
| Is the mention and citation scoring method documented and held fixed? | Changing only the judge instruction can change the score |
| Do the composite score's weights match your goal? | Weights can flip rankings |
| Is a rising citation count being read as performance? | A citation is not evidence of contribution |
RanketAI's tools follow the same order. A page structure check first confirms that AI crawlers can read your pages, and AI brand visibility analysis then measures mentions and citations in ChatGPT, Perplexity, and Gemini answers to questions that do not include your brand name. To see competing brands side by side under the same conditions, use competitor comparison. Whatever tool you use, the seven questions above apply the same way.
Three lines to hold when citing this paper
The paper's value lies in its framework for measurement design, not in numbers. Cite it with these limits attached.
- There are no new experiments. The author states: "No new experimental collection from users or generative systems is reported." The ranking-reversal example is explicitly labeled fictional. Quoting its numbers as real measurement results would be a factual error.
- It is a single-author preprint that has not passed peer review. The empirical validity of the proposed protocol is also left for future work.
- It is a selective, predominantly English-language survey. The author notes that the same vocabulary appearing across studies is not an accumulation of independent replications. There is no evidence yet that it transfers unchanged to Korean-language questions or Korean AI services.
Our own measurement deserves the same reading. The table above comes from one brand, one category, a small sample, and one day; it shows the direction in which the measurement model moves a score, not which model is better.
FAQ
Isn't measuring with the newest model more accurate?▾
It depends on what "accurate" means. If the goal is the answers a particular group of users sees, the model that group uses is the reference. In our test, the newer Gemini version actually produced a lower mention rate than the earlier one (40% versus 74%). A new model shows you that model's answers; there is no guarantee they match what the users you want to represent see.
If the measurement model changes, can I still compare against earlier scores?▾
It is safer to read before and after as separate segments. The paper itself recommends segmenting collection when a break is suspected. Treat the first few rounds after the change as a new baseline, and read changes as a trend only from there.
Different tools give different scores — which one should I trust?▾
A difference in scores is not a red flag in itself. If the question sets, measurement models, scoring methods, or weights differ, each tool may be measuring something different. To compare two tools, ask each of them the seven questions above and compare only the parts where conditions match.
We have little budget — can we measure it ourselves?▾
Yes, as long as you record the conditions so the results can be interpreted. At minimum, save the exact question, date and time, the app and plan used (free or paid), the model name shown on screen, and the full answer, and repeat under the same conditions several times.
Our citation count went up — does that mean it worked?▾
Citation counts alone cannot tell you that, because a citation is not evidence of contribution. Check together whether the way the answer describes your brand actually changed, and whether you now appear in the recommendation list, before calling it a result.
Related reading
- The 40% GEO Gain Has a Condition — Inside the 2026 Critical Survey of 45 Studies — the same author's earlier survey, on the conditions behind technique effects.
- Why Brand-Name Prompts Can't Measure AI Visibility — criteria for designing the question set.
- Monthly Change Was Smaller Than Same-Day Measurement Noise — measured variance from date as an execution condition.
- Don't Measure AI Visibility Once: Why You Need 7 Runs a Day and 2–4 Week Windows — the statistical baseline for repeated measurement.
- The Demand Side AI Visibility Misses — another side of the answer market drifting from real demand.
Execution Summary
| Item | Practical guideline |
|---|---|
| Core topic | What Does a GEO Score Measure? Prompts, Models, and Judges Change the Number (2026) |
| Best fit | Prioritize for geo workflows |
| Primary action | Standardize an input contract (objective, audience, sources, output format) |
| Risk check | Validate unsupported claims, policy violations, and format compliance |
| Next step | Store failures as reusable patterns to reduce repeat issues |
Data Basis
- Primary source: Olivier Martinez, "Measuring GEO Visibility: Prompt Corpora Define the Answer Market" (arXiv:2609.06811, 2026-09-06), full text — the five components of the measurement instrument (situation annotation, prompt formulations, execution conditions, weights, scoring rules), the definition of the "answer market," the judge-instruction effect, reporting admissible score ranges, the distinction between citation and contribution, and the seven-step measurement protocol are cited from the original. Every English quotation was checked character by character against the paper's HTML. It is a methods survey with no new experiments, and the ranking-reversal example is explicitly fictional. The author is the same researcher behind the 2026-07-15 GEO critical survey (arXiv:2607.14035).
- RanketAI measurement: on 2026-09-04, while reviewing our measurement conditions, we ran the same set of category questions (no brand name included) on the same day, changing only the model, over repeated rounds (target: our own brand; category: AI search visibility diagnostics). Mention rate is the share of completed answers that mention our brand. Single brand, single category, small sample, one day — the magnitudes are not generalized, and model identifiers and question counts are not disclosed.
- Model availability: the scope of GPT-6 Sol and Luna was checked against the OpenAI Developer Community announcement (2026-09-22) and OpenAI's announcement expanding GPT-5.6 Luna to free users. The default model in a free account's regular chat on 2026-09-24 was checked directly.
- Repeated-measurement and same-day re-measurement figures are quoted from existing posts on this blog (the 2026-07-12 St. Gallen study analysis and the 2026-08-12 re-measurement benchmark), with details linked internally.
Key Claims and Sources
This section maps key claims to their supporting sources one by one for fast verification. Review each claim together with its original reference link below.
Claim:With the same category questions asked on the same day, our brand mention rate in ChatGPT was 43% on model A and 5% on the lighter model B, with an average of 4.3 versus 1.2 web searches per answer
Source:RanketAI measurement (2026-09-04, first published in this post)Claim:Under the same conditions, our brand mention rate in Gemini was 74% on the earlier version and 40% on the newer version
Source:RanketAI measurement (2026-09-04, first published in this post)Claim:The weighted set of answer opportunities produced by prompt selection and weights (the answer market) does not necessarily represent actual user demand
Source:Martinez (arXiv:2609.06811, preprint)Claim:When answer scoring is assigned to a language model, the judge instruction can change the score even when the evaluated answers remain identical
Source:Martinez (arXiv:2609.06811, preprint)Claim:In a fictional example with discovery and verification success rates of 0.30 and 0.90 for A and 0.50 and 0.60 for B, A leads at 0.78 with a discovery weight of 0.2, B leads at 0.52 with a weight of 0.8, and the ranking flips at 0.6
Source:Martinez (arXiv:2609.06811, fictional example)Claim:GPT-6 Sol and Luna began rolling out on September 22, 2026 in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, while Free and Go users can try GPT-6 Luna in the desktop app
Source:OpenAI Developer Community announcement (2026-09-22)Claim:Measuring 12 Korean B2B SaaS companies twice on the same day produced a median individual score change of 6 points, while the median change over a 20-day interval was 5.5 points
Source:RanketAI re-measurement benchmark (2026-08-12)
External References
The links below are original sources directly used for the claims and numbers in this post. Checking source context reduces interpretation gaps and speeds up re-validation.
- Martinez — Measuring GEO Visibility: Prompt Corpora Define the Answer Market (arXiv:2609.06811, 2026-09-06, preprint)
- OpenAI Developer Community — Announcing GPT-6 Sol and GPT-6 Luna in the API, Codex and ChatGPT (2026-09-22)
- OpenAI — Introducing GPT-6 Sol and Luna
- OpenAI — Improving GPT-5.6 Sol in ChatGPT and expanding access to GPT-5.6 Luna for free users
- Martinez — Optimizing Visibility in Generative Engines: A Critical Survey of GEO (arXiv:2607.14035, 2026-07-15, preprint)
Is your site visible in AI search?
See for free how ChatGPT, Perplexity, and Gemini describe your brand.
Start Free Diagnosis →Keep reading
Related Posts
These related posts are selected to help validate the same decision criteria in different contexts. Read them in order below to broaden comparison perspectives.
The Demand Side AI Visibility Misses — Which Sources Your Buyers Trust
AI visibility measurement checks one direction: whether your pages got cited. An August 2026 preprint releases a million-persona buyer corpus arguing the missing half is which sources buyers trust. Here is the axis you can apply today, and its limits.
Perplexity Never Skips the Web — What the Answer Stream Reveals About Citations (2026)
A stream-level observation of Perplexity: all seven test queries triggered a live web search, yet only 5 to 9 of the 10 to 15 retrieved sources earned a citation. Local cited place entities, how-to cited video, comparison cited the vendor's own page.
The 40% GEO Gain Has a Condition — Inside the 2026 Critical Survey of 45 Studies
The famous claim that GEO lifts AI-answer visibility by up to 40% was measured with the source already inside the answer context. A 2026 critical survey grades 45 GEO studies (2023-2026) and shows what reproduces, what backfires, and what to measure first.
Google Preferred Sources: The One AI Visibility Lever Readers Hand You
Google extended Preferred Sources into AI Overviews and AI Mode in every supported language. Designated sites get labeled inside AI answers, and Google reports twice the click-through. What the docs say about eligibility, setup, and that figure's limits.
Pew: 10% of the Web Is Written With AI — Authorship Detection Is Not GEO Detection
Pew Research analyzed 490,000 webpages and found AI authorship signs on 10% of the July 2026 crawl, and over a third among post-ChatGPT dated pages. Authorship detection and GEO detection answer different questions — neither decides whether you get cited.