Monthly Change Was Smaller Than Same-Day Measurement Noise (August 2026 Benchmark)
We re-measured 12 Korean B2B SaaS domains 20 days apart and compared the result with two rounds run on the same day: month-over-month scores moved a median of 5.5 points versus 6 points within a single day. What to read as signal, and what to discard as noise.
This blog content may use AI tools for drafting and structuring, and is published after editorial review by the RanketAI Editorial Team.
Key takeaways: We re-measured 12 Korean B2B SaaS domains under identical conditions 20 days after the previous issue. Before interpreting the result, we first computed how far scores move when the same measurement runs twice on the same day. Same-day re-measurement moved individual scores by up to 17 points (median 6), and 2 of 12 domains changed grade. Month-over-month, individual scores moved a median of 5.5 points. In other words, a month of change was smaller than a single day of measurement variance. So the movement we report as real in this issue is not the individual scores — it is the composition of cited sources, which noise does not explain.
A problem that appeared before we could read the results
RanketAI published an anonymized aggregate of AI answer visibility for 12 Korean B2B SaaS domains in July 2026, and has now published the August 2026 issue, re-measured with the same sample under the same conditions. The sample was held fixed, and inputs and measurement conditions were unchanged. Only the measurement date differs.
The aggregate looks like this.
| Metric | 2026-07 | 2026-08 |
|---|---|---|
| Median score (p50) | 51 | 53 |
| Top 25% (p75) | 59 | 58 |
| Top 10% (p90) | 67 | 73 |
| Grades (A/B/C/D) | 5 / 5 / 0 / 2 | 6 / 4 / 0 / 2 |
The common reading is: "the median rose 2 points and one more domain reached grade A, so visibility improved." But writing that sentence requires answering a prior question. Are those 2 points a real change, or the normal variance of the measurement itself?
So we measured the variance first — same day, two rounds
This issue ran two rounds on the same day. Because the timing is identical, the difference between rounds is not "what happened over a month" but variance produced by the measurement itself.
| Same-day, two rounds | Value |
|---|---|
| Maximum individual score change | 17 points |
| Median individual score change | 6 points |
| Domains that changed grade | 2 of 12 |
The same brands, the same questions, the same day — and two domains still landed on a different grade. This is not a quirk of our sample. Researchers at the University of St. Gallen analyzed the repeatability of AI search visibility measurement and summarized the risk of single-run interpretation this way.
Visibility rankings are not stable, and a substantial share of the differences observed in a single measurement is explained by sampling variation rather than by real differences between brands. — Schulte, Bleeker & Kaufmann, Don't Measure Once: Measuring Visibility in AI Search (GEO), Univ. of St. Gallen, 2026 (arXiv:2604.07585)
Month-over-month change fell inside that band
With a noise baseline in hand, we compared July-to-August individual score changes against the same yardstick.
| Comparison | Median individual score change |
|---|---|
| Same day, two rounds (noise) | 6 points |
| July to August (20 days) | 5.5 points |
Individual movement over 20 days was smaller than movement within a single day. That is why this issue does not describe per-domain increases or decreases as trends, and why the report's limitations state the same. The 2-point median rise and the one additional grade A also fall inside that band, so we do not call them an improvement.
This does not mean the measurement is unreliable. It means we now have a boundary between what counts as signal and what counts as variance. Without that boundary, listing numbers every month invites strategy changes driven by noise.
What noise did not explain — which documents get cited
While individual scores stayed flat, one axis moved clearly: how often a brand's own domain is cited as the basis for an answer.
| Platform | Brand mention rate (Jul → Aug) | Own-domain citation rate (Jul → Aug) |
|---|---|---|
| ChatGPT | 33% → 33% | 8% → 25% |
| Perplexity | 83% → 92% | 42% → 75% |
| Gemini | 83% → 83% | 83% → 67% |
Mention rates held steady on two of the three platforms. Own-domain citation rates, by contrast, rose from 8% to 25% on ChatGPT and from 42% to 75% on Perplexity, and fell from 83% to 67% on Gemini. Whether a brand appears in an answer and which documents that answer rests on moved independently of each other.
The top cited sources were reshuffled as well. The video platform that ranked first in the July issue dropped out of the August top 10, and blog platforms such as Naver Blog, Tistory, and Brunch, along with startup and IT trade media, took the leading positions. The full list is in the August issue report.
The same restraint applies here. With 12 domains in the sample, one domain equals 8.3 percentage points. ChatGPT's 8% to 25% means one domain to three. Read the direction, not the magnitude.
What to watch in practice
Below is the review order derived from this data. It is not a prescription that guarantees causation, but a sequence that reduces misreading.
- Establish your own variance baseline first. Measuring twice on the same day gives you the movement band, and every later increase or decrease can be judged against it. Without a baseline, a first measurement tells you nothing about what counts as signal.
- Read grades and distribution rather than individual scores. Scores move easily near grade boundaries, while grade bands and position within the distribution are less sensitive.
- Track changes in the list of cited documents. The axis noise did not explain in this issue was the composition of cited sources. Which document types get adopted as evidence in your category connects to execution far more directly than a single-digit score move. Why surfaces outside your own site matter is covered in our source gap analysis.
- Do not change strategy on a single result. Reacting to one run without repeated measurement leads to mistaking variance for performance, or overreacting to a decline that sits inside the normal range.
To find where you stand, AI visibility diagnostics such as RanketAI can run a brand visibility measurement under the same conditions and show where your grade (A-D) sits in the category distribution.
Frequently asked questions
Our score dropped. Did we do something wrong?
First check whether the drop sits inside your own variance range. In this sample, same-day re-measurement alone moved individual scores by a median of 6 points and up to 17. If the drop is inside that range, re-measure before hunting for a cause. If it exceeds the range, separate two questions: did the mention disappear, or did a document that used to be cited drop out?
How many measurements does it take to trust the number?
There is no single correct count. The minimum condition is clear, though: you need your own variance band, established by repeated measurement at the same point in time. The St. Gallen researchers likewise warned against interpreting a single run and recommended designing the measurement itself. Related statistical criteria are summarized in our AI visibility measurement statistics analysis.
What if budget and time are limited?
If you cannot increase the number of runs, lower the unit of interpretation instead. Read grades rather than scores, monthly rather than weekly, and position in the distribution rather than individual rank. Moving to units that are less sensitive to variance reduces misreading on its own.
If the change is noise, why publish a monthly report at all?
Individual scores get buried in noise, but aggregate metrics and citation composition behave differently. In this issue, mention rates stayed flat while own-domain citation rates and the leading cited sources were clearly reshuffled. The value of monthly tracking is not "what is our score this month" but "how is the mix of documents adopted as evidence changing."
Is it not enough for a brand to be mentioned in the answer?
Mentions and citations are separate events. In this measurement, mention rates on two platforms held steady while own-domain citation rates moved sharply. If you are mentioned but the evidence comes from competitors or third-party documents, control over how your brand is described sits outside your hands. How long an answer holds up across follow-up questions is covered in our analysis of AI recommendation persistence.
Full numbers, sample criteria, and limitations are published in the August 2026 benchmark report, and the findings from the previous issue are summarized in our July 2026 issue analysis.
Execution Summary
| Item | Practical guideline |
|---|---|
| Core topic | Monthly Change Was Smaller Than Same-Day Measurement Noise (August 2026 Benchmark) |
| Best fit | Prioritize for geo workflows |
| Primary action | Standardize an input contract (objective, audience, sources, output format) |
| Risk check | Validate unsupported claims, policy violations, and format compliance |
| Next step | Store failures as reusable patterns to reduce repeat issues |
Frequently Asked Questions
What problem does "Monthly Change Was Smaller Than Same-Day…" address, and why does it matter right now?▾
Start with an input contract that requires objective, audience, source material, and output format for every request.
What level of expertise is needed to implement AI visibility effectively?▾
Teams with repetitive workflows and high quality variance, such as geo, usually see faster gains.
How does AI visibility differ from conventional geo approaches?▾
Before rewriting prompts again, verify that context layering and post-generation validation loops are actually enforced.
Data Basis
- Own research measurement: the same 12 Korean B2B SaaS domains used in the July 2026 issue were re-measured on 2026-08-12 with identical inputs and identical measurement conditions. Two rounds were run on the same day, and only the latest valid run per domain entered the distribution (the same rule as the July issue).
- Noise baseline: the per-domain score movement between the two same-day rounds was computed first, so that issue-to-issue changes could be classified as inside or outside that band. Results are fully anonymized aggregates; individual brands and domains are not disclosed.
- Limitations: with 12 domains, one domain equals 8.3 percentage points, and the gap between the two issues is 20 days rather than a full month. No statistical significance testing was performed.
Key Claims and Sources
This section maps key claims to their supporting sources one by one for fast verification. Review each claim together with its original reference link below.
Claim:Across two same-day rounds, individual scores moved by up to 17 points (median 6) and 2 of 12 domains changed grade
Source:August 2026 benchmark data.json (limitations)Claim:Month-over-month individual score change had a median of 5.5 points, inside the same-day re-measurement band
Source:August 2026 benchmark data.json (limitations)Claim:Own-domain citation rates moved from 8% to 25% on ChatGPT and 42% to 75% on Perplexity, while Gemini fell from 83% to 67%
Source:July and August 2026 benchmark data.jsonClaim:Interpreting a single AI visibility run without repeated measurement risks mistaking statistical noise for signal
Source:Univ. of St. Gallen (2026), arXiv:2604.07585
External References
The links below are original sources directly used for the claims and numbers in this post. Checking source context reduces interpretation gaps and speeds up re-validation.
- Korean SaaS AI Visibility Benchmark, August 2026 (full report)
- August 2026 benchmark data file (data.json)
- Korean SaaS AI Visibility Benchmark, July 2026 (previous issue)
- Schulte, Bleeker & Kaufmann (Univ. of St. Gallen, 2026): Don't Measure Once — Measuring Visibility in AI Search (GEO), arXiv:2604.07585
Is your site visible in AI search?
See for free how ChatGPT, Perplexity, and Gemini describe your brand.
Start Free Diagnosis →Related Posts
These related posts are selected to help validate the same decision criteria in different contexts. Read them in order below to broaden comparison perspectives.
Korean B2B SaaS AI Visibility Benchmark: ChatGPT Mentions 33% vs 83% Elsewhere
We measured 12 Korean B2B SaaS domains in real ChatGPT, Gemini, and Perplexity answers for the July 2026 benchmark: a median visibility score of 51/100, a 2.5x platform gap in brand mention rates (33% vs 83%), and citations dominated by third-party sources.
Why Brand-Name Prompts Can't Measure AI Visibility — The "AI Knows Us" Illusion
With web search on, AI mentions a brand almost every time its name is in the prompt — near-zero signal as a KPI. Baselines by query type, the three real failure conditions, and query-design criteria for choosing an AI visibility tool.
Perplexity Never Skips the Web — What the Answer Stream Reveals About Citations (2026)
A stream-level observation of Perplexity: all seven test queries triggered a live web search, yet only 5 to 9 of the 10 to 15 retrieved sources earned a citation. Local cited place entities, how-to cited video, comparison cited the vendor's own page.
89% of AI Search Demand Has No Clear Brand Owner — What 1,094 Categories Reveal (2026)
Kevin Indig tracked 1,094 US categories in ChatGPT for six months: only 15.2% had a clear brand owner and 89.3% of estimated demand sat in unowned categories, yet ownership held at a 90.4% monthly rate once claimed. How to claim a category before it locks.
Can AI Explain What Your Company Does? The 6-Axis Entity Footprint Audit (2026)
Ask an AI to describe your company and the gaps in your evidence surface at once. Across 231,347 LLM responses, entities with all three signals — authority, third-party validation, community discussion — averaged 7.8x more mentions. Score yours on six axes.