Home / Methodology

Disclosed

How we measure AI visibility.

The complete procedure, including the rules for borderline cases and the limits of the method. We disclose it because a measurement nobody can verify is not a measurement.

meit is the GEO and SEO agency for ambitious companies in Zurich. We measure before we make an offer.

Prompt set

Where the questions come from.

A prompt set is only as good as its origin. Invented questions produce invented results.

Source 1

Real customer enquiries

Wording taken from the client’s enquiries, pitch conversations and tenders. These questions are demonstrably real, because they were actually asked.

Source 2

Search terms with purchase intent

Taken from classic search and converted into natural language. They connect the measurement with what people actually search for in Google and Bing.

Source 3

The uncomfortable questions

Trustworthiness checks, direct provider comparisons, shortlist building. These are exactly the questions that produce the most revealing — and for the client most uncomfortable — answers.

Scope depends on the mandate: 40 to 60 prompts in stage 01, 80 to 120 in stage 02, 150 to 250 multilingual in stage 03. The set is not changed between measurements.

Systems

Four answer systems, two search engines.

Every system sources its content differently. Measuring one and inferring the others would not be serious work.

AI answer systems

  • ChatGPTThe most used assistant, historically connected through the Bing index.
  • Google AI OverviewThe widest reach, because it appears inside normal search results. Live in Switzerland since 25 March 2025, in four languages.
  • PerplexityThe highest citation density, its own crawler, the strongest overlap with Google’s top 10.
  • ClaudeUsed above average in B2B settings, with its own retrieval behaviour.

Measurement conditions

  • No logged-in historyEarlier conversations and saved preferences distort the answer.
  • From a Swiss IPLocation influences provider searches considerably. For German and Austrian mandates, from the country in question.
  • Comparable weekdaysTo keep weekend and update effects constant.
  • Classic search in parallelGoogle and Bing, desktop and mobile separately, 30 to 300 core terms.

Sources: Google, announcement of the AI Overview rollout in Switzerland, 25 March 2025 · StatCounter Switzerland, Bing desktop share, June 2026

How a measurement is produced
Every month, same set
What gets asked40–60real buyer prompts

  • Taken from real customer enquiries
  • Search terms with buying intent
  • The question of credibility

Unchanged for the whole engagement.

Where we measure

  • ChatGPT
  • Claude
  • Perplexity
  • Google AI Overview
  • Google
  • Bing

30 core terms with buying intent.

What you receive

  • The wording of every answer
  • Your sources, counted
  • Positions in Google and Bing
  • All of it as a table with timestamps

The raw data belongs to the client.

Always under the same conditions
  • From a Swiss IP
  • No logged-in history
  • Every answer checked by hand
  • The same procedure each month

An unchanging procedure is what makes change measurable in the first place.

The measurement setup. 40 to 60 real buyer prompts, unchanged for the whole engagement, measured monthly across four answer systems plus Google and Bing — from a Swiss IP, with no logged-in history, every answer checked by hand. You receive the wording, the counted source base and all raw data with timestamps.

Rules

What counts and what does not.

Six rules fixed before the first measurement. Anyone who sets them afterwards sets them in their own favour.

What counts as a mention

The company name appears in the answer, spelled correctly or unambiguously attributable. A mention inside a list counts the same as an explicit recommendation. Mentions in paid elements do not count.

What counts as a citation

The company is linked as a source or explicitly given as evidence. A bare mention with no source reference does not count as a citation.

What counts as an independent source

Sources the company does not control itself. The company’s own website, its own social media profiles, paid directory listings and press releases on its own channels are not counted.

How competitors are selected

Confirmed by the client, not guessed by us. We propose a list, the client removes and adds. Only confirmed competitors enter the benchmark, so the comparison does not flatter our result.

How prompts are created

From three sources: the client’s actual customer enquiries, search terms with purchase intent from classic search, and the uncomfortable questions people really ask — about trustworthiness, say, or in a direct provider comparison. We do not phrase questions that favour a result.

How often we measure

Monthly, on comparable weekdays, with an unchanged set. Changes to the set are documented and the affected values are reported separately from the month of the change onwards.

Metrics

Six values, reported separately.

We deliberately do not combine them into an index. A visibility score hides which part actually moved.

Metric What it measures How it is collected
Mention rate Share of answers in which the company name appears Per prompt, per system, monthly
Citation rate Share of answers in which the company is given as a source Per prompt, per system, monthly
Position in the answer Where in the list the company is named Only for lists, otherwise not reported
Source density Number of independent sources about the company Quarterly, checked manually
Classic position Ranking in Google and Bing per core term Monthly, desktop and mobile separately
Local visibility Presence in the map pack and the state of the Google Business Profile Monthly, per location

Limits

What this method cannot do.

The section most providers leave out. Without it, a methodology is an advertising claim.

  • We measure outputs, not causes. Why a system picks a source cannot be observed from outside. Any statement about causes is a reasoned assumption and is marked as one.
  • Individual values fluctuate. There is a random component in answer generation. We assess quarters, not weeks, and never report monthly values in isolation.
  • Models change without notice. An update can shift values overnight without anything changing at the company. We flag such breaks in the report.
  • Our sample is your market, not the market. The figures apply to your prompt set and your competitors. They do not transfer to other industries.
  • No measurement is free of judgement. Whether an ambiguous mention counts is decided by a person, following the rules above. We document borderline cases.

Verifiability

Everything you need to check it yourself.

Every client receives these five components. They belong to them, including after the contract ends.

  • The complete prompt setAs a spreadsheet, with the origin noted for each question.
  • The confirmed competitor listWith the date the client confirmed it.
  • All raw data per measurementAnswer text, system, date, assessment, borderline cases flagged.
  • The counting rules in the version that appliesChanges documented with dates.
  • The analysis logicHow the six metrics are derived from the raw data, as a formula you can follow.

Read on

The definition, with the distinctions and the three levers.

Where the two overlap, evidenced with figures.

The same methodology, applied once and without obligation.

Frequently asked

Common questions about this.

With a fixed prompt set: an unchanged list of real customer questions, put to the same systems in the same wording at the same intervals. We record mention rate, citation rate, position within the answer and the sources the answer was built from. Without a fixed set, no statement about change is possible, because even small differences in wording shift the result.

A logged-in account influences the answer through earlier conversations and saved preferences, so the result would not be reproducible. Location influences the answer as well, particularly for provider searches with a local element. Both have to be held constant, otherwise you are measuring your own session rather than the market.

The mention rate measures in how many answers the company name appears at all. The citation rate measures in how many answers the company is linked as a source or explicitly given as evidence. The second number is considerably harder to achieve and commercially more valuable, because it produces a click and an attribution.

In our experience at least 40; 80 to 120 is sensible. Below 40, the variance of individual answers comes through so strongly that changes cannot be told apart from chance. Above 120, the gain in insight slows down, except across several markets or languages.

Yes, and that is the most important caveat in any measurement of this kind. The same question can be answered differently on two days, because models get updated and because there is a random component in answer generation. That is why we assess quarter on quarter rather than week on week, and always report a monthly value together with the month before.

Yes. You receive the complete prompt set, the list of competitors checked, the counting rules and all raw data as a spreadsheet. If you rebuild the procedure, pay attention to identical wording, the same location and a logged-out session. Deviations of a few percentage points are normal.

Because almost one search in six on Swiss desktops is a Bing search, and Bing’s index feeds Microsoft Copilot. One study found that 87 per cent of ChatGPT citations matched Bing’s top results, against 56 per cent for Google. The sample was small, but the finding is relevant enough for the Swiss market not to leave Bing out.

Merag Shahzad

Merag Shahzad

Founder of meit, in digital marketing since 2006. Forbes Agency Council, LinkedIn Top Voice. He carries out every first analysis personally. Reachable directly at [email protected] – no account team in between, Monday to Friday, 09:00 to 18:00.

Next step

Test the method on your own case.

The free first analysis applies exactly this procedure to your company, once. You receive the video analysis, all raw data and the revenue model calculation, whether or not we end up working together.