Benchmark protocol, not results

The exact 100 questions are frozen, but 900 response captures and independent source coding are still required. This page remains noindex until the evidence and completion gates pass.

We Asked ChatGPT, Gemini and Google the Same 100 Questions — Here’s How Their Sources Differ

Status: Benchmark protocol and 100-question registry complete; platform responses required. No comparative result has been measured.

Updated: 12 August 2026 Author: VITON13 Research editorial desk Category: AI Search / Source Benchmark Expected reading time after results: 18–23 minutes

Direct answer

VITON13 does not yet have evidence showing how the sources surfaced by ChatGPT Search, Gemini Apps and Google AI Mode differ. The exact 100-question set and coding rules are frozen, but the required 900 consumer-product responses and their link archive have not been collected. Publishing percentages now would be fabrication.

The benchmark asks the same questions across seven intents: informational, commercial, comparison, technical, local, product and research. Each question runs three times on each product in a fresh conversation. That produces 100 × 3 products × 3 replicates = 900 response records. The study measures visible link patterns, not hidden retrieval systems and not which model is universally “best.”

The frozen 100-question registry is public before collection begins. Questions will not be rewritten after the team sees which domains appear.

What the products officially expose

OpenAI’s ChatGPT Search documentation says Search can activate automatically or be selected by the user, and that responses can include inline citations. Its Sources view can contain cited sources and other relevant links. The study therefore records inline citations and Sources-panel links as different surfaces.

Gemini Apps documentation says a response may include sources or related content, inline or in a Sources panel, and that some responses contain no links. A Gemini related link will not be relabelled as proof that the model used that page for a claim.

Google’s AI Mode documentation describes supporting web links and a query fan-out process that divides a question into subtopics. Google also says AI Mode can return only a set of web links when it lacks confidence in an AI response. Supporting links and ordinary result blocks will be preserved separately.

These descriptions justify comparing visible link surfaces. They do not reveal the platforms’ private ranking, retrieval or generation systems. The report will use “first visible link position,” not “internal ranking position.”

The 100-question benchmark

The registry contains complete, natural-language questions rather than keyword fragments. It avoids medical diagnosis, legal advice, personal finance and breaking emergencies. Time-sensitive questions are labelled so source age can be interpreted against intent.

IntentQuestionsWhat it probes
informational15explanations and reference material
commercial15evaluation before buying a service
comparison15trade-offs between approaches or technologies
technical15implementation and troubleshooting evidence
local14geographically specific, changing information
product13product-selection criteria without a sponsored shortlist
research13studies, datasets and evidence synthesis
Total100seven pre-registered strata

Category results will be reported separately and as an equal-weight macro average. The uneven number of questions cannot make a 15-question category silently dominate the overall profile.

Products, setting and collection window

The named objects of study are the English desktop consumer interfaces for:

  1. ChatGPT with Search visibly selected;
  2. Gemini Apps with web access available and no connected Workspace sources;
  3. Google Search in AI Mode.

Before collection, a public run manifest will freeze the visible product label, plan, model label when shown, browser version, country, interface language and personalization controls. All 900 runs occur inside one 48-hour window. If a product is not officially available in the selected country, the test pauses; regional controls will not be bypassed.

Every question is run in three independent rounds. Product order rotates inside each question block using a published random seed. Each run uses a new conversation or new AI Mode session. Memory, conversation history, connected apps and personal-results settings are disabled where the product exposes such controls. Any control that cannot be made equivalent is a named limitation.

The prompt is submitted exactly as registered. No follow-up asks the product to add citations, and no regeneration replaces an inconvenient first answer. Refusals, link-only responses, errors and zero-link answers remain in the denominator when a complete evidence capture exists.

Evidence preserved for every response

Each record must contain:

  • query ID, replicate, platform and UTC timestamp;
  • exact visible product, plan and model label when available;
  • country, language, sign-in state and relevant personalization controls;
  • search-active state and the unchanged prompt;
  • full response text plus a full-page screenshot or permitted export;
  • every visible outbound URL in display order;
  • the UI surface: inline citation, Sources panel, related link, supporting link, conventional result block or direct link in answer text;
  • visible link title and destination before and after redirect resolution;
  • errors, warnings, rate limits and deviations.

The archive stores raw evidence separately from the normalized analysis. A row without a response capture cannot be converted into a zero-link answer.

URL and domain normalization

Raw hrefs are retained. For analysis, the pipeline removes fragments and known analytics parameters, resolves safe public redirects, lowercases the hostname and records the canonical destination when independently verifiable. Query parameters that change page identity are preserved.

The analysis reports both hostname and registrable domain using a frozen Public Suffix List version. Subdomains remain available in the public data. Multiple links to one URL count as multiple visible link occurrences but only one unique URL in the diversity measure.

Links to uploaded files, private connected documents, account pages and unsafe destinations are excluded from public release and coded as non-public. The original capture is retained under the project’s access policy.

Source type rubric

Two reviewers use one hierarchy:

Source typeDefinition
government / intergovernmentalan official public authority or treaty body
academic / scholarlyjournal, university repository or identifiable scholarly publisher
standards / documentationstandards body or primary technical documentation
nonprofit / instituteregistered nonprofit, research institute or professional body
newsroomeditorial publication with named publishing responsibility
brand-owneda company explaining its own product, service or position
retailer / marketplacea seller, listing platform or booking marketplace
community / user-generatedforum, social platform, community answer or user review
reference / databaseencyclopedia, structured reference or public database
aggregator / directoryprimarily collects or redirects to third-party material
other / uncleardoes not meet one definition or evidence is insufficient

Ownership, not page tone, decides the type. A brand blog is brand-owned, even when it resembles a magazine. A university press journal is scholarly; a university marketing page is not automatically research.

A stratified 20% sample of link occurrences is coded independently by two reviewers before reconciliation. Pre-reconciliation agreement and Cohen’s kappa are reported for source type and link-surface coding. All disagreements are resolved against the archive, with the original labels preserved.

How source age is measured

Source age is calculated from the response timestamp to the earliest credible publication date visible in the page, publisher metadata or stable scholarly record. dateModified, an HTTP Last-Modified header and a recently changed template do not replace the original publication date.

The dataset keeps firstPublished, lastUpdated, observedAt, the date source and a confidence label. If no defensible publication date is available, age is “unknown” rather than zero. Planned buckets are:

  • 0–30 days;
  • 31–90 days;
  • 91–365 days;
  • more than one and up to three years;
  • more than three years;
  • unknown.

Source age will be interpreted by query freshness. An older standards document can be excellent evidence for an evergreen technical question; a stale transit schedule can be poor evidence for a local query.

Metrics

For each product and category: the share of valid responses with at least one public link, with a 95% interval and the number of valid responses. A Sources button is not assumed when the capture does not show one.

Report total visible link occurrences, unique URLs, unique registrable domains, median links per response and repeat rate. Surfaces are never pooled without a breakdown. The wording “citation rate” is reserved for links presented by the interface as citations.

3. First visible position

Within each surface, record the displayed order beginning at one. Report which source types and domains appear first most often. This is a UI observation, not the platform’s internal ranking.

4. Source composition and age

Report the source-type and age-bucket distribution by product and intent. Show unknown dates explicitly. Category macro averages accompany response-weighted totals.

5. Domain concentration and overlap

For each product, publish the top-domain share and Herfindahl–Hirschman index of link occurrences. For every product pair, publish Jaccard overlap across unique domains and the shared-domain list. These are descriptive concentration measures, not quality scores.

6. Brand mentions

Reviewers record brands named in response prose using a frozen alias dictionary and context check. The benchmark distinguishes an unlinked mention, a linked brand-owned domain and a recommendation. Generic words that happen to match a brand name are rejected. No sentiment is inferred automatically.

7. Claim support audit

For a stratified sample, reviewers inspect whether the linked page supports the specific nearby claim: supports, partially supports, does not support or cannot be determined. A relevant domain is not automatically evidence for the exact sentence beside it.

Planned result tables

ProductValid / 300Responses with linksMedian linksUnique domainsTop-domain shareUnknown source age
ChatGPT SearchDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIRED
Gemini AppsDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIRED
Google AI ModeDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIREDDATA REQUIRED
PairShared domainsUnionJaccard overlap
ChatGPT Search × Gemini AppsDATA REQUIREDDATA REQUIREDDATA REQUIRED
ChatGPT Search × Google AI ModeDATA REQUIREDDATA REQUIREDDATA REQUIRED
Gemini Apps × Google AI ModeDATA REQUIREDDATA REQUIREDDATA REQUIRED

No empty cell will be filled with an estimate.

Publication and comparison gates

The result remains noindex and cannot use the headline’s “here’s how” language as a completed claim until:

  • the registry contains exactly 100 unique frozen questions in all seven strata;
  • every product has at least 95% valid archived responses out of 300;
  • valid completion rates differ by no more than three percentage points;
  • every counted link can be traced to a response capture;
  • the double-coded sample covers every product and category;
  • source-type and surface agreement reach at least 90% before reconciliation;
  • URL normalization, exclusions and deviations pass independent review;
  • tables are generated from the locked analyzer, not copied by hand.

The study does not crown an overall winner. More citations can mean better traceability, repeated links, an interface convention or all three. Newer sources are not automatically better. The publishable conclusion describes source profiles, category-specific differences and uncertainty.

Data required from VITON13

DATA REQUIRED

Run environment
- visible consumer-product, plan and model labels
- country, language, browser, sign-in and personalization settings
- 48-hour collection window and public randomization seed
- three fresh-session replicates for every query and product

Response archive
- 900 complete response records or explicit archived deviations
- full text, screenshot/export and every raw outbound href
- link surface and visible order
- safe redirect and canonical resolution record

Source review
- registrable domain using one frozen Public Suffix List version
- source type, publication date evidence, date confidence and age bucket
- brand mention alias review
- independent labels for the stratified 20% sample
- claim-support review sample and reconciliation log

Analysis and governance
- locked normalized dataset and analyzer output
- missingness, exclusions, incidents and product-change log
- privacy, legal, methods and editorial sign-off

Limitations

This is a dated snapshot of three consumer interfaces in one country and language. Product labels, search partners, ranking systems, interfaces and available settings can change. Repeating an answer can produce different links. The UI does not expose all sources used internally, and a related link may not be an evidentiary citation.

The question set covers seven useful intents, not every language, country, profession or sensitive domain. Three replicates estimate ordinary variability but cannot reveal a platform’s complete distribution. Source type, claim support and brand context require judgement even with a rubric. Publication dates are missing or ambiguous on many pages.

The benchmark compares visible source behavior. It does not independently score the completeness, truth or usefulness of every answer, and it cannot prove why a platform selected a domain.

Quality score before data

CriterionCurrent scoreReason
originality9/10public 100-question registry and response-level coding design
methodology9/10frozen prompts, replicated runs, explicit surfaces and gates
source transparency9/10operator documentation and source register are public
data completeness0/10no platform responses have been captured
reproducibility9/10registry, schema and analyzer are versioned
publication readiness3/10protocol is reviewable; comparative claims are not

Editorial conclusion

The responsible answer is “not measured yet.” The protocol now makes later results auditable: anyone can inspect the same 100 questions, see which links counted, distinguish citations from related links and reproduce the tables. The page will not become indexable until the 900-response archive and independent review are complete.