ChatGPT vs Gemini vs Google AI Mode: Which One Finds New Websites Faster?
Status: Experiment pre-registration complete; controlled sites and captured responses required. No winner has been measured.
Updated: 12 August 2026 Author: VITON13 Research editorial desk Category: AI Search / Controlled Experiment Expected reading time after results: 15–19 minutes
Direct answer
VITON13 does not yet know whether ChatGPT Search, Gemini Apps or Google AI Mode finds new websites faster. No controlled test corpus, synchronized publication record or complete response archive exists in the research package. Naming a winner now would be fabrication.
This page pre-registers the comparison before data collection. The primary outcome is not whether a system says a site name. It is the first scheduled observation in which the system places a direct, working link to a canonical test page in response to a blind topical query that contains neither the site name, domain, URL nor a unique phrase copied from the page.
The design uses matched new sites, a frozen prompt set, fresh conversations, documented account and location settings, archived responses and a public coding rule. If the platforms do not produce enough valid observations, the result will be “no clear winner.”
What the products officially say
OpenAI documents ChatGPT Search as a web-search experience that may search automatically or when the user selects Search; responses can show inline citations and a Sources panel. OpenAI’s publisher guidance says public sites can appear and that allowing OAI-SearchBot helps content become discoverable, but does not guarantee placement.
Gemini Apps documentation says responses sometimes include sources or related public links and that not every response has a Sources button. Google also warns that Gemini can be inaccurate about how it works, so a model’s self-description will not be used as methodology evidence.
Google AI Mode can divide a question into subtopics, search across multiple data sources and return supporting web links. For website owners, Google says AI Mode and AI Overviews use ordinary Search eligibility: a supporting page must be indexed and eligible for a snippet, there is no special schema or AI file requirement, and inclusion is not guaranteed.
These official descriptions establish that all three interfaces can expose web links. They do not establish which one discovers an unseen site first.
What “finds a new website” means
The experiment separates four outcomes that are often collapsed into one:
| Outcome | Operational definition | Included in primary result? |
|---|---|---|
| direct page discovery | response links the canonical test page under a blind topical prompt | yes |
| site-level discovery | response links another canonical page on the same test domain | secondary |
| unlinked recognition | response names the test site without a working citation | secondary |
| prompted retrieval | response finds the site only after its name, domain, URL or unique token is supplied | separate diagnostic |
A link appearing in a conventional result block outside the generative answer will be recorded separately. A citation to a third-party page that mentions the test site is indirect discovery and does not count as a direct hit.
“Faster” means earlier observed direct citation under the scheduled checks. It does not reveal when a platform first crawled, indexed or stored the URL internally. The true event can occur between two checks, so time is reported as an interval.
Experimental corpus
Sites and pages
The study will use six newly registered, independently hosted domains after checking public registration and web-archive history. Each domain receives four English HTML pages, producing 24 test pages across four neutral, non-sensitive topics:
- one factual explainer;
- one structured comparison;
- one small original dataset with methodology;
- one practical reference page.
The page templates, word-count bands, heading depth, structured data, image count, internal-link depth and performance budgets are matched. Site names and topics are randomly assigned before publication. No site uses VITON13 branding, an existing audience, paid promotion or backlinks acquired by outreach during the 30-day window.
Using several domains reduces the risk that one domain’s topic or technical failure decides the comparison. The study still remains a bounded experiment, not a universal map of every AI search system.
Day 0 eligibility gate
A test page enters the cohort only when it:
- returns HTTP 200 on the canonical HTTPS URL;
- is accessible without login or cookie consent blocking the content;
- permits relevant crawlers in robots and infrastructure controls;
- has no
noindexand supplies a self-referencing canonical; - is linked from its site homepage and XML sitemap;
- renders the same substantive text to an ordinary browser and crawler;
- passes a stored content hash and screenshot check;
- contains no prompt injection, hidden text or platform name.
All 24 pages publish within a pre-registered 15-minute window. If a page fails the gate, the launch pauses rather than silently changing the denominator.
Prompt design
Each page receives three prompts written before publication:
- Blind exact-intent prompt: asks for the page’s specific answer without a copied phrase.
- Blind comparison prompt: asks for multiple relevant sources or options.
- Blind recency prompt: asks for recently published evidence on the topic.
Prompts cannot contain the site name, domain, URL, author name, title, test token or a sequence of eight words from the page. A script will reject prompt rows that leak those fields.
After the primary 30-day test, a separate diagnostic asks first for the site by name and then by exact URL. Those trials measure user-directed retrieval, not organic discovery, and cannot make a platform the winner.
Run schedule and environment
Blind queries run at 12, 24, 48 and 72 hours, then on Day 5, 7, 10, 14, 21 and
- Each platform receives the same prompt set within a randomized block. The order rotates so one platform is not always tested first.
Every observation records:
- UTC timestamp and scheduled day;
- platform and visible product/surface label;
- subscription tier and model label, if shown;
- country, interface language and device class;
- signed-in state and personalization/history settings;
- exact prompt ID and fresh-conversation ID;
- whether web search was visibly active;
- full response text, screenshot and exported/share record where permitted;
- inline citations, Sources-panel links and conventional result links;
- errors, refusals, rate limits and unavailable features.
The experiment begins only where all three products are officially available in the selected country and language. It will not bypass regional restrictions. Fresh conversations are used for every observation. Memory, personal search history and connected applications are disabled where controls exist. The published limitations will state where the interfaces do not offer equivalent controls.
Response preservation and coding
Two reviewers independently code normalized destination URLs without seeing the platform summary. Disagreement is resolved against the archived response and logged. Redirects are followed in a safe, non-authenticated environment; the final canonical destination determines direct versus indirect discovery.
The coding hierarchy is:
- direct canonical page citation;
- direct same-domain citation;
- indirect third-party citation;
- unlinked site mention;
- no discovery signal;
- invalid observation because the product, search mode or archive failed.
An answer that merely repeats a false URL or hallucinates a site name does not count. A link must resolve to the controlled domain and must have been visibly present in the preserved response.
Primary endpoint and winner rule
Primary endpoint
For each platform, calculate the share of 24 test pages directly cited at least once by Day 30 under blind prompts.
Speed endpoint
For each platform and page, record the first scheduled day with a valid direct citation. Pages not detected by Day 30 remain censored. Report milestone detection curves and a 30-day restricted mean observation time; do not replace undetected pages with a fictional discovery date.
Pre-registered winner rule
A platform is labelled the descriptive winner only if all conditions hold:
- at least 90% of scheduled observations are valid for every platform;
- completion rates differ by no more than five percentage points;
- it has the highest Day-30 direct-page detection coverage;
- it has the lowest 30-day restricted mean observation time;
- both independent reviewers agree on at least 95% of primary classifications before reconciliation.
If one platform has broader coverage but another has earlier detections, if the sample is too incomplete, or if the metrics tie, the conclusion is no clear winner. With 24 pages, the report will emphasize effect sizes and uncertainty, not manufacture statistical certainty.
Planned result table
| Platform | Valid observations | Pages directly cited by Day 30 | Day-30 coverage | Restricted mean observed days | First observed hit | Result |
|---|---|---|---|---|---|---|
| ChatGPT Search | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | not measured |
| Gemini Apps | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | not measured |
| Google AI Mode | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED | not measured |
Planned milestone chart
| Scheduled check | ChatGPT Search | Gemini Apps | Google AI Mode |
|---|---|---|---|
| 12 hours | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 1 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 2 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 3 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 5 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 7 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 10 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 14 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 21 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
| Day 30 | DATA REQUIRED | DATA REQUIRED | DATA REQUIRED |
Each value will be the cumulative number of cohort pages directly cited, not the number of citations or mentions. Repeated links to one popular page cannot inflate coverage.
Contamination controls
- Test prompts and domains remain private until the primary window closes.
- Reviewers do not share or publicly link test pages during collection.
- Exact page URLs are never pasted into the blind trial interfaces.
- Referral logs are inspected for unexpected public traffic and backlinks.
- The study archives sitemap, robots, DNS, deployment and content changes.
- Any crawler blocks, outages or platform incidents are timestamped.
- Pages are not rewritten to chase early results.
- Prompt order is randomized and the seed is published.
- A conventional Google result check is stored separately and never substituted for AI Mode evidence.
The research team cannot prevent a platform update, third-party indexing feed or unsolicited link. Those are recorded as real-world events rather than hidden.
Data required from VITON13
DATA REQUIRED
Experimental assets
- six eligible new domains with archived history checks
- 24 matched public pages and their content hashes
- randomized site/topic assignment and random seed
- frozen prompt registry with three blind prompts per page
- Day-0 publication, DNS, sitemap and deployment timestamps
Platform observations
- scheduled runs at 12h, 24h, 48h, 72h, Day 5, 7, 10, 14, 21 and 30
- exact product/surface, account, tier, locale, country and settings
- complete response text, screenshots and permitted exports/share links
- normalized inline, Sources-panel and conventional-result URLs
- search-active state, errors, rate limits and missing runs
Operational evidence
- server/CDN crawler and referral aggregates
- robots.txt and sitemap snapshots
- backlink and unexpected-traffic audit
- platform incident and experiment-deviation log
Review
- two independent classifications
- reconciliation log
- privacy, legal and editorial review before publicationNo test-site credentials, private account identifiers or session tokens belong in the public dataset.
Limitations
The interfaces are not identical search engines. ChatGPT Search may use search partners and rewritten queries; Gemini Apps does not attach sources to every answer; AI Mode is part of Google Search and can personalize responses when history settings are enabled. Product models, retrieval systems and interfaces can change during the 30 days.
The sites’ topics, language, country, technical stack and lack of established links constrain generalization. Repeated scheduled queries may themselves create user-interest signals. Detection through a citation does not reveal the underlying crawler, index or retrieval path. Absence from one preserved response does not prove that a system had never encountered the site.
The experiment measures visible retrieval under a declared environment. It cannot measure every user, geography, account tier or hidden product state.
FAQ
Is ChatGPT Search the same as OAI-SearchBot?
No. OAI-SearchBot is a crawler identity described for search discoverability; ChatGPT Search is the user-facing product that can return web-linked answers. Crawler access can support eligibility but does not guarantee a citation.
Are Gemini Apps and Google AI Mode the same product?
No. They are measured as separate user-facing surfaces. Gemini Apps may attach sources or related links to some answers, while AI Mode is a Google Search experience using Search systems and query fan-out. Their visible outputs and controls must be archived separately.
Why not paste each URL and ask whether the platform can open it?
That tests prompted retrieval after the user supplies the location. It does not show that the system independently discovered the page. Exact-URL retrieval is useful, but it is a secondary diagnostic.
Does the first citation reveal the first crawl or index date?
No. It is only the first positive result at a scheduled observation. Internal discovery may have happened earlier, and the page may have entered a response through an undisclosed retrieval path.
Why use 24 pages instead of one VITON13 page?
One result could be driven by topic demand, domain history or a technical error. A matched multi-domain cohort gives a more stable descriptive comparison while remaining feasible to inspect and archive.
What happens if no platform finds most pages?
That is a valid negative result. The report will publish detection coverage, missingness and confidence limits. It will not change the success definition after seeing the data.
Sources and editorial accountability
Product-behavior claims are linked to current OpenAI and Google help or Search Central documentation. The dated source register is stored in sources.md. Every primary result will link to a privacy-reviewed observation record and coding decision. VITON13 has no affiliate relationship with the compared products in this study.
Update plan
Freeze the protocol, pages, prompts and random seed before Day 0. Append a signed run record after each scheduled block. Do not inspect aggregate platform rankings until the primary coding pass is complete. After Day 30, reconcile reviewer labels, run the locked analysis, publish deviations and then decide whether the evidence qualifies for indexing. Preserve the version and visible product labels because later reruns may use different systems.
Quality score before data
| Dimension | Score | Note |
|---|---|---|
| Originality | 15/20 | matched multi-domain and blind-prompt protocol is defined |
| Information gain | 16/20 | separates independent discovery from prompted retrieval |
| Evidence | 14/15 | primary product documentation supports capability boundaries |
| Search intent | 14/15 | directly defines “finds” and a fair comparison |
| First-hand experience | 0/10 | controlled pages and observations do not yet exist |
| Structure | 9/10 | corpus, schedule, coding and winner rule are pre-registered |
| Freshness | 5/5 | official sources checked 12 August 2026 |
| Technical SEO | 4/5 | metadata/schema prepared; results page remains noindex |
| Total | 77/100 | Not eligible for a completed experiment label or indexing |
