By Manuel Hürlimann for GaryOwl.com | Published: July 23, 2026 | Last updated: July 23, 2026
Expertise: Digital Authority Engineering | AI citation pipeline diagnostics | Source-Trust Mechanics
Time to read: ~12 minutes · ~2,800 words
Series: DAE Mechanics — DAE Glossary
ChatGPT is not a search engine. When it searches the web, the browser-visible result payload reveals several distinct sourcing pipelines — and a single field in that payload, result_source, labels which infrastructure fetched every source in the observed records. What that field proves is firm. What people conclude from it about contracts and backends is not. This article draws that line precisely.
📌 Navigate
Authority Intelligence Lab · DAE Framework · DAE Glossary
📌 Glossary: Key DAE Terms in This Article
Sourcing Pipeline · Stage-Dependent Authority · Evidence Tiers
TL;DR — Key Takeaways
ChatGPT is not a search engine. How ChatGPT fetches its sources is visible in its own authenticated web traffic. As documented in late June 2026, every retrieved source carried a label called result_source with one of four observed values: serp, labrador, bright, and oxylabs. In early July a fifth value, bing, surfaced in a second observer’s instance and was reproduced by the primary analyst — cohort-gated at the time of observation, and treated below with that caveat. The field’s existence and its values are firm: independently documented by one primary analyst and at least two independent verifiers, one of them a working browser tool that could not function if the field were not there (Mohanadasan, 2026; verifiers below) [Tier D], replicated practitioner analyses.
The pattern the labels describe is a layered sourcing system: a publisher-heavy pipeline consistent with licensed-content access that delivers substantially longer text excerpts from named news organizations, a bulk retrieval layer that delivers only short metadata snippets in the observed records, a regionally skewed third fetcher, and a conventional SERP channel. That structure has a direct consequence for anyone who wants to be cited: large publishers do not merely look more authoritative to ChatGPT; in the observed records they sit in a privileged sourcing tier whose result entries carry more of their text. Everything beyond that — who operates the pipelines commercially, what search index sits underneath them, whether any contract exists — is inference, and this article marks it as such.
What the network traffic actually shows
The finding originates in packet-level observation, not in speculation. In late June 2026, Suganthan Mohanadasan documented roughly 1,240 source records from his own logged-in ChatGPT Pro account by reading the conversation JSON in browser developer tools. Each web source the system retrieved carried a field named result_source, and across all searches observed in the June window that field took exactly four values: serp, labrador, bright, and oxylabs (Mohanadasan, 2026) [Tier D], documented method, single account.
Two independent parties looked at the same traffic and found the same thing. Synscribe read the conversation payloads through the backend API and logged the same four values, plus an orchestrator codename (“Sonic”) in adjacent debug fields (Synscribe, 2026) [Tier D]. And Resoneo built a Chrome extension that reads result_source live and aggregates the pipeline mix per conversation — a tool whose sourcing-pipeline feature depends directly on parsing that field (Resoneo, 2026) [Tier D], functional verification. One primary observer plus two verifiers who looked at the wire themselves: that is what makes the core finding firm rather than folklore. Re-sharing without looking is not verification. The previous article in this series treats that distinction at length.
One boundary matters before anything else: the field lives only in the authenticated web traffic (the browser’s conversation endpoint). OpenAI’s official Responses API exposes web-search tool activity, citations and source information, but no documented result_source sourcing label (OpenAI Responses API web-search docs) [Tier B]. Whoever wants to observe the pipelines must look at the product, not the API — the two surfaces answer different questions, which is a recurring theme in how these systems are measured.
Everything in this article is a snapshot of an undocumented, internal mechanism observed in June–July 2026. These fields carry no guarantees; they can be renamed or removed in any weekly release. The primary analyst’s own follow-up re-inspection in early July found the field alive but changed underneath — one new value (bing), one formerly empty field (supporting_websites) now populated, and two of his own June readings corrected. That is the perishability thesis confirmed in miniature: the structure held, the details moved. Treat every specific value in this article as firm for its observation window and perishable beyond it.
The sourcing labels: serp, labrador, bright, oxylabs — and a late arrival, bing
What does each label mean? The names themselves are internal and undocumented, so everything here rests on observed behavior — what kind of source arrives under which label, and in what form.
labrador delivered content from named news and reference organizations (the observed examples include Reuters, WSJ, Wikipedia and arXiv-hosted material), with substantially longer text excerpts of up to roughly 1,080 characters in the primary analysis and full paragraphs in the verifier’s schema description, plus adjacent type labels like “news” and “academia” (Mohanadasan, 2026) [Tier D]. The pattern is consistent with a licensed-content pipeline (more on that below).
bright did the bulk of the fetching in both independent samples, especially on commercial, shopping, finance and weather queries — and, critically, delivered short metadata records: URL, title, a brief snippet (around 150 characters in Synscribe’s inspection), a date (Synscribe, 2026) [Tier D], Synscribe and Mohanadasan concordant.
oxylabs skewed regional and local in the primary sample — Gulf-region press appeared under this label for an analyst located in Dubai (Mohanadasan, 2026) [Tier D], single-account observation.
serp appeared least distinctly, carrying predominantly news-type results (Mohanadasan, 2026) [Tier D].
bing is the late arrival, and its evidence status differs from the four above, so it deserves its own paragraph. In early July 2026, GEO researcher David Konitzny posted network captures from his own account showing "result_source": "bing" on ordinary review pages, a value absent from the primary June dataset. The primary analyst then aimed his own instance at the same pages and reproduced the value, while confirming it did not appear in his account’s regular traffic (Mohanadasan, Part 2, 2026) [Tier D], two independent first-hand observers, different accounts. The pattern is consistent with a cohort-gated rollout: pages served from Bing’s ecosystem for some users and not others, beginning in late June. Two caveats keep this honest. First, this article’s author has not verified bing in his own traffic; the value rests on two independent external observations, not on first-party inspection. Second, cohort-gated features can be widened or withdrawn in any release; the analyst’s own framing — file it under watch, not panic — is the right one. If it holds, the practical consequence is direct: pages missing from Bing’s index would be missing from those users’ ChatGPT, which makes Bing Webmaster Tools and a sitemap check cheap insurance.
Whether bright and oxylabs are literally the commercial scraping providers Bright Data and Oxylabs is the obvious reading of the names — and it is exactly where evidence ends and inference begins. The names are suggestive; no contract is visible in the traffic, and the section on what this does not prove treats that line explicitly.
The runner-up layer: supporting_websites
The July re-inspection also promoted a field the June analysis had dismissed as empty. supporting_websites now arrives populated: when ChatGPT cites a source for a claim, that citation carries an array of other pages that supported the same claim but did not win the visible slot: runner-up citations, each with its own result_source, sitting invisibly under the winner (Mohanadasan, Part 2, 2026) [Tier D]. The analyst calls it quietly the most useful thing in the payload, and the reason is competitive: the wire now shows, per claim, exactly which page beat yours and which pages you are beating — not at the vague level of “competitor X gets cited more,” but this sentence, this claim, this winner. For the themes of this series it carries a second meaning: the visible citation list understates how many sources the system actually weighed, and the competition for a citation happens claim by claim, not page by page.
Why some sources arrive as full text and others as snippets
The most consequential observation is not which pipeline fetches what, but how much source text is visible at this observed retrieval stage. Under labrador, result records carried substantially longer text excerpts (roughly 1,080 characters in the primary analysis; full paragraphs in the verifier’s schema). Under bright, the record carried a URL, a title, and a short snippet — little more (Mohanadasan, 2026) [Tier D], two concordant independent analyses.
That asymmetry matters for what a source can contribute. At this observed stage, a record carrying several paragraphs offers the model more of the source to work with than a record carrying one short snippet. What the observation shows is a large difference in excerpt length in the visible result records; it does not by itself prove that no further fetch or open step occurs, nor that the visible record is the totality of text the model can access. Even read conservatively, the pattern points to tiers of how much of a source appears at this stage — set by infrastructure, not by content quality.
Why the labrador pattern is consistent with a licensing tier
The labrador pattern does not float free of context. OpenAI has publicly announced content-licensing partnerships with, among others, News Corp (WSJ), Axel Springer, the Associated Press, The Guardian and the Financial Times, in official announcements by the parties themselves [Tier B]. A pipeline that delivers longer excerpts from exactly this class of publisher is consistent with what a licensed-content tier would look like in the result records — though the mapping from the label labrador to any specific licensing arrangement is inference, not something the traffic proves.
Here is the interpretation, marked as ours: this is a candidate mechanistic explanation for why large publishers dominate ChatGPT citations, one that does not require the word “authority” at all [Tier DAE]. In the framework of this series, it is stage-dependent authority in its purest form: the advantage is conferred at the sourcing stage, by infrastructure, before any ranking or relevance judgment sees the content. A licensed publisher is not necessarily judged more credible — it enters the pipeline through a wider door, with more of its text intact. If you compete with such publishers for citations, you are not losing a credibility contest; you are standing in a different queue.
What this does not prove
This is the section most re-tellings of the finding skip, and it is the reason this article exists.
It does not prove that OpenAI pays Bright Data or Oxylabs. No contract is visible in network traffic, and the primary analyst says so himself. The field values are internal strings whose names suggest those providers; “the names point to” is the strongest claim the evidence supports. Formulations like “ChatGPT’s search is powered by Bright Data” convert an inference into a fact the wire does not contain.
The name collision is real. Both Bright Data and Oxylabs sell products called “ChatGPT scrapers” — tools their customers use to scrape ChatGPT itself. The existence of those products is evidence of nothing about ChatGPT’s backend; it is a naming coincidence that has actively fed the confusion.
What sits underneath is a black box. Even granting the middleware reading, the layer below it, whether the fetchers query Google, Bing, a proprietary cache, or a blend, is invisible in the traffic. One verifier’s attempt to reproduce ChatGPT’s result sets by running the same query strings through Google and Bing did not cleanly match (Green, via Search Engine Land, 2026) [Tier D]. The honest description is: a confirmed retrieval layer with an unconfirmed interior.
The distributions are shape, not measurement. Counts like “484 bright results against 5 labrador” come from a handful of conversations on single accounts, with query mixes that reflect their authors’ work and location. They establish that the bulk pipeline dominates in those samples — not a population-level ratio. Anyone quoting these numbers as ChatGPT-wide statistics has left the evidence behind.
What it means for being cited
Once you can see how ChatGPT fetches its sources, the practical consequences follow from the mechanics, and they are narrower than the optimization industry would like.
First: for most of the web, ChatGPT sees metadata, not pages. If your key facts live deep in body text, a short-snippet record may not carry them into the visible result. This suggests a practical hypothesis: facts that can survive as a title, a heading, or a first-sentence summary may travel further through a metadata-only record than facts that require the full page [Tier DAE], hypothesis from the fetch mechanics above.
Second: presence in the places the pipelines actually fetch matters more than on-page perfection. The bulk pipeline concentrated on commercial and comparative queries, where aggregators, listicles and community threads dominate the retrieved set. This suggests a further hypothesis: a parseable fact on a source that repeatedly enters the retrieved set may carry more citation opportunity than a stronger first-party page that does not appear in that set [Tier DAE], hypothesis from the observed retrieval mix.
Third: the citation pool this feeds into is shifting under everyone’s feet. In one vendor study of 150 conversations on a single US Plus account, first-party citation rates fell from 56.8% on GPT-5.4 Thinking to 47% on GPT-5.5 Thinking, with Reddit leading the cited-domain count in that sample (Writesonic, 2026) [Tier E], vendor study, single-account scope, COI: tool provider. The model layer itself keeps moving: the GPT-5.6 family (Sol/Terra/Luna) entered general availability on July 9, 2026, with ChatGPT access varying by model, plan and product surface, with no sourcing-pipeline impact documented either way (OpenAI, 2026) [Tier B], platform release note. The sourcing pipelines are stable enough to describe; what the model does with them is not. Which is one more reason to measure citation share as an interval over repeated runs (the subject of the measurement article in this series) rather than as a fact checked once.
Honest Limitations
Everything here describes an undocumented internal mechanism, observed by a small number of independent analysts over a few weeks on logged-in consumer accounts. The field is invisible in the official API, so none of this is measurable at scale through sanctioned channels. The distributions are single-account snapshots. The pipeline-to-provider mapping is inference from field names. The interior of the retrieval layer is a black box, explicitly. And the observation window is narrowing with two untested variables, not one: interface changes in early July 2026 removed adjacent data from the web payload (fan-out queries vanished between GPT-5.3 and 5.4), and the GPT-5.6 family entered general availability on July 9, 2026 (ChatGPT access varying by plan) — whether result_source itself persists after that rollout has not been publicly verified at the time of writing. If the field disappears, the mechanics described here do not become false — they become unobservable, which is a different thing, and worth remembering as this article ages.
Frequently Asked Questions
Does ChatGPT pay Bright Data and Oxylabs?
Not provable from the traffic. The field values bright and oxylabs suggest those providers by name, but no contract, invoice or partnership is visible in network data, and the primary analyst explicitly declines that claim. Treat every confident assertion of a paid backend as having exceeded its evidence.
Is ChatGPT using Bing or Google under the hood?
Unknown — deliberately so, from the outside. The observable layer ends at the fetching middleware. A reproduction attempt running identical queries through Google and Bing did not cleanly match ChatGPT’s result sets, which argues against a thin pass-through but proves nothing about the blend underneath.
Can I see this myself?
Yes, and you should — the finding is reproducible, which is its greatest strength. In a logged-in ChatGPT session with web search active, open browser developer tools, trigger a search-invoking question, and inspect the conversation JSON for result_source. If you can read a network tab, you can verify the core claim of this article in five minutes.
Will these findings stay true?
Assume not. This is a snapshot of an internal mechanism that changes weekly; adjacent observability has already narrowed twice in 2026, and the field’s status after the July 9 GPT-5.6 rollout was still unverified at publication. The durable takeaway is not the field name — it is the structure it revealed: layered sourcing, a licensed full-text tier, and metadata-only fetching for most of the web.
Does this apply to the API or only the ChatGPT product?
Only the product traffic exposes the field. The official Responses API returns queries and citations without any sourcing label — so API-based measurement, whatever its other virtues, cannot see pipelines at all.
Evidence Tiers Used in This Article
| [Tier A] | Peer-reviewed academic research |
| [Tier A*] / [Tier B*] | Work of that tier’s methodological quality, peer review pending (star drops on acceptance/replication) |
| [Tier B] | Primary platform statement or large-scale/credible study, pending or outside peer review |
| [Tier C] | Independent meta-analysis (aggregates ≥ 10 external sources, transparent methodology, vendor affiliation disclosed) |
| [Tier D] | Reputable journalism or industry study with documented methodology, not vendor-self-published |
| [Tier E] | Vendor study (self-published, regardless of sample size or methodology quality); COI disclosed inline |
| [Tier DAE] | Framework term (synthesized from empirical sources, attributed to DAE) |
Sources & Methodology
[Tier D] Mohanadasan, S. (2026). “How ChatGPT Actually Picks Sources.” suganthan.com, June 24, 2026 (updated June 27). Primary analysis; ~1,240 source records, one logged-in Pro account, documented method, author-disclosed sampling bias. suganthan.com/blog/how-chatgpt-picks-sources (Accessed: July 14, 2026)
[Tier D] Mohanadasan, S. (2026). “ChatGPT Changed How It Picks Sources While You Were Reading My Last Post.” suganthan.com, early July 2026. Follow-up re-inspection ~10 days after the primary analysis: fifth result_source value bing (first observed by David Konitzny on a separate account, then reproduced by the author — two independent first-hand observers), supporting_websites runner-up layer now populated, two of the author’s own June readings corrected. suganthan.com/blog/how-chatgpt-picks-sources-part-2 (Accessed: July 14, 2026)
[Tier D] Synscribe (2026). “The Search Engine Powering ChatGPT Is Bright Data.” June 2026. Independent backend-API analysis; documented method; sample counts are single-snapshot. Title overstates: confirmed is a middleware retrieval layer, not “the search engine.” synscribe.com (Accessed: July 14, 2026)
[Tier D] Resoneo (2026). ChatGPT Search & Fan-outs Capture, Chrome extension V3.4. Functional verification: the tool’s sourcing-pipeline feature depends on parsing result_source live per conversation. think.resoneo.com/scrap-chatgpt-plugin (Accessed: July 14, 2026)
[Tier B] OpenAI content-licensing partnerships, official announcements (2023–2025): News Corp (openai.com/index/news-corp-and-openai…), Axel Springer (openai.com/index/axel-springer-partnership), The Associated Press (openai.com/index/openai-and-the-associated-press), The Guardian (openai.com/index/the-guardian), Financial Times (openai.com/index/content-partnership-with-financial-times). Establishes the licensed-publisher landscape; the mapping to the labrador label is inference, not established by these deals. Reuters is deliberately omitted: no OpenAI–Reuters content-licensing agreement was verifiable in this pass. (Accessed: July 14, 2026)
[Tier E] Writesonic (2026). “GPT-5.5 Cites Brand Sites 47% of the Time.” GPT-5.5 citation study, April 27, 2026: 150 conversations (50 prompts × 3 model versions), single US Plus account. First-party citation rate 56.8% (GPT-5.4 Thinking) → 47% (GPT-5.5 Thinking); Reddit led the cited-domain count. Vendor study, single-account scope, COI: published by a tool provider. writesonic.com/blog/gpt-5-5-citation-study-vs-gpt-5-4 (Accessed: July 14, 2026)
[Tier B] OpenAI (2026). GPT-5.6 (Sol/Terra/Luna) general-availability announcement and model release notes, July 9, 2026. Basis for the model-context statement; ChatGPT access varies by model, plan and product surface per the release notes. Deep link: openai.com/index/gpt-5-6 and the OpenAI model release notes. community.openai.com — GPT-5.6 announcement (Accessed: July 14, 2026)
Methodology: This article separates three layers of claim strength — wire-observed facts (multiple independent traffic analyses), officially documented context (platform announcements), and inference (marked as such in the text). Distribution figures are reported as sample shapes, never as population statistics. The snapshot caveat in the opening section governs the whole article. Access dates reflect the publish-day verification pass.
Update Log
V1.0 (July 23, 2026) — Initial publication.
[Future updates logged here.]
About the Author
Manuel Hürlimann runs GaryOwl.com and the octyl AI Authority Lab, where the DAE (Digital Authority Engineering) framework is developed and tested against live AI systems. This article belongs to the mechanics track of the DAE series: tier-disciplined analyses of how AI citation systems actually work, with the evidence line drawn where it actually sits. Disclosure: no relationship with OpenAI, Bright Data, Oxylabs, or any tool vendor cited here.