Rage Against The Machine: The AI Citation Paradox

September 22, 2026

Nicole Dobernig, Founder of Data Insight, says AI visibility depends less on a single prestigious placement than on a current, credible trail of evidence machines can find.

Rage Against The Machine: The AI Citation Paradox
Credit: Agentic Marketing News

Blocking AI training does not mean disappearing from AI search. For brands chasing citations, confusing the two can be an expensive mistake. The argument sounds logical enough: major publishers are blocking AI crawlers, therefore paying for press releases, earned media, or coverage in authoritative publications is increasingly useless for AI visibility. There is one problem: It confuses AI training with AI retrieval.

The distinction has become one of the most important, and most misunderstood, parts of AI search strategy. A publisher can prevent a company from vacuuming up its archive to train a foundation model while still allowing its journalism to appear in search, AI-generated answers, or licensed retrieval systems.

Cloudflare has now formalized this distinction in its crawler controls. Its framework separates search, training, and agent activity and specifically allows publishers to disallow AI training without necessarily sacrificing search discoverability. Cloudflare notes that companies including OpenAI separate search and training crawlers, allowing the training crawler to be blocked without automatically blocking search.

That changes the conversation around digital PR. The question is no longer simply, “Does this publisher block AI bots?” The better question is: “Which bots does it block, for what purpose, and through what other channels can its information still reach an AI system?” That is the AI citation paradox.

One publisher. Several doors.

Nicole Dobernig is the Founder of Data Insight, an AI-native marketing intelligence agency helping iGaming and other highly competitive brands compete across traditional search and AI discovery. Dobernig spent close to two decades in marketing across small and mid-sized businesses, most recently leading marketing for an international freight-shipping company. She says her team consistently competed well above its weight against multi-billion-dollar companies in organic search and a muti-category leader in AI recommendations .

She moved deeper into AI search after recognizing that discovery was fragmenting. Brands were no longer competing for ten blue links on one search engine. They were competing across search engines, answer engines, visibility systems and increasingly agentic interfaces, each assembling information differently.

That fragmentation also exposes a flaw in one of the industry's emerging assumptions: that a publisher blocking an AI crawler has somehow become invisible to AI. It hasn't necessarily. A modern publisher may have one policy for model training, another for search indexing, another for user-initiated retrieval and yet another relationship established contractually through licensing.

Cloudflare describes precisely this emerging architecture. Its controls distinguish training from search, including the ability to communicate a “no-training preference while remaining discoverable in search.” The implications for marketers are substantial.

GPTBot is not the entire machine

For marketers, "AI crawler" has become a dangerously broad category. GPTBot, for example, is associated with OpenAI's collection of web content that may be used to improve its models. But AI companies can operate separate mechanisms for search and retrieval. Blocking a training crawler therefore does not establish that the information cannot appear in an AI-generated answer.

Cloudflare's current framework explicitly recognizes this separation. It says training-only crawlers operated by companies including OpenAI can be blocked without affecting their separate search functions. And crawling is only one route into the machine.

In 2024, OpenAI and News Corp announced a multi-year global agreement giving OpenAI access to current and archived journalism from News Corp publications. The agreement includes content from The Wall Street Journal, Barron's, MarketWatch, Investor's Business Daily, The Times, The Sunday Times, The Australian, the New York Post and other properties.

Under the agreement, OpenAI has permission to display News Corp content in response to user questions and use that content to enhance its products. That is an important counterexample to the idea that a publisher's robots.txt file tells marketers everything they need to know about AI visibility. A closed crawler door does not mean a closed information pipeline.

The press release isn't dead

The same mistake appears in discussions about press releases. If major publishers restrict AI scraping, the argument goes, why pay for distribution? Because distribution networks themselves can be highly accessible sources.

PR Newswire has publicly taken the opposite position from publishers attempting to restrict AI access. In October 2025, the company said it maintains an open-access policy for legitimate AI explorers and LLMs and intentionally allows legitimate AI and search bots to crawl, scan and index its public content. That matters because the wire service is not merely a doorway to another publication.The release hosted on the wire is itself a machine-readable source.

PR Newswire now explicitly markets press releases around AEO and GEO, arguing that structured releases can be discovered and cited by systems including ChatGPT, Gemini and Perplexity. The economics therefore need to be evaluated differently. A press release should not be purchased simply because somebody promises hundreds of syndicated backlinks. That is an old SEO argument wearing new clothes.

The more interesting value is “entity distribution”. A properly structured release can establish a machine-readable relationship between a company, executive, product, category, event, geography and claim. Distribution then replicates those associations across multiple independent locations on the web. For AI visibility, that creates something more useful than another backlink. It creates corroboration.

Authority is becoming a network, not a masthead

This does not mean every expensive media placement is worthwhile. Nor does it mean every press release creates AI authority. The mistake is thinking in absolutes.

A major publication may be valuable because its journalism enters an AI ecosystem through licensing. Another may remain accessible to search crawlers but reject training. Another may sit behind technical or commercial restrictions that make retrieval difficult. A specialist publication may be far easier for answer engines to retrieve and may provide stronger topical context despite having a fraction of the traditional domain authority.

The value of coverage therefore cannot be determined from the prestige of the masthead alone. Dobernig's approach is to think in terms of ‘citation architecture’: creating enough independent, credible and retrievable evidence around an entity that an AI system has multiple reasons to understand and trust it. That can include national journalism, trade publications, wire services, industry interviews, original research, expert commentary, review platforms, company websites and other independently accessible sources.

The objective isn't to appear everywhere. It is to create a consistent entity narrative across the sources machines can actually retrieve. That also changes how marketers should evaluate PR. Traditional SEO trained an entire industry to look for the hyperlink. AI systems have another problem to solve: entity resolution. They need to determine whether a company exists, what it does, which subjects it is associated with, whether independent sources corroborate those associations and whether the information is current enough to trust.

A clear company or executive mention in a headline or opening paragraph can therefore have strategic value even when the publication does not provide the perfect followed backlink. For Dobernig, this is why PR, SEO and reputation management are beginning to collapse into the same discipline. The unit of value is increasingly not simply the link. It is the verifiable association.

Reviews: Another layer of corroboration

Reviews operate on the same principle. Dobernig treats platforms such as Trustpilot and Clutch as part of the evidence layer surrounding a company rather than as isolated reputation channels. A collection of detailed reviews can help establish what customers associate with a business: its services, category, geography, expertise and outcomes.

The language matters. "Great company" says almost nothing. A detailed review naturally describing what was purchased, the problem being solved, the market involved and the result creates considerably more contextual information. This is not an argument for manufacturing keyword-stuffed reviews. Authenticity matters. It is an argument for recognizing that customer language contributes to the broader entity footprint AI systems encounter when researching a brand.

Dobernig still treats technical SEO as the foundation beneath the entire system. A brand can accumulate press, reviews, interviews and third-party mentions while simultaneously making its own website difficult for machines to interpret. Broken canonicals, duplicate URLs, inconsistent entities, poor internal linking, contradictory schema, inaccessible content and crawler restrictions can weaken the first-party source that should provide the clearest description of the company.

"Everything else works better when the technical base is solid," Dobernig says. In the AI-search era, technical SEO is no longer only about helping Google crawl a website. It is about creating a clean, machine-readable “source of truth”.

Two discovery layers, always running

Underneath the tactics is a larger structural change. Discovery now operates across at least two overlapping layers: traditional search and AI-mediated discovery. The first ranks documents. The second can assemble an answer from documents, indexes, licensed datasets, retrieval systems, knowledge graphs and other sources before the user ever visits a website.

Brands therefore cannot optimize for one and assume the other will follow automatically. Dobernig describes the work as continuous optimization across both discovery layers. The website establishes the first-party entity. Technical SEO makes it interpretable. PR distributes the entity into independent sources. Reviews provide customer corroboration. Specialist publications establish topical relevance. Major journalism can contribute authority. Wire services distribute structured information.

AI retrieval systems then encounter those signals through a mixture of crawling, search indexes, partnerships, licensing and live retrieval. The system is considerably messier than "AI crawled my page." That is precisely the point.

Recency becomes part of authority

Authority used to accumulate slowly. AI discovery introduces another variable: freshness. A company with strong historical authority but very little recent information may be less useful for a query requiring current evidence than a company generating credible new signals across several sources.

Dobernig therefore argues for a steady publishing and PR cadence rather than occasional visibility campaigns. Not noise, evidence. New research. New interviews. New case studies. New customer experiences. New announcements. New expert commentary. New third-party corroboration.

The objective is to keep reinforcing what the entity is, what it knows and why it should remain part of the answer.

The shift goes beyond search. Dobernig calls AI an accelerant, not a replacement. “It removes the heavy R&D friction that used to slow everything down.” Research, competitive scans, content modelling and technical audits that once took weeks can now be orchestrated across multiple systems in days.

Speed does not remove judgment. It multiplies the need for it. “The CMO role has turned into something closer to an air-traffic controller,” she says. The modern marketing leader coordinates agents, datasets, specialists and platforms, setting strategy and guardrails while machines handle more of the repetitive work. The advantage is not simply producing more content. It is watching the discovery environment, spotting gaps, deploying resources and continuously reinforcing the signals that matter.

Rage against the machine, or learn how the machine reads?

Publishers are trying to restrict machines and become more valuable to them at the same time. They can block training, permit search, license archives, negotiate commercial access, allow one crawler and reject another. AI companies, in turn, can pull information through indexes, retrieval systems, partnerships and direct licensing rather than scraping everything in sight.

Cloudflare’s granular crawler controls point to where the web is heading: access is becoming purpose-specific, not binary. For marketers, the old rule “publisher blocks AI bots, therefore PR is wasted” no longer holds. Neither does the claim that a prestigious placement automatically delivers AI visibility. The real question is whether a brand is building a credible, consistent, current and retrievable body of evidence across the wider information ecosystem.

That still requires owned content, solid technical architecture, reviews, specialist authority and, when the economics make sense, major media. The future of authority is not one citation on one famous site. It is a network of corroborating sources that machines can find, reconcile and trust.

Brands that understand the difference between being used to train the machine and being retrieved by it will be playing a different game from those still trying to block or please every bot that arrives.

Don’t miss the next one.

Get our latest pieces in your inbox. Leave whenever you want.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.