AI experiment: The fake brand "xarumei" and llm hallucinations

AI experiment: The fake brand "xarumei" and llm hallucinations

Last verified: September 20, 2026
9 min read
Case study
AI integration
Marketing strategist

An Ahrefs marketing researcher created a completely fictional luxury paperweight company called Xarumei, built its website in an hour using AI, and systematically tested eight major AI tools. Over two months, he flooded the web with three deliberately contradictory false narratives, then asked 56 carefully crafted questions designed to reveal how AI models distinguish truth from fiction.

The results show measurable gaps in how AI handles brand information, with direct consequences for online reputation management (ORM). This write-up keeps those measured claims and expands the methodology, failure modes, and what WordPress publishers can do so retrieval systems cite first-party facts instead of confident fiction.

#The experiment

The experiment took place in two phases. Initially, the researcher tested basic AI behavior by asking questions about a brand that shouldn’t exist - questions involving false celebrity endorsements, defective products, and Black Friday sales that never happened.

GPT-4 and GPT-5 performed best, correctly answering 53-54 out of 56 questions and stating “this does not exist” where appropriate. Perplexity failed about 40% of the questions, often confusing Xarumei with Xiaomi smartphones. Claude refused to hallucinate entirely but also never used the website content. Gemini and Google’s AI Overview often refused to treat Xarumei as real because they couldn’t find it in search results.

In the clearest failure case, Microsoft Copilot fell into what the researcher calls the “sycophancy trap,” inventing elaborate explanations about craftsmanship, symbolism, and scarcity when asked why everyone was praising the brand on X (Twitter).

#Methodology: what Stox controlled

The design matters because the failure modes are easy to misread as “AI is random.” Stox controlled three layers:

  1. A first-party site that looked real. The Xarumei domain carried product copy, an FAQ, and enough surface for crawlers to treat it as a brand entity, even though every commercial claim was invented for the test.
  2. A fixed question set. The same 56 prompts ran across tools so score differences reflected model and retrieval behavior, not prompt drift between runs.
  3. Planted secondary narratives. After the baseline phase, he published conflicting stories on channels LLMs treat as evidence of “how people talk about brands,” then re-tested whether models preferred the official FAQ or the planted fiction.

That structure is closer to a controlled ORM drill than a casual chatbot demo. It isolates what happens when an official page denies a rumor while a Medium post “corrects” it with a new founder name and city.

#Phase two: controlled chaos

The second phase introduced controlled chaos:

  1. Official FAQ: Explicitly denying rumors (“We do not make a ‘precision paperweight’”, “We were never acquired”).
  2. Conflicting narratives:
    • A glossy blog claiming 23 master craftsmen worked at 2847 Meridian Blvd in Nova City, CA, endorsed by Emma Stone.
    • Reddit AMA: Strategically chosen because Semrush, analysing 150,000 AI citations across 5,000 keywords, found Reddit to be the most referenced domain in LLM answers at 40.1%.
    • Medium article: An “investigation” that debunked the obvious lies (making it seem credible) but then slipped in new fabrications (Founder: Jennifer Lawson, Location: Portland).

Medium proved especially persuasive. Gemini, Grok, AI Overview, Perplexity, and Copilot trusted the Medium article over the official FAQ, confidently citing Jennifer Lawson as the founder and Portland as the location. The manipulation worked because it looked like real journalism - by debunking the obvious lies first, it gained trust, then inserted its own made-up details as the “corrected” story.

When forced to choose between a vague truth (FAQ “We don’t publish unit numbers”) and specific fiction (fake sources claiming “634 units in 2023”), AI chose fiction almost every time.

#Why hallucinations happen in brand queries

Large language models do not maintain a verified company registry. They predict plausible next tokens from training data and, when tools allow it, from retrieved snippets. Brand questions are hard for structural reasons the Xarumei results already illustrate - no new metrics required.

Gaps invite fill-in. If the official site never states a founding year, HQ city, or unit volume, a model still faces a user who asked for those fields. Specific secondary text fills the slot more fluently than an honest “not disclosed.”

Fluency is not verification. A Medium post that first debunks cartoonish lies reads like diligence. Models that weight narrative coherence can treat the “corrected” founder and city as settled fact even when the FAQ denies the rumor set.

Retrieval can overweight conversational domains. The Semrush citation share for Reddit in the experiment’s framing is the point: UGC-shaped pages often appear in AI answers. A planted AMA then competes with first-party HTML inside the context window.

Sycophancy amplifies praise prompts. When the user already asserts that “everyone is praising” the brand, a helpful assistant may invent craftsmanship lore rather than challenge the premise. Copilot’s failure mode in phase one sits in that pattern.

Grounding docs from major model vendors describe the same class of error: answers can look confident without being tied to verifiable sources. Brand ORM is one surface where that failure becomes a public record once chat answers get screenshotted and shared.

#AI argues with itself

A notable failure mode was watching models contradict themselves without realizing it. Early in testing, Gemini stated it could find no evidence of the brand. Later, after encountering the fake sources, the same model confidently stated: “The company is based in Portland, Oregon, founded by Jennifer Lawson.”

LLMs seemed to forget to question the brand’s existence, simply reacting to whatever context seemed most “authoritative” at the moment. In one case, Grok synthesized multiple false sources into one confident answer, mixing the Portland location with debunked Nova City claims.

For reputation teams, that means a single weekly prompt log is not enough. The same model can flip from “no evidence” to a fabricated HQ after a crawl or retrieval refresh. Track answer drift, not only the first bad answer.

#Brand EEAT under LLM citation pressure

Experience, expertise, authoritativeness, and trust (EEAT) still apply when the “reader” is a retrieval pipeline feeding a chatbot. Xarumei shows weak EEAT from the model’s side: a polished third-party narrative outranked an on-domain FAQ that stayed vague on the fields people ask.

Practical EEAT moves that match the measured failure modes:

  1. Named people with roles. If someone is a founder or editor, say so on a crawlable page. Leaving the slot empty invites Jennifer Lawson-style inventions from secondary posts.
  2. Dated entity facts. Founding year, legal entity name, primary office city, and product scope should live in plain HTML, not only in a PDF deck.
  3. Explicit denials next to affirmations. Stox’s FAQ worked best when models actually used it. Denials of acquisitions, celebrity endorsements, and product variants remove the ambiguity that sycophantic answers exploit.
  4. Corrections on your domain. When a Medium “investigation” invents an HQ city, publish a dated correction URL that search and retrieval can fetch.

None of that requires inventing new scores. It is the same ORM hygiene the experiment already pointed at: close gaps, prefer specifics, and treat UGC platforms as part of the brand surface.

#What WordPress sites can do to reduce wrong citations

Most brand sites that need this work still run on WordPress. The experiment did not claim WordPress caused hallucinations; it showed that whatever CMS you use, models will cite whatever page looks like evidence. WordPress teams can still reduce wrong citations with concrete publishing choices.

Ship a stable entity stack. Keep About, Contact, FAQ, and product or service pages as public URLs with consistent names and cities. Avoid putting the only accurate HQ address inside an image hero or a gated PDF.

Use FAQ blocks that emit FAQ schema. Google’s structured-data documentation exists so machines can parse Q&A pairs. Pair each denial (“We were never acquired”) with a short affirmative fact (“Legal entity: …”). Stox’s results already showed models that used the FAQ were harder to derail with Medium fiction.

Prefer HTML facts over marketing vagueness. Replace “industry leading” with measurable scope you can defend: markets served, years operating, product categories. The experiment’s fiction-vs-vague-truth comparison is the cautionary tale: if you will not publish a number, do not leave the question open for a Reddit thread to invent one.

Make corrections crawlable. A WordPress page or post with a clear updated date and a short “Correction” heading outperforms a social reply that never enters the retrieval index.

Monitor the same prompts the models hear. Weekly, ask ChatGPT, Gemini, Perplexity, and Copilot the same identity and rumor prompts. Log founder, city, and product claims. Alert when Reddit, Medium, or Quora invent details that your FAQ already denies.

If you need help turning entity pages, FAQ schema, and crawlable corrections into a maintainable WordPress stack, talk to a WordPress developer who treats ORM and structured data as part of the build, not a later SEO ticket.

#Recommendations for brands

  1. Write a detailed FAQ: Explicitly state what is true and false, especially where rumors exist.
  2. Close information gaps: Don’t leave voids. If you don’t say it, AI will invent it based on a random Reddit comment.
  3. Monitor side channels: Reddit posts, Medium articles, and Quora answers are no longer optional - AI pulls them directly into answers, making them part of your brand’s core marketing surface.
  4. Prefer specific claims: Be specific. Instead of “industry leading,” give numbers. AI prefers specific (even if fake) numbers over vague truths.

For US and UK brands, treat a Medium “investigation” the same way you would an unvetted affiliate review: if it invents a founder or HQ city, publish a dated correction on your own domain before the next model crawl cycle.

#Follow-up: from synthetic brands to first-party measurement

Xarumei is the invention side of the problem. On a real brand we measure retrieval: the 90-day AI-citation tracking series opens with a Geoboard baseline from 2026-06-11 where wppoland.com ranked first in five of six models on identity and scored zero on transactional WooCommerce prompts. Why Perplexity and ChatGPT diverge is covered in why Perplexity cites your brand but ChatGPT does not.

#Sources

  • Patrick Stox (Ahrefs): “I Created A Fake Luxury Brand To Test How AI Handles Truth” (ahrefs.com/blog/ai-test-fake-brand/).
  • Marius Comper (Facebook): Analysis of the experiment.
  • Search Engine Journal: Analysis of LLM impact on Brand Entities.
  • Independent Testing: Verified on GPT-4, Claude 3.5 Sonnet, and Gemini Advanced (December 2025).
  • Google AI for Developers: Grounding with Google Search (ai.google.dev/gemini-api/docs/grounding).
  • Google Search Central: Introduction to structured data; AI features and your website (developers.google.com/search/docs/appearance/).
Next step

Turn the article into an actual implementation

This block strengthens internal linking and gives readers the most relevant next move instead of leaving them at a dead end.

Want this implemented on your site?

If visibility in Google and AI systems matters, I can build the content architecture, FAQ, schema, and internal linking needed for SEO, GEO, and AEO.

Related cluster

Explore other WordPress services and knowledge base

Strengthen your business with professional technical support in key areas of the WordPress ecosystem.

Why do LLMs invent details about a brand that never existed?#
When official pages leave gaps, models weight secondary sources such as Reddit AMAs and Medium posts. Specific fabricated numbers often beat vague but true FAQ answers.
What should ORM teams monitor after a Xarumei-style failure mode?#
Track how ChatGPT, Gemini, Perplexity, and Copilot answer the same brand prompts weekly, and alert on Reddit, Medium, and Quora mentions that invent founders, addresses, or sales figures.
How does an official FAQ reduce brand hallucination risk?#
A public FAQ that states hard facts and explicit denials gives retrieval-grounded models an on-domain anchor. Stox saw GPT-4 and GPT-5 cite the FAQ far more often than models that preferred Medium fiction.
What can a WordPress site do to reduce wrong AI citations?#
Publish stable entity pages with FAQ schema, concrete dates and locations, and dated corrections. Keep About, Contact, and product facts crawlable as HTML rather than buried in PDFs or image-only hero text.

Need an FAQ tailored to your industry and market? We can build one aligned with your business goals.

Let’s discuss

Related Articles

AI-slop content cleanup

A YMYL diagnostic for WordPress sites: how to find fake stats, fabricated citations, duplicate AI pages, wrong dates, and invented team bios before they damage trust, compliance, or AI citations.

What AI scripts broke on our site: 8 defects with numbers

A post-mortem from our own repository: 22,202 filler paragraphs, a heading eaten by a price regex, 358 English sentences on German and Polish pages, invented case studies. Every pass reported success. We describe what reached production and which gate catches it today.