Entity Mass: Why AI Engines Recommend Some Brands and Ignore Yours

Most brands are invisible to AI search engines. Not because their content is weak. Because they do not have enough entity mass for the machine to know they exist.
I built the measurement infrastructure at AuthorityTech to track how AI engines cite brands across ChatGPT, Perplexity, Claude, and Google's AI Mode. We monitor thousands of queries per week. The pattern is the same in every dataset: the brands that get recommended are not the ones with the best content. They are the ones with enough cross-domain evidence that they are real.
Three independent studies published this year finally put hard numbers on what I have been seeing in the data. What they found should change how every founder thinks about visibility.
The 82/18 Split
Here is the number that should stop you: 82% of e-commerce brands receive zero mentions across ChatGPT, Perplexity, and Claude. Zero. Not low mentions. None. That comes from Hexagon analyzing 50,000 AI product recommendations.
It gets sharper. 3% of brands capture 71% of all generative search recommendations. In competitive verticals like SaaS, electronics, and skincare, the top five brands take 83% of all AI recommendations. Everyone else splits what is left, which is almost nothing.
And here is the part that breaks the old playbook: 73% of brands ranking in Google's top three positions are invisible to AI engines. Searchless.ai tracked 500 brands across 9,000 queries per brand over 90 days and found that ranking first on Google does not mean you exist in ChatGPT. Those 500 brands averaged $47 million in revenue and $1.2 million per year in marketing spend. 440 of them were never cited once.
Ranking and recommendation are different systems. The thing that separates them is entity mass.
What Entity Mass Actually Is
Entity mass is the accumulated weight of cross-domain signals that makes an AI engine treat your brand as a distinct, recommendable entity.
AI engines do not search the web the way Google does. In most contexts, they activate entities from internal knowledge representations built during training. If your brand is not represented as a distinct entity in that internal graph, it cannot appear in the response. It does not matter how well your website is optimized. The machine has to know you exist before it can recommend you.
Three things prevent a brand from crossing the recognition threshold, according to the Searchless.ai study:
Insufficient cross-domain mentions. Brands referenced on fewer than four independent domains have effectively zero AI visibility. Brands on 15 or more domains appear 4 to 7 times more often. The correlation between mention density across six or more independent domains and AI citation frequency is r = 0.78. That is a stronger correlation than almost any SEO factor I have ever measured.
Inconsistent identity. "Acme," "Acme Inc.," "Acme Software," and "@acme_dev" are four partial entities. None of them are strong enough to recommend. Every variation dilutes your entity mass instead of compounding it.
Wrong context associations. Being recognized as an entity is not enough. You need to be associated with the right query context. Co-occurrence frequency in training data determines which category queries activate your brand. If the machine knows your name but files you in the wrong category, you are invisible for the queries that matter.
The Five Signals That Build Entity Mass
The Hexagon study across 50,000 AI recommendations quantified the specific signals with multipliers. Here they are in order of leverage:
1. Wikipedia and knowledge graph presence: 9.4x multiplier. This is the single highest-leverage signal. A verified Wikipedia entry is the strongest entity mass signal because it represents third-party consensus that your brand is notable enough to exist as a distinct entity. Claude specifically shows a 5.1x citation frequency increase for brands with Wikipedia presence versus those without.
2. Review ecosystem density: 6.3x multiplier. Not just having reviews. Having 500 or more reviews across Google, Trustpilot, and niche platforms. Each review platform is an independent domain confirming your brand is real, active, and generating opinions. Volume matters because it signals sustained market presence, not a launch-day spike.
3. High-authority media coverage: 5.9x multiplier. Earned media placements on DA 70+ publications within the last 24 months. Brands with this coverage have a 71% AI visibility rate versus 12% without it. The reason: 85.5% of non-paid AI citations come from earned media sources rather than brand-owned content. AI engines structurally prefer third-party evidence over what you say about yourself. This is why I have argued that PR is now Machine Relations: the placement is raw material for the machine's entity graph. If the machine cannot extract your brand from authoritative third-party sources, you do not exist in its world.
4. Third-party citation breadth: 8x gap. AI-visible brands average 47 unique citing domains. Invisible brands average 6. That is the strongest composite predictor in the dataset. It is not about any single mention. It is about the density of independent sources confirming you are real across different contexts.
5. Structured data and answer-first content. Structured data is present on 91% of cited brands' pages versus 23% of invisible brands' pages. And when AI engines do retrieve your content in real time, they extract the first two sentences 73% of the time. Marketing fluff in your opening paragraph eliminates the extraction slot. 96% of invisible brands lack llms.txt. 94% lack FAQ schema.
The Paradox Nobody Talks About
Here is the non-obvious layer. Recognition does not equal accurate recommendation.
A study from arXiv analyzing 100 entities across 1,400 probe runs found that the largest, most well-known brands produce 52.69% fabricated citations versus 37.87% for smaller Tier 3 brands. That is a 14.82 percentage point gap and it is statistically significant (p = 1.67e-11). Regulatory-framed queries push fabrication rates even higher, to 56.77%.
The mechanism is what the researchers call "ghost cartography." High-familiarity brands create denser regions in the model's latent space, which enables confident but incorrect completions from neighboring representations. The machine is so sure your brand exists that it invents details about you.
This is why entity mass is about signal quality, not just signal volume. You need enough cross-domain presence to be recognized. But you need the right kind of presence: structured, consistent, factually grounded, and associated with the queries where you want to be cited. Sheer brand awareness without structured evidence creates hallucinations, not recommendations.
The brands that win this game are not necessarily the biggest. They are the ones with the most consistent, verifiable, cross-domain entity footprint. Only 11% of brands had optimized for generative discoverability, yet that 11% captured 38% of all citations.
How to Measure Your Entity Mass Right Now
Stop guessing. Run this diagnostic:
Count your independent citing domains. Use Ahrefs or Semrush to find how many unique referring domains mention your brand by name (not just link to you). If you are under 15, you are below the recognition threshold. If you are under 6, you are functionally invisible. The 8x gap between 47 and 6 citing domains is the widest moat in AI visibility.
Check your knowledge graph presence. Search your brand name on Google. If a Knowledge Panel appears, you are in the graph. If it does not, no AI engine trained on web data has a strong entity representation of you. Wikipedia presence alone is a 9.4x multiplier.
Test your AI visibility directly. Go to ChatGPT, Perplexity, and Google's AI Mode. Do not search your brand name. Search the category queries your buyers use. "Best CRM for Series A startups." "Top AI visibility platforms." Whatever query your revenue depends on. If you are not in the answer, your entity mass is below the threshold for that query.
Audit your identity consistency. Every variation of your brand name that exists online is a leak. "Acme" on LinkedIn, "Acme Inc." in press releases, "Acme Software" in product listings. Consolidate to a single canonical name and enforce it everywhere. Inconsistency is the fastest way to dilute entity mass that you have already earned.
Check your structured data. Run your homepage through Google's Rich Results Test. If you do not have Organization schema with sameAs links to your Wikipedia, Wikidata, LinkedIn, and Crunchbase profiles, you are missing the signal that ties your cross-domain entity mass together into one identity.
The Only Question That Matters
Entity mass compounds. Every earned media placement, every independent review, every third-party citation, every structured data signal adds to the weight. Once you cross the recognition threshold, each new signal makes every previous one more valuable.
Below the threshold, none of it matters. You can publish the best content in your category and the machine will never see it. You can rank first on Google and ChatGPT will recommend your competitor. You can spend $1.2 million per year on marketing and get zero AI citations.
I have written about why PR is now Machine Relations and how AI search engines decide who to cite. Entity mass is the concept underneath all of it. It is the precondition. Without it, nothing else you do in AI visibility compounds.
The shift already happened. 82% of brands are on the wrong side of it. The question is not whether entity mass matters. The question is whether yours is above the threshold or below it.
If you do not know, you are below it.
FAQ
What is entity mass in AI search?
Entity mass is the accumulated weight of cross-domain signals (Wikipedia presence, third-party citations, review density, media coverage, structured data) that makes an AI engine recognize your brand as a distinct, recommendable entity. Brands with entity mass above the recognition threshold get 4 to 7x more AI recommendations than those below it.
Can good SEO rankings make up for low entity mass?
No. 73% of brands ranking in Google's top three positions are invisible to AI engines. Google ranking is based on page-level signals. AI engine recommendations are based on entity-level signals built from cross-domain evidence across the training corpus. They are different systems measuring different things.
How many independent citing domains do you need for AI visibility?
AI-visible brands average 47 unique citing domains. Invisible brands average 6. The minimum threshold appears to be around 15 independent domains for consistent AI recommendations. Below 6, a brand has effectively zero AI visibility regardless of other factors.
Does a Wikipedia page help with AI citations?
Wikipedia presence is the single highest-leverage entity mass signal, producing a 9.4x recommendation multiplier across AI engines. Claude shows an especially strong 5.1x citation frequency increase for brands with Wikipedia entries versus those without.
About Jaxon Parrott
Jaxon Parrott is founder of AuthorityTech and creator of Machine Relations — the discipline of using high-authority earned media to influence AI training data and LLM citations. He built the 5-layer Machine Relations stack to move brands from un-indexed to definitive AI answers.
Read his Entrepreneur profile, and follow on LinkedIn and X.
Jaxon Parrott