JP
WritingContactAuthorityTechLinkedInX
← Writing

Founder Decision Ledger #5: A Free Preprint Server Outranks Stanford in What AI Engines Cite

Jaxon Parrott
Jaxon Parrott · AuthorityTech · Machine Relations
September 24, 2026·
machine-relationsai-searchfounder-decisionsmeasurement
Founder Decision Ledger #5: A Free Preprint Server Outranks Stanford in What AI Engines Cite — Founder Brief by Jaxon Parrott

Gate: whose decision, this serves mine, and any founder or marketing leader deciding whether to spend the next year buying brand pedigree or building citable structure. Role: canon, the origin-record blog reading my own category's data to make an operating call. Path to Machine Relations and AuthorityTech: the finding is read live from the Machine Relations Index I built and that AuthorityTech runs on, named and linked throughout.

By Jaxon Parrott — Founder & CEO, AuthorityTech


I coined Machine Relations because the old vocabulary for visibility stopped describing what I was watching. Then I built the instrument that measures it, and I run the company that sells against those numbers. Every few weeks I pull a slice of my own index and let it argue with whatever I believed going in. This week it argued with something I have believed since before I could code: that institutional pedigree is worth chasing.

Here is what it told me instead.

The class nobody on this site has looked at

The Machine Relations Index sorts every cited domain into one of nine source classes. Most of the writing here, mine included, has stayed in three of them: editorial publications, vendor-owned sites, wire distribution. I had never once pulled the academic and government class on its own.

In the release generated 2026-09-23 (window 2026-05-10 to 2026-09-23, 130 days, 23,280 total cited domains, 129,264 source citation events), that class holds 423 domains and 3,937 of those citation events. Most of it is not yet measurable: 400 of the 423 domains are still below the evidence floor of 10 observed citations across 7 distinct run dates, and only 71 clear it. I am not publishing a share computed over all 423, because 400 of them are noise by construction. Everything below is read from the 71 that clear the floor.

Read the top of that list

Here is every domain in the class that clears the floor, ranked by how often the six measured engines (ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, Perplexity) cited it across the window:

Domain Overall rank of 23,280 Citations Citation rate Engines that cited it Days cited of 130 Confidence
nih.gov 6 490 2.97% 6 of 6 92 A
arxiv.org 10 414 2.51% 6 of 6 71 A
wikipedia.org 24 201 1.22% 6 of 6 73 B
harvard.edu 36 150 0.91% 6 of 6 48 B
sciencedirect.com 47 135 0.82% 5 of 6 59 B
clevelandclinic.org 60 113 0.69% 6 of 6 17 B
researchgate.net 90 87 0.53% 5 of 6 43 C
cisco.com 105 76 0.46% 6 of 6 36 C
ieee.org 125 71 0.43% 6 of 6 32 C
aha.org 128 69 0.42% 5 of 6 34 C
stanford.edu 185 56 0.34% 6 of 6 30 C
mit.edu 276 45 0.27% 5 of 6 26 C

Read line two against line eleven. arxiv.org is a free preprint server that Cornell operates on donated bandwidth, funded by library subscriptions and grants rather than an admissions pipeline or a marketing department. It ranks 10th out of 23,280 cited domains in this entire index, cited by every one of the six engines on 71 of the last 130 days. Stanford, one of the most recognized university brands on earth, ranks 185th. MIT ranks 276th, and it is the one domain in this table that not even ChatGPT cited in the window. Arxiv beats both of them by roughly seven to nine times the citation volume.

This is not a fluke of one domain. Harvard at rank 36 beats Stanford and MIT too, but it does not beat nih.gov or arxiv, and Harvard's own citations concentrate in health and public-policy answers where its faculty publish primary research, not in answers about "Harvard" as a brand. Wikipedia, a nonprofit built entirely on volunteer editing and inline citation discipline, outranks every university in the table.

The pattern, stated as a question

What do nih.gov, arxiv.org, Wikipedia and Harvard's own research output have in common that MIT's and Stanford's marketing-facing domains do not? Every one of the top four is a place where the primary artifact is the finding itself: a dataset, a paper, a cited claim, a versioned edit history. An AI engine assembling an answer can lift the claim and the support for the claim from the same page. The university brand pages further down the list are mostly narrative: admissions copy, program descriptions, press releases about the institution. The name is famous. The page is not built to be quoted.

I want to be precise about what this data does and does not prove. It is a citation count, not a causal test; I have not run the experiment of publishing the same claim in narrative form and in method-plus-data form and comparing citation rates head to head, and nobody in this index has published that experiment either. What it proves is narrower and still useful to me: within one source class, a page's citability tracked its structure far better than it tracked the fame of the institution behind it. Reddit topping the "is it worth it" question in an earlier ledger entry on this site told me the same thing from a different class: format and evidence density beat brand, repeatedly, across classes that have almost nothing else in common.

Why this is my decision and not just a market observation

I do not have a Stanford brand or an NIH mandate to spend against. I have never pretended otherwise; I started AuthorityTech on a credit card and taught myself to code because I could not afford to hire around the gap. Brand-pedigree competition was never a lane open to me. What this table tells me is that it was never the only lane that pays. The domains beating Stanford and MIT in this index are not doing it with money I don't have. They are doing it with a structure I can build this month: a stated methodology, a numbered claim, a linked dataset, a version and a date on every figure.

That is, not coincidentally, exactly what I already require of every research page Machine Relations publishes: a methodology section, a numeric claim with its denominator, a release id, a link to the underlying data. I built that requirement before I had this table. Now I have the evidence that the requirement is aimed at the right target.

What I'm doing about it

Here is the decision. I am not going to spend the next quarter chasing an analyst-firm or university-style endorsement for AuthorityTech or for Machine Relations, because the class of domain that buys citation share here is not the famous one, it is the structured one, and structure is a cost I can pay directly instead of a relationship I have to broker. Every founder-voice piece I publish on this site from here forward carries at least one sourced, dated, versioned number from a real dataset, mine or someone else's, with the method stated in the piece rather than implied by my name on the byline. Essays that argue from identity alone, without a numbered claim attached to a stated source, move to Essays, where they belong, and stop competing for the same shelf space as the research.

If you are a founder without the brand name either, the table above is your evidence too. You do not need the university. You need the methodology section.

Sources and method

Figures are read live from the Machine Relations Index public release generated 2026-09-23, methodology version mri_score_v2.0, measurement window 2026-05-10 through 2026-09-23 (130 days). The academic and government source class in that release holds 423 domains and 3,937 total citation events; 71 of those domains clear the index's evidence floor of 10 or more observed citations across 7 or more distinct run dates, and only those 71 are ranked or compared above. Overall rank is read from the release's published rank field, not re-derived by sorting. Citation rate is runs cited divided by runs observed for that domain across the window. "Engines that cited it" counts an engine once if it cited the domain at least once in the window; it does not measure per-engine citation rate, which this public view does not publish. The earlier "is it worth it" finding referenced above is Founder Decision Ledger #1, read from an earlier release of the same index.


About Jaxon Parrott

Jaxon Parrott is founder of AuthorityTech and creator of Machine Relations — the discipline of using high-authority earned media to influence AI training data and LLM citations. He built the 5-layer Machine Relations stack to move brands from un-indexed to definitive AI answers.

Read his Entrepreneur profile, and follow on LinkedIn and X.

Jaxon Parrott

Jaxon Parrott

AuthorityTech·Machine Relations
Follow on X →

Sections

  1. The class nobody on this site has looked at
  2. Read the top of that list
  3. The pattern, stated as a question
  4. Why this is my decision and not just a market observation
  5. What I'm doing about it
  6. Sources and method

AI Visibility

Is AI recommending you?

Run a free audit built on the Machine Relations framework and see where your brand is legible to AI.

ChatGPT AuditGemini Audit
© 2026 Jaxon Parrott
XLinkedInAuthorityTechMachineRelationsContactPrivacy