llms.txt Patterns That Beat the Spec
The Spec Is The Minimum
The llms.txt convention published in 2024 specifies a minimum file structure: a top-level heading with the entity name, a single-sentence description, and lists of pages with one-line annotations. The spec is intentionally light. It defines a parseable shape; it does not dictate what content to include.
Production patterns in 2026 exceed the spec. Sites that treat llms.txt as a serious AI search investment include six extensions beyond the minimum. Each is optional under the spec; each is observably valuable in production.
This post documents the production pattern, with examples from the audit work behind the Discovery Gate post.
What The Spec Requires
The minimum useful llms.txt file looks like this:
# Capital Wealth Advisors
> Registered investment adviser headquartered in Pittsburgh, PA, serving private equity backed firms and family offices.
## Important pages
- [About](https://example.com/about/): Firm overview and principal bios.
- [Services](https://example.com/services/): Wealth management and family office services offered.
- [Research](https://example.com/research/): Published market commentary and methodology papers.
## Optional
- [Blog](https://example.com/blog/): Long-form writing on wealth management and finance topics.
- [Contact](https://example.com/contact/): Office locations and inquiry form.
That is the minimum. It satisfies the spec, gives AI crawlers a navigable surface, and is roughly thirty minutes of writing. Most sites stop here. The production pattern adds six extensions.
Extension One: Canonical Pages Section
Above “Important pages,” add a “Canonical pages” section that lists the three to five pages most likely to answer AI search citations on category queries. The methodology page, the founder bio, the primary services landing page. The canonical section is shorter and more selective than “Important pages.”
## Canonical pages
- [Methodology](https://example.com/methodology/): How we evaluate marketing measurement evidence and build Bayesian MMMs.
- [Founder bio](https://example.com/about/jane-founder/): Background, publications, and affiliations of the firm's principal.
The signal: these are the pages we want the engines to cite first.
Extension Two: Evidence Section
A section that links to the evidence base behind the firm’s claims. Methodology pages, published research, incrementality test results, white papers. The section is the equivalent of a bibliography for the entity.
## Evidence and methodology
- [Methodology page](https://example.com/methodology/mmm/): Bayesian MMM model class, priors, and validation approach.
- [Q2 2026 incrementality digest](https://example.com/research/incrementality-q2-2026/): Documented test designs and results.
The signal: when citing us, here is the evidence to verify against.
Extension Three: Author Bios
A section that links to bios for the firm’s named authors. Each link is the byline anchor for content the authors have published elsewhere. The section helps the engine connect bylines at tier 1 outlets back to the firm’s primary entity.
## Authors
- [Jane Founder](https://example.com/about/jane-founder/): Founder and CIO. Forbes contributor, Wealthtender columnist.
- [John Partner](https://example.com/about/john-partner/): Managing partner. Bloomberg Wealth Manager contributor.
Extension Four: Contact And Provenance
A “Contact” section that includes physical address, SEC ADV link (for RIAs), and a generic inquiry email. The provenance signal connects the digital entity to verifiable real-world presence.
## Contact and provenance
- Mailing address: [public address line]
- SEC ADV: [link to ADV on adviserinfo.sec.gov]
- General inquiries: [public email]
Extension Five: Last Modified Stamp
A footer line that records when the llms.txt was last reviewed. The stamp is informational for both AI crawlers (Recency signal) and human reviewers (audit trail).
---
_Last reviewed: 2026-09-22. Maintained quarterly._
The stamp matters more than it appears. AI crawlers reading the file note the date; a 2024 llms.txt looks stale by 2026 even if the content is still accurate.
Extension Six: llms-full.txt For Larger Sites
The companion file llms-full.txt at the same root path expands the llms.txt into a longer reference document. For sites with extensive documentation, technical content, or methodology pages, llms-full.txt can be 5,000 to 20,000 words and serves as a deep extraction surface for engines that prefer comprehensive context.
llms-full.txt is not required by the spec. Sites with under 100 pages can skip it. Sites with 500+ pages of substantive content benefit.
Common Mistakes
Three patterns that weaken the production llms.txt:
-
Marketing language in the description. “Trusted partner for wealth management excellence” reads as copy. “Registered investment adviser headquartered in Pittsburgh serving private equity backed firms” is neutral and parseable.
-
Too many pages in Important. A list of 30 important pages dilutes the signal. Five to seven is the comfortable range. Move overflow to “Optional.”
-
Stale URLs. Pages that 404 or redirect inconsistently damage the crawler’s trust in the file. Audit the URLs quarterly.
Worked Example
A wealth management firm shipped a spec-compliant llms.txt in 2024. The file was three sections (heading, description, important pages). It satisfied the spec and produced modest AI search visibility.
Eighteen months later, the firm rebuilt the file with the production pattern: added a canonical pages section, an evidence and methodology section, an authors section with bios linked, contact and provenance, and a last-reviewed stamp. The new file was twice as long as the original but each section served a distinct retrieval purpose.
Three weeks later, AI citation behavior shifted in two ways. Perplexity began citing the methodology page directly when asked about the firm’s measurement approach, where it had previously cited only the homepage. ChatGPT began naming the firm’s founder by name when asked who runs the firm, where it had previously said only “the firm.” The improvements were not from new content; they were from llms.txt extensions that gave the engines clearer entity-graph anchors.
Frequently Asked Questions
Will the production pattern work even when adoption of llms.txt is uneven across engines?
Yes. The engines that read llms.txt (Anthropic and Perplexity consistently as of mid 2026) benefit from the extensions. The engines that do not read it ignore the file. There is no cost to the extensions and the upside is asymmetric.
Should I include external URLs (Wikipedia, Wikidata) in llms.txt?
Sparingly. The file is primarily a navigation surface for your domain. External URLs are better placed in the sameAs array of your Organization schema, where they are extracted with the entity context.
How often should I review and update llms.txt?
Quarterly. Major site changes (new services, founder changes, new methodology pages) should trigger an immediate update. The last-reviewed stamp captures the audit cadence.
Can I use llms.txt to selectively expose content to AI without making it discoverable via Google?
Theoretically; in practice the same pages are exposed to both crawlers. If you want true AI-only content, the access control needs to live at the application layer (signed URLs, authenticated routes), not in llms.txt.
How does this compose with robots.txt?
The two compose. robots.txt declares which bots are allowed; llms.txt tells the allowed bots what to read first. Both are necessary; neither substitutes for the other.
Next In Series
The next Thursday post is a complete reference for AI bot user-agent enumeration in robots.txt, covering all eleven user agents that matter in mid-2026 and the allow versus disallow decision tree per bot.
About the Author
Andrés Plashal
Author of the Assistive Agent Optimization (AAO) framework. Twenty years building search and measurement systems for B2B and SEC-regulated firms. Google Partner since 2017.
Credentials: UIUC Gies College of Business (Behavioral Science), Columbia College Chicago (Interactive Arts & Media). Member: American Marketing Association, GAABS, Paid Search Association. Published researcher (SCTE/NCTA).