FAQ
Common questions
Start with 11 core questions. Local businesses, optional tools and site history follow below.
Start here
What are discovery files, and what does this guide cover?
“Discovery files” is our informal label for supporting files, not a universal AI standard. This guide focuses on useful pages, AI search access and measurement, with the SEO foundations those need.
How do I start?
Choose a few customer questions, publish useful answers, check access, add formats with a clear purpose, then review results. Follow the five steps in the Quick guide.
How do SEO and GEO differ?
SEO means search engine optimisation; GEO means generative engine optimisation, focused on AI-generated answers. Google and Bing connect their AI search guidance with established search foundations.
Is AI search optimisation mainly a technical task?
Technical access and useful content both matter. Fix access problems and answer customer questions with original evidence and clear expertise.
Where should I publish business facts, brand guidance and answers?
Put identity and contact facts on About and Contact pages, and service details and answers on the relevant product or service pages. Use a brand guide for specialised brand guidance.
Which files does my website need for AI discovery?
Start with useful public pages. Use robots.txt for crawl rules, a sitemap for URL discovery and relevant structured data. llms.txt is optional; choose it for a tool or workflow that uses it.
Will publishing llms.txt improve my visibility in AI answers?
It depends on the system. Google says llms.txt provides no visibility benefit in its Search AI features. Other tools may use it for navigation or context; that alone does not establish improved search visibility.
Is what a model learned the same as what it searches for?
No. Training changes a model’s parameters; search retrieves information while answering. A product can use both. A citation does not establish that a page was used in training.
Do Google, Bing and AI assistants all search the same way?
No. Search indexes, retrieval tools and controls vary by product and setting. A model name alone does not identify the sources an answer used.
Which AI crawlers should I allow or block?
Choose separately for search, training and user-triggered retrieval. Consult each provider’s documentation; the crawler reference links to it.
Do AI systems have to follow preferences published on my website?
Do not assume they do. Use documented provider controls for crawling, and check whether a particular tool supports other published preferences.
Local businesses
What are local SEO and local GEO?
Local SEO concerns searches connected to a place or nearby need. Here, local GEO applies AI search work to location-related answers; GEO does not mean geolocation.
Why can the same search show different businesses in different places?
Results can depend on the searcher’s location and the place named in the question. Record both, together with the product and exact question, before comparing businesses shown.
Should I connect my website to company registers and business profiles?
Yes, where relevant and supported. Accurate records and eligible profiles help people check your identity. They do not establish a universal AI trust score or guarantee citations.
Is a service area the same as having a business location there?
No. A service area describes where you visit or deliver to customers, not a branch in every town. State your real coverage and whether customers can visit your premises.
Does every website qualify for Google and Apple business listings?
No. Eligibility depends on the provider, feature and region. Check the current requirements in the local-search reference; do not invent a physical location to obtain a listing.
Optional files and tools
How does llms.txt differ from robots.txt?
robots.txt gives cooperating crawlers access instructions. llms.txt provides context and links to useful content. It does not control crawler access.
Where should I publish an optional llms.txt file?
The upstream proposal supports /llms.txt or a file under a relevant path, such as /docs/llms.txt. Follow the intended tool’s requirements and check the published response.
Do I need a developer to publish llms.txt?
It depends on your host. Some let you upload a text file directly; others need configuration or a developer.
How often should I update llms.txt and other content summaries?
Update summaries when their facts or links change. Generate them from the same source as your pages where practical.
My file is published but nothing seems to read it. What is wrong?
Check the body, HTTP status and content type, including whether the URL returns an error page with status 200. If it works, investigate whether your intended tool actually requests it.
How do I check a file is valid?
Check the current format specification, the intended tool’s requirements and the published response. Format validity is separate from visibility.
Can llms.txt and other text summaries be indexed by search engines?
Yes. Google says it can fetch and index text files as ordinary content, without special treatment. Keep summaries accurate and avoid creating unnecessary copies to maintain.
What information should I keep out of public website files?
Do not publish credentials, API keys, internal or staging URLs, confidential information, or private personal data in public files.
How do AGENTS.md, CLAUDE.md and GEMINI.md relate to discovery files?
They give project instructions to supporting coding tools. Keep them with the project; publishing one on your domain does not establish that search engines will read it as instructions.
Do I need all three agent instruction files?
No. Use the files your tools actually load. Where supported, share instructions through imports or filename configuration; check each tool’s loading rules and scope.
Access and evidence
Did ChatGPT switch from Google to Bing?
Microsoft announced Bing integration with ChatGPT in May 2023. That does not establish a switch from Google or exclusive Bing use today. OpenAI also documents its own search crawler.
What else affects how AI systems represent my business?
For answers with search enabled, check the specific product and cited pages. Correct outdated facts on your own pages and identify mistakes in other cited sources. These checks assess retrieved information; they do not establish what a model learned during training.
About this site and its history
What is discoveryfiles.ai?
A free practical guide to website discovery, operated by Generell AI.
Who maintains this guide?
Generell AI operates this site. The llms.txt proposal is maintained independently at llmstxt.org.
Is the discovery-files concept obsolete, or has the specification just become smaller?
The original eleven-file proposal is archived. We have not replaced it with a smaller specification. The current guide explains individual formats by their purpose and documented support.
Does this guide define an official web standard?
No. This is an independent guide. An earlier experimental file proposal is retained in the project archive.
What happened to files such as identity.json?
We retired the site-specific experimental files. Their URLs redirect to the archive explanation, where the original templates remain available as historical examples.
How current is the provider information?
Each source shows its fetch and review dates. These record documentation checks, not live access tests. See the source-date explanation for how they differ from publisher dates.
Does it cost anything?
No. The website code and original text are MIT licensed, including for commercial reuse. Keep the copyright and permission notices. Logo and brand assets are excluded; third-party materials retain their own terms. See the license page for details.
How do I contact the operators of this site?
General enquiries: info@discoveryfiles.ai. Privacy and data protection requests: privacy@discoveryfiles.ai.
Not answered here? Email info@discoveryfiles.ai, or read the Quick guide .
Can AI search tools read Reddit, LinkedIn, Facebook, Instagram and TikTok?
Sometimes. Search results, direct requests and licensed APIs offer different routes. Published crawl rules do not prove successful access, and a citation does not establish access to private feeds or complete videos.
Social platforms and the crawler-rule snapshot