Investor Relations Websites

How Search Engines and AI Systems Find Public Company Information

VendorGroup

Understand how crawling, indexing, retrieval, and user-requested access affect discovery of public-company information on investor relations websites.

Search engines and AI services can encounter public-company information through links, crawlers, indexes, retrieval systems, or a user's request to open a page. Those routes are related but not identical. Understanding the difference helps an IR team diagnose why information is missing, outdated, or cited incorrectly.

Start by asking which stage is failing. A page may be discoverable but blocked from crawling, accessible but not indexed, or indexed without being selected for a particular answer.

Discovery identifies a possible source

Links and sitemaps can point a service toward public URLs. The website should expose important investor destinations through ordinary navigation and contextual links rather than relying only on a visitor's interaction with a filter.

Check the path from the IR homepage to a results period and then to its documents. Make each step descriptive and usable. The SEO guide explains how this connects to the site's content structure.

Discovery does not mean the page's contents have been processed. It is the beginning of the route, not the outcome.

Crawling retrieves the page

A crawler requests content from the website. The server, content-delivery network, firewall, and robots directives can affect that request. Test the public response and the relevant operator logs when investigating an access problem.

Google describes crawling, indexing, and serving as separate stages. Not every discovered page proceeds through all three. [1] A successful browser visit is useful evidence, but it does not prove that a crawler received the same content or response.

Do not expose drafts to solve a discovery problem. Use real access controls for confidential information, and manage crawl permissions for public material deliberately.

Indexing organizes information for later use

After retrieval, a search engine may analyze and index the content. Technical access alone does not guarantee inclusion. Duplicate content, page directives, canonical choices, and the engine's processing decisions can affect what is represented.

If the site uses noindex, the crawler needs access to read it; blocking the URL in robots.txt can prevent that. [2] Review the intended behavior of each control rather than treating all exclusion settings as interchangeable.

Use inspection tools to check a known URL's status and processed content where available. Compare a problem page with a similar working page to narrow the investigation.

Retrieval selects sources for a question

A search or AI experience can select indexed or otherwise available information relevant to a request. Eligibility is not a promise that a particular page will be chosen. Google documents Search-based requirements for supporting links in its AI features, while Bing's guidelines connect indexing and content clarity with grounding eligibility. [3] [4]

For IR teams, the practical response is to keep the official page understandable on its own. Identify the issuer, subject, date, and original supporting material. A clear source reduces ambiguity in the information you publish, even though external interpretation remains outside your control.

Distinguish OpenAI search from other access

OpenAI identifies OAI-SearchBot as its search crawler, GPTBot as a crawler for potential training use, and ChatGPT-User as an agent for certain user-initiated visits. ChatGPT-User is not its automatic search crawler, and robots rules may not apply to those user-requested actions. [5]

That distinction matters when reviewing an “AI bot” policy. Choose controls for the intended use case and verify them against current documentation. A training opt-out should not be assumed to be a search opt-out, or vice versa.

Use a disciplined troubleshooting sequence

Identify the affected URL and expected use. Verify the live page, access rules, server response, rendered content, indexing state where available, and links from relevant pages. Then test the specific search or AI question and record the result.

Coordinate infrastructure issues with the hosting and security owner. Use the AI visibility guide for the editorial and measurement work once the technical path is understood.

Questions companies ask

Does a crawler visit mean the page is indexed?

No. Retrieval and indexing are separate events, and a log entry alone does not establish inclusion or ranking.

Does indexing guarantee an AI citation?

No. Source selection depends on the service and request. Treat citation outcomes as observations to test.

Can robots.txt keep confidential information private?

Do not use crawler instructions as the confidentiality mechanism. Protect private content with access controls.

Why might an old page still appear?

Investigate the live content, redirects, indexing, and the service's processing state. Do not assume an edit immediately updates every external system.

Use this sequence in a VendorGroup technical review so the work addresses the stage where discovery is actually failing.

Related VendorGroup resources

Primary sources

  1. Google — How Search Works
  2. Google — Block Search Indexing with noindex
  3. Google — AI Features and Your Website
  4. Bing — Webmaster Guidelines
  5. OpenAI — Overview of OpenAI Crawlers

Contact

VendorGroup®