AI search visibility starts with technical accessibility
A website cannot be selected, summarised or cited by an AI search platform if the platform cannot reliably retrieve and understand its public pages. The technical foundation is similar to conventional SEO: the site must be crawlable, return valid responses, expose indexable HTML and provide consistent signals about each page's identity and purpose.
AI search is not one universal index. Google, Bing, ChatGPT, Claude, Perplexity and other systems use different crawlers, search indexes and retrieval processes. Some automatically discover pages, while others fetch a page in response to a user request. A technically sound website gives all legitimate systems the best opportunity to access the same authoritative source material.
This does not guarantee a citation or ranking. It establishes eligibility. Relevance, source quality, authority, freshness and the user's exact question still determine which sources are used.
Separate crawling, indexing and retrieval
Crawling is the process of requesting a URL and following links to discover other URLs. Indexing is the process of analysing an accessible page and storing information about it for later search. Retrieval happens when a search or AI system selects information that may answer a particular query.
These stages require different controls. A robots.txt rule manages crawler access; it does not reliably remove an already known URL from search. A robots meta tag or X-Robots-Tag header controls indexing, but a crawler must be allowed to fetch the URL before it can read that instruction. A canonical link identifies the preferred version of equivalent content, but it is not a substitute for redirects or consistent internal linking.
Treating all of these signals as interchangeable creates conflicts. Public pages intended for discovery should be crawlable, indexable, canonicalised to themselves and linked from the website's normal navigation or content structure.
Configure robots.txt for legitimate search and AI crawlers
Serve robots.txt from the root of the production hostname, such as https://www.example.com/robots.txt. Return HTTP 200 with plain-text content and use only supported directives such as User-agent, Allow, Disallow and Sitemap. Avoid unsupported instructions or crawler-specific groups unless a genuine policy difference requires them.
A simple User-agent: * group that allows public content applies to traditional search crawlers and AI search crawlers unless a more specific group overrides it. Relevant crawler tokens can include Googlebot, Googlebot-Image, bingbot, Applebot, OAI-SearchBot, PerplexityBot, Claude-SearchBot and Amazonbot. User-initiated fetchers can include ChatGPT-User, Perplexity-User and Claude-User.
Training controls are separate from search visibility. GPTBot and ClaudeBot relate to model development, while Google-Extended controls certain Gemini training and grounding uses. Allowing or blocking a training crawler should be a deliberate publishing-policy decision, not described as a search ranking technique.
A production robots audit should confirm that:
- The homepage and every public service, location, article and case-study URL are allowed
- Public images, CSS, JavaScript and API responses required for rendering are not blocked
- Crawler-specific groups do not accidentally override important wildcard restrictions
- The XML sitemap uses an absolute production URL
- Staging and development hosts have their own access controls and cannot be indexed
Return stable HTTP responses without bot challenges
Every indexable URL should return HTTP 200 directly. Permanent moves should return 301 or 308 to a single final URL. Removed pages should return 404 or 410. Redirect loops, long redirect chains, soft 404 pages and intermittent 5xx responses reduce crawl reliability and can prevent a source from being retrieved when an answer is generated.
A content delivery network, firewall or bot-protection service can block a crawler even when robots.txt allows it. Test representative URLs using the documented user-agent strings, then verify genuine crawler traffic using the provider's published IP ranges or reverse-DNS process. Check for 403, 429 and challenge responses in CDN, WAF and origin logs.
Do not require a CAPTCHA, cookie consent interaction, geolocation prompt, login or client-side action before the main public content is delivered. Search systems need a stable URL that produces the same substantive information available to a normal visitor.
Deliver indexable HTML before relying on JavaScript
Server-rendered or statically generated HTML gives crawlers immediate access to the title, main heading, body copy, internal links and structured data. Client-side rendering can work for some search engines, but it introduces an additional rendering stage and is less reliable across the full range of AI search and user-fetch systems.
Inspect the raw HTTP response, not only the browser's rendered Document Object Model. The initial HTML should contain the primary content and links. JavaScript can enhance navigation, forms and interactive tools, but it should not be the only mechanism that inserts the page's meaning.
Keep essential CSS, JavaScript, fonts and image resources crawlable. If a crawler cannot load the resources needed to interpret the layout, important content may be treated as hidden, incomplete or unavailable.
Use precise indexing and canonical signals
Public pages should not emit noindex through a meta robots tag or X-Robots-Tag header. Check application code, reverse-proxy rules, hosting configuration and CMS plugins because an indexing header can be added outside the page template.
Each indexable page should have one self-referencing canonical URL using the preferred HTTPS hostname and a consistent trailing-slash policy. Redirect alternative hosts and URL variants to that version. Internal links, sitemap entries, Open Graph URLs and structured-data identifiers should use the same canonical address.
Do not canonicalise distinct service or location pages to a broad parent page merely because their layouts are similar. If a page serves a separate search intent and contains unique, useful information, it normally needs its own canonical URL.
Publish a clean XML sitemap
The XML sitemap should contain only canonical, indexable URLs that return HTTP 200. Include the sitemap URL in robots.txt and submit it through Google Search Console and Bing Webmaster Tools. Large websites should use a sitemap index and split child sitemaps by content type or update pattern.
Use lastmod only when it represents a meaningful content change. Updating every URL's date during each deployment creates a noisy freshness signal and makes it harder for crawlers to prioritise genuinely updated pages. Do not include redirects, 404 pages, noindex pages, duplicate parameters, private previews or staging URLs.
A sitemap supports discovery, but it does not replace normal internal links. Important pages should be reachable through crawlable HTML links from relevant hubs and related content.
Create an information architecture that exposes meaning
AI retrieval works best when a website has a clear subject hierarchy. Give each major service, location, project and customer question a stable URL with a single primary purpose. Use descriptive paths, concise title tags, one visible H1 and logical H2 sections that correspond to the questions a visitor might ask.
Connect related pages with contextual internal links. A service page can link to relevant case studies, technical explanations, service areas and frequently asked questions. An article should link back to the commercial page that establishes who provides the service. This creates a crawl path and helps search systems understand the relationships between topics.
Avoid fragmenting one answer across thin pages or placing every service on one undifferentiated page. The ideal unit is a page that completely addresses a recognisable intent without duplicating another URL.
Write content that can be extracted and verified
AI search systems often need to identify a direct answer, supporting detail and the source responsible for it. State important facts in complete sentences near the relevant heading. Use tables, ordered steps and concise lists when they are the clearest format, but support them with explanatory context.
Define specialist terms, specify scope and include dates where freshness matters. Name the service area, eligibility conditions, limitations, methodology and evidence behind claims. Replace vague superlatives with verifiable proof such as licences, project details, measured outcomes, named authors, review sources or documented processes.
Keep the page updated and show material revision dates when appropriate. Contradictory addresses, phone numbers, service descriptions or business names across the website and external profiles weaken entity confidence.
Implement structured data that matches visible content
Use JSON-LD to describe entities and page types with Schema.org vocabulary. Common types include Organization or LocalBusiness for the business, Service for service offerings, Article for editorial content, BreadcrumbList for hierarchy and WebSite for the site itself. Use the most specific valid type supported by the page.
Structured data should identify the canonical URL, organisation name, logo, contact details, service area, author, headline, image and publication dates where those properties are relevant. Stable @id values help connect the same organisation or page across multiple data blocks.
Markup must reflect information that a visitor can see and verify. Structured data reduces ambiguity; it does not create authority, guarantee rich results or force an AI platform to cite the page. Validate syntax and eligibility with the appropriate testing tools after every template change.
Strengthen entity and source signals
A search system needs to resolve who published the information and why that source is credible. Maintain a complete About page, consistent business identity, physical or service-area details, clear contact information, author attribution and links to authoritative external profiles.
Case studies should identify the problem, work completed, location, constraints and outcome. Service pages should explain the provider's process and qualifications. Articles should have a named author or accountable organisation and link to primary evidence when making technical, legal, medical or statistical claims.
Consistency matters across the website, Google Business Profile, major directories and industry references. These sources help systems distinguish the business from similarly named entities and associate it with the correct services and locations.
Optimise images and non-HTML resources
Use descriptive filenames, accurate alternative text, explicit image dimensions and modern compressed formats. Important images should have stable crawlable URLs and be placed near relevant explanatory text. Do not embed essential facts only inside an image because text extraction and accessibility will be weaker.
Create dedicated social and article images with absolute URLs in Open Graph, Twitter and Article structured data. If PDFs contain valuable standalone information, make sure they return the correct content type, have crawlable links and use X-Robots-Tag deliberately. Where possible, also publish critical PDF information as accessible HTML.
Maintain performance, mobile usability and security
Fast server response and efficient rendering allow crawlers to retrieve more content reliably and give visitors a better experience. Optimise Core Web Vitals, compress media, cache static assets, reduce unused JavaScript and prevent layout shifts. Test templates on mobile because mobile-first crawling is the standard for major search engines.
Use HTTPS everywhere and redirect HTTP requests to the canonical HTTPS host. Keep certificates valid, avoid mixed content and ensure security headers do not prevent legitimate rendering. Performance and security are not replacements for relevant content, but technical instability can make otherwise useful content unavailable.
Do not rely on experimental shortcut files
Files proposed specifically for language models may be useful for experiments, but they are not substitutes for crawlable HTML, robots.txt, XML sitemaps, canonical URLs or structured data. Do not assume that an llms.txt file, an AI-only summary or hidden machine-facing copy will improve visibility across major search platforms.
Maintain one accurate public source of truth. Content shown to crawlers should be consistent with what visitors receive. If an experimental discovery file is added, treat it as supplementary and monitor whether any verified platform actually requests or uses it.
Monitor crawler access and index coverage
Technical optimisation is an ongoing operational task. Use Google Search Console and Bing Webmaster Tools to monitor sitemap processing, indexing, canonical selection, crawl errors and important query visibility. Review server, CDN and WAF logs for verified search and AI crawler activity.
Run automated checks against representative public and private URLs. Test robots.txt interpretation, HTTP status, redirect destination, canonical value, robots metadata, rendered HTML, structured data and sitemap membership. Include at least the homepage, a service page, an article, a case study, an image and a deliberately non-indexed page.
When a crawler is missing, determine whether the failure occurs at robots.txt, DNS, TLS, CDN, WAF, origin response, rendering or indexing. Fix the earliest failing layer rather than adding more metadata to a page the crawler cannot retrieve.
Organic AI visibility and ChatGPT advertising are separate channels
The technical work in this guide is designed to make public website content eligible for organic discovery, retrieval and citation. It does not purchase placement and cannot guarantee that ChatGPT or another AI platform will select the website for a particular answer.
ChatGPT advertising is a separate paid channel. It can give a business another way to appear while potential customers are researching services, comparing options and deciding what to do next. Businesses considering that channel should treat crawler access, landing-page quality, conversion tracking and campaign management as part of the advertising setup rather than as organic ranking signals.
Explore ChatGPT Ads managementTechnical eligibility is the foundation, not the finish line
A technically accessible website gives Google and AI search platforms permission and a reliable method to retrieve public content. Clear architecture, consistent entities, structured data and extractable answers then make that content easier to interpret.
The final selection still depends on whether the page is relevant, current, trustworthy and useful for the question being asked. The strongest strategy is to combine clean technical delivery with original service knowledge, real evidence and pages that resolve a customer's decision more completely than competing sources.

