GEO guides

Technical GEO checks for reading a website's HTML content

Inspect service-page content with curl, original HTML and browser rendering, then check crawl rules, search status and actual AI citations separately.

When working on GEO for a business website, start with a service page that explains an actual offering. A browser may show the company name, service description and contact information correctly, but you still need to check how that content reaches the browser. If JavaScript has to call an API before the main text appears, a program that downloads the document without executing scripts may receive only a loading indicator and script tags or references.

Google's JavaScript SEO documentation provides a concrete basis for investigating this. Google separates crawling, rendering and indexing, and can execute JavaScript. The documentation also notes that not every bot can run scripts. A useful engineering check therefore compares the HTML returned by the server with the page produced after a browser executes JavaScript. Google's JavaScript SEO documentation

Save the final response and its content

Choose a publicly accessible service page and use its complete public URL. A homepage usually introduces the whole company; a service page is a better place to check whether a system can retrieve the details of a specific offering.

The following example is for Windows PowerShell. Replace the example URL with your own page and run the commands in a new diagnostic directory to avoid overwriting existing files with the same names.

curl.exe -sS -L --max-redirs 5 --connect-timeout 10 --max-time 30 --compressed `
  -D geo-headers.txt -o geo-page.html `
  -w "final=%{url_effective} status=%{http_code} bytes=%{size_download}\n" `
  "https://www.example.com/services/ai-visibility/"

if ($LASTEXITCODE -ne 0) {
  throw "The request did not complete; check the curl error first"
}

Get-Content -LiteralPath .\geo-headers.txt
Select-String -Path .\geo-page.html -Pattern '<h1', 'Actual company name', 'Actual service name'

Calling curl.exe explicitly avoids differences between commands with the same name in PowerShell environments. -L follows redirects, -D saves response headers and -o saves the response body. The url_effective value is the last URL fetched. The header file may contain several stages of a redirect chain, so read it alongside the final URL and the last set of headers. curl's option reference

First check whether the request completed, then inspect the final HTTP status. Even with a 200 response, open the HTML file and confirm that it contains the intended service page. Login and verification pages can also return 200. The byte count can flag an unusual response, but cannot establish that the content is complete. A HEAD request alone cannot inspect the response body.

Searching for a company or service name helps locate clues, but a matching string is not enough. It may appear only in script data or JSON-LD. Inspect the surrounding markup and confirm that the actual page content includes the service description, intended customers, use cases and contact route. Plain text searches may also miss names encoded as HTML character references; inspect the source or use an HTML parser in that situation.

Compare the original HTML with the rendered page

Open the same URL in a browser. In the developer tools' Network panel, find the document request and inspect its Response. Then use the Elements panel to inspect the DOM after scripts have run. These views show what the server initially sent and what the browser subsequently added.

Chrome offers a useful supplementary check. Keep developer tools open, press Ctrl + Shift + P, run Disable JavaScript, and reload the page. Run Enable JavaScript afterwards to restore it. This check shows how the page behaves without scripts in that browser; it does not simulate every AI system's crawler. Chrome's instructions for disabling JavaScript

Use a small record of what you observed to decide where to investigate next.

ObservationFirst place to investigate
The download fails, or the final response is 403 or 5xxCDN, firewall, origin response and redirects
The response is 200, but contains a verification or login pageAccess policies and edge protection rules
The original HTML contains a loading placeholder, while the rendered page has the business contentRendering strategy and client-side data fetching
The HTML already contains the full contentCrawl rules and the search platform's reported state

The first three observations can already guide a specific repair. If the text exists only in the rendered DOM, consider server rendering or static generation for the core business content, while keeping interactive controls on the client. Google also recommends considering server rendering or prerendering to reduce dependence on bots executing JavaScript. Google's guidance on rendering

The goal is for the initial HTML of a public page to contain its real business information. Service coverage, pricing or contact details that change should come from a reliable data source, so that static text and interactive components do not present conflicting information.

Check crawl permissions separately from indexing directives

Once you have retrieved the full HTML, check the site's robots.txt, the page's robots meta tag and any X-Robots-Tag response header. These controls address different questions. Google can discover a page's indexing directives only when it is allowed to crawl and actually retrieves the page. Blocking the URL in robots.txt may prevent it from reading a noindex directive on that page. Google's robots meta and response header rules

For ChatGPT search, identify the specific crawler. OpenAI provides independent controls for OAI-SearchBot, which is used for search, and GPTBot, which is used for model training. A site can allow the former while restricting the latter. ChatGPT-User, used for user-triggered visits, has a different purpose, and its records do not replace search crawler records. OpenAI's crawler documentation

If the business decides to allow ChatGPT search crawling, check whether the existing rules cover the service page, then check whether the CDN allows the official request sources. Preserve existing restrictions on administrative, private and other protected paths when changing the rules.

Changing curl's User-Agent to a bot name only tests whether the server returns different content for that string. Verify actual crawler sources against the platform's published IP information as well; a name in an access log does not establish the visitor's identity. OpenAI publishes an IP list for its search crawler. OAI-SearchBot IP list

Recheck the same business page after a repair

Continue using the original URL and save a fresh copy of its headers and HTML. Confirm that the request reaches the correct page, the main content appears in the initial HTML, and the browser does not show a different business description after executing scripts.

For Google's actual search state, use the URL Inspection tool in Search Console. A live test can help inspect the current page and its rendering; review indexing status separately against the indexed version. Google's JavaScript troubleshooting documentation

For Google AI Overviews and AI Mode, a page must be indexed and eligible to appear with a search snippet before it is eligible as a supporting link. Google adds no separate technical requirements for these features, and meeting the requirements does not guarantee display. Google's requirements for AI search features

When observing later AI answers, record the original question, platform, time and actual cited URL. A technical repair can make page content easier to retrieve; brand mentions and citations to that service page still need to be observed in subsequent answers. End the diagnostic record with a specific page and defect, so that a developer knows what to change in the next release.

To connect these technical checks to the page's business content and a repeatable review, see the guide to improving an existing page.

Key references

Sources checked on October 8, 2026.