The Keigeo crawler.

If you found this address in your server logs, this page says what we were doing, how to block us, and how to reach a person.

What it is

Keigeo is a tool people use to check a website: whether search engines can read it, and whether AI assistants can quote it. Someone asked us for a report on a site, and this crawler fetched the public pages needed to answer that.

It sends this User-Agent, on every request, from both halves of our system:

Keigeo/1.0 (+https://keigeo.com/crawler)

What it requests

Public pages, the way a browser would. A free scan reads your home page and up to ten more. A paid deep report crawls up to 250 pages of the same site, one request at a time.

It does not log in, submit forms, or try to reach anything a visitor could not reach. It keeps the measurements taken from a page, not a copy of the page.

If your server answers with 429, it waits as long as your Retry-After header asks, up to ten seconds, slows to one request a second, and stops entirely after the third refusal.

How to block it

We read /robots.txt before fetching, and we obey it. To turn us away, add this:

User-agent: Keigeo
Disallow: /

A blanket User-agent: * with Disallow: / stops us too. Either takes effect on the next scan, with nothing to wait for and nobody to email.

How to stop it completely

Email hello@keigeo.com from an address at the domain, or tell us how else you can show the site is yours, and we will add it to a list the scanner checks before every run. After that we refuse the scan, including to anyone who asks us for a report on you.

We do this on request and we do not ask why. It takes effect within a minute of us adding it, not at some later release.

Who asked for the scan

A person who enters a domain must first confirm that they own the site, have its owner’s permission, or are researching it as a competitor for a lawful purpose, and accept our Terms. The report goes only to that person. Our Terms forbid them from publishing a report about a site they do not own without the owner’s written agreement.

If a report about your site is being used in a way you think is wrong, email us. We would rather hear about it.

Which sites it requests

We request pages on the domain that was entered, and its subdomains, and nothing else. Links from a scanned site to other sites are read from the page. We do not request them. If a scanned site redirects to another domain, the scan stops.