StoreCanaryBot

Public identity and operating policy for the StoreCanary Shopify SEO crawler.

Operator
StoreCanary, operated by LAMOTTE LLC
Bot name
StoreCanaryBot
User-Agent
Mozilla/5.0 (compatible; StoreCanaryBot/1.0; +https://storecanary.io/bot; hello@getstorecanary.com)
Signature-Agent
https://storecanary.io/.well-known/http-message-signatures-directory
Public keys
HTTP Message Signatures directory
Contact
hello@getstorecanary.com

Purpose

StoreCanaryBot audits public Shopify storefront pages for search visibility, structured-data, indexability, catalog, and product-feed problems. It supports reports requested by merchants, recurring monitoring for StoreCanary customers, and limited qualification scans that identify a public issue before contacting a merchant.

What the bot fetches

The bot may request the public homepage, robots.txt, XML sitemaps, public catalog JSON endpoints, collection pages, product pages, and a bounded byte range from a representative public product image. It does not log in, use Shopify Admin, access customer accounts, add products to carts, begin checkout, submit forms, or modify storefront data.

Crawl rate and error handling

Requests are paced separately for each hostname. StoreCanaryBot currently starts at one HTML request per second and immediately reduces its rate after HTTP 429 or 430 responses. It respects Retry-After, retries at most once, and stops a scan when the response pattern no longer supports a reliable report. Higher rates are used only after Shopify explicitly grants a signed-traffic tier and only while the observed 429/430 rate stays at or below the one-request-per-second reference.

robots.txt and crawl directives

Before catalog pages are crawled, StoreCanaryBot reads robots.txt. A group for StoreCanaryBot takes precedence over the wildcard group. The bot applies longest-match Allow and Disallow rules, honors Crawl-delay, and declines the scan when the file forbids a required product path or when the requested delay cannot fit within a reliable scan. A temporary 429, 430, or 5xx response from robots.txt also stops catalog crawling.

Data use and retention

Fetched content is processed on StoreCanary infrastructure to produce technical findings and supporting page URLs. Storefront HTML is not sold, used to train models, or republished as a content corpus. Reports contain derived audit findings, not copies of product-page content. StoreCanary does not rotate source addresses or disguise its User-Agent to bypass a merchant's controls.

Opt out or report a problem

A store operator can disallow StoreCanaryBot in robots.txt. For an immediate manual block, removal request, rate concern, or security report, email hello@getstorecanary.com or use the contact page. Include the store hostname so the request can be applied accurately.