StoreCanaryBot
Public identity and operating policy for the StoreCanary Shopify SEO crawler.
- Operator
- StoreCanary, operated by LAMOTTE LLC
- Bot name
- StoreCanaryBot
- User-Agent
Mozilla/5.0 (compatible; StoreCanaryBot/1.0; +https://storecanary.io/bot; hello@getstorecanary.com)- Signature-Agent
https://storecanary.io/.well-known/http-message-signatures-directory- Public keys
- HTTP Message Signatures directory
- Contact
- hello@getstorecanary.com
Purpose
StoreCanaryBot audits public Shopify storefront pages for search visibility, structured-data, indexability, catalog, and product-feed problems. It supports reports requested by merchants, recurring monitoring for StoreCanary customers, and limited qualification scans that identify a public issue before contacting a merchant.
What the bot fetches
The bot may request the public homepage, robots.txt, XML sitemaps, public catalog JSON endpoints, collection pages, product pages, and a bounded byte range from a representative public product image. It does not log in, use Shopify Admin, access customer accounts, add products to carts, begin checkout, submit forms, or modify storefront data.
Crawl rate and error handling
Requests are paced separately for each hostname. StoreCanaryBot currently starts at one HTML request per second and immediately reduces its rate after HTTP 429 or 430 responses. It respects Retry-After, retries at most once, and stops a scan when the response pattern no longer supports a reliable report. Higher rates are used only after Shopify explicitly grants a signed-traffic tier and only while the observed 429/430 rate stays at or below the one-request-per-second reference.
robots.txt and crawl directives
Before catalog pages are crawled, StoreCanaryBot reads robots.txt. A group for StoreCanaryBot takes precedence over the wildcard group. The bot applies longest-match Allow and Disallow rules, honors Crawl-delay, and declines the scan when the file forbids a required product path or when the requested delay cannot fit within a reliable scan. A temporary 429, 430, or 5xx response from robots.txt also stops catalog crawling.
Data use and retention
Fetched content is processed on StoreCanary infrastructure to produce technical findings and supporting page URLs. Storefront HTML is not sold, used to train models, or republished as a content corpus. Reports contain derived audit findings, not copies of product-page content. StoreCanary does not rotate source addresses or disguise its User-Agent to bypass a merchant's controls.
Opt out or report a problem
A store operator can disallow StoreCanaryBot in robots.txt. For an immediate manual block, removal request, rate concern, or security report, email hello@getstorecanary.com or use the contact page. Include the store hostname so the request can be applied accurately.