- User-Agent: mcp-risk-crawler-worker/1
- Snapshot User-Agent: MCPRiskBot/1.0 (+https://mcprisk.dev/crawler)
- Contact: crawler@mcprisk.dev
- Purpose: Collecting source-scoped MCP server identity, version and lifecycle evidence for later review.
Source intake and opt-out.
The Go v2 crawler collects source observations for the MCP Risk threat feed. It is run manually; no crawler schedule is enabled. An hourly Official Registry schedule, capped at 100 records per run, is defined but disabled. The separate importer validates and writes published artifacts to the database. A retained snapshot worker has a separate User-Agent and can make MCP metadata requests when its own job is invoked.
These are the four supported Go v2 source keys and their request hosts.
official-registry · registry.modelcontextprotocol.io · Manual runs; hourly schedule defined but disabled
- Reads: HTTPS GET of the public /v0.1/servers listing, including deleted records and cursor pages. Version claims come from listing records; detail URLs are not fetched.
- Cadence: When an operator starts a Go v2 run; included in the default selection. An hourly schedule of at most 100 records per run is defined but disabled.
github-registry · api.mcp.github.com · Available for manual runs
- Reads: HTTPS GET of the public /v0.1/servers listing, including deleted records and cursor pages. Version claims come from listing records; detail URLs are not fetched.
- Cadence: Only when an operator starts a Go v2 run; included in the default selection.
docker-catalog · desktop.docker.com · Available for manual runs
- Reads: HTTPS GET of /mcp/catalog/v3/catalog.json.
- Cadence: Only when an operator starts a Go v2 run; included in the default selection.
mcpservers-org · mcpservers.org · Manual opt-in only
- Reads: HTTPS GET of /sitemap.xml and its server-sitemap XML shards only. Listed server pages and /api/ are never fetched.
- Cadence: Only when an operator explicitly selects this source and acknowledges its access and licensing constraints.
- The default manual run selects the official registry, GitHub registry and Docker catalog, with at most 100 records per source. Operators can select fewer sources or change that bound.
- A run is capped at 50,000 records, 500 HTTP requests, 250 MiB downloaded and 15 minutes. Each response also has an adapter-specific size cap.
- Requests use HTTPS GET to explicitly allowed source origins. Redirects must stay on an allowed origin.
- The Go v2 crawler does not fetch robots.txt or enforce a per-host request rate. That applies to manual runs and to the schedule below. Please email us to request exclusion.
- Disabled by default. It is not enabled, and no other source has a schedule.
- If an operator enables it, it runs hourly against registry.modelcontextprotocol.io only and reads at most 100 records per run, continuing from a cursor stored in our database.
- Each scheduled run is capped at 20 HTTP requests, 32 MiB downloaded and 3 minutes, on top of the per-response size cap.
- Each run's artifact is stored privately and deleted after 90 days. Failed imports are retried a bounded number of times.
- The mcpservers.org sitemap contributes weak directory pointers only; it does not prove that a listed server belongs to any repository or package.
- The crawler records source claims and attribution in a v2 artifact. Import and identity resolution happen separately; a source claim is not a verified security finding.
- The intake crawler does not connect to listed MCP endpoints, execute tools or run code from a listed repository.
- The separate snapshot worker may contact a declared MCP endpoint for initialize, notifications/initialized, tools/list, prompts/list and resources/list. It never invokes a listed tool. Its User-Agent is shown above.
Email crawler@mcprisk.dev with the source or host to exclude. We will stop selecting it for future manual and scheduled runs and confirm. The current Go v2 crawler does not read robots.txt, so a robots.txt rule alone does not change its behavior.