# Crawler Source Inventory and Legal Pre-review
AI Summary
Purpose:
- Preserve the 2026-08-20 source inventory and pre-review that grouped current
library, vulnerability, OS-package, and license collection paths by access and licensing risk.
Key points:
- The review separates three levels: low (public data or an OSS/data license),
medium (collection is possible only while published conditions such as rate limits, API keys, current-version requirements, or bot restrictions are followed), and high (terms restrict automated access or commercial reuse, so a written agreement or a different source is required).
- High-priority cases on the published one-page summary are Cloudera and
Atlassian Maven, GitLab gemnasium-db, and Amazon/Oracle OS-package data. The proposed routes are written agreement/official protocol, the MIT-licensed advisories-community, and vendor repodata or written permission.
- Medium examples include Maven Central, npm (published crawler limit: 1
request/second and 5M requests/month), NVD/GitHub API key or token limits, Swift Package Index bot restrictions, OSADL's current-version condition, and Photon terms that still need confirmation.
- The operating rules are to prefer public APIs/files/repos, set per-source
request intervals/concurrency/retry policy, identify the client with an appropriate User-Agent, obey robots and access controls, and never bypass a block.
- The customer project is framed as building a collection system that the
customer operates, not selling a copied dataset. That framing still depends on the customer owning the relevant terms acceptance/API keys and does not settle LabradorLabs' separate third-party redistribution question.
- Attribution requirements and possible maintainer names/emails in package
metadata need an explicit handling rule.
- The one-page summary excludes file/function/binary components, code-level
vulnerability sources, exploit/PoC data, and Snyk. It is a scoped briefing, not the full internal crawler inventory.
- This is a technical pre-review, not a legal opinion. Final legal approval and
the disposition of restricted sources remain Needs confirmation.
Relevant when:
- Changing a crawler source, adding a new upstream, writing a customer-facing
collection plan, or preparing portfolio material about source governance.
Do not read full document unless:
- The exact per-source URL, cadence, terms link, or unresolved source list is
needed; then open the Confluence origin.
Linked documents:
ai/wiki/projects/crawler-source-governance.mdai/repo-notes/labrador-scrapers.mdai/worklog/2026/2026-W34.md
Open Questions
- Whether written agreements will be obtained for every high-risk source.
- Whether every medium-risk source has an enforced runtime limit rather than a
documentation-only rule.
- Whether production redistribution requires supplier agreements beyond the
customer-operated project scope.
- How attribution and maintainer personal data will be represented downstream.
Details
- Captured from Confluence DT/4211146807 and its one-page child summary
DT/4212457542 on 2026-08-20.
- Page comments contained private discussion and personal identifiers; they
were used only to understand scope and were not copied into this source note.