For data platforms, enrichment tools, and dataset builders
License the subprocessor disclosure corpus.
Evidence-backed vendor-stack data on 3,800+ entities — growing with every scan. Deterministic extraction, with the source URL and quoted disclosure row on every record, ready to resell with provenance intact.
Why this data
First-party signal your customers can audit.
Highest-precision public stack signal
Subprocessor disclosures are legally mandated, first-party statements — not JavaScript fingerprints or job-posting inference. When a record says an entity uses Twilio, the entity wrote that itself.
Provenance on every record
Each match carries the disclosure source URL and the verbatim quoted row it was extracted from. Your customers (and their auditors) can verify any claim in one click — resale-grade lineage.
Deterministic extraction
No LLM in the pipeline: the same page always parses to the same output, and no vendor is ever emitted that isn't on the page. Reprocessing is reproducible, diffs are meaningful.
13 categories, refresh timestamps
Vendor stacks are normalized into 13 categories against a curated slug taxonomy, and every entity carries scanned_at so you always know data age.
Access patterns
Pull it, license it, or subscribe to changes.
Bulk API
The same flat pricing at volume: 1 credit per lookup, 1 credit per reverse-lookup page. rpm_limit is raised per key for pipeline workloads, and unindexed domains live-scan automatically at no extra cost — point your backfill at us and the corpus grows with it.
Corpus licensing
Full or per-category dataset exports (CSV/Parquet) with provenance columns — domain, vendor slugs by category, evidence source + quote, scanned_at, confidence tier — refreshed on a schedule you set. Licensing terms sized to your redistribution model.
Change feeds
Webhook push when an entity is re-scanned, a new disclosure appears, or a stack changes — the delivery plumbing exists in the schema today and is being productized. Until it ships, scheduled re-pulls plus scanned_at diffing cover the same need.
What a record looks like
Every field ships with its receipts.
| Field | Contents |
|---|---|
| domain | Canonical entity domain — the join key |
| vendor_stack | Vendor slugs grouped into 13 categories (sms_messaging, email, payments, …) |
| evidence[] | Per match: disclosure source URL + the verbatim quoted row |
| subprocessor_urls | Where the entity publishes its disclosure(s) |
| scanned_at | Last successful scan timestamp — data age, per record |
| confidence | vendor_keyword (row names vendor + function) > vendor (name match) — evidence strength, not model certainty |
Precision framing: a vendor_keyword match means the disclosure row itself names the vendor and describes the category function — the strongest public evidence a vendor relationship can have. Records that couldn't be grounded in a quoted row simply don't exist; we never backfill from inference.
Tell us what your platform needs.
Category coverage, export cadence, redistribution terms — send the shape of it and we'll come back with a concrete proposal. Volume API pricing and contract terms live under Enterprise.
Contact us — pick "Partnerships & data licensing"