Data Sources

No single vendor owns our truth.

We acquire facts through official APIs first, licensed feeds second, and governed web collection third — through providers that stay replaceable. What we own permanently is our warehouse, knowledge graph, provenance ledger, and scoring.

Our focus right now

Building on Apify

Apify's marketplace of reusable Actors, scheduling, datasets, and APIs makes it the fastest place to prove a source: we can stand up collection for a new buying trigger in days, not months. That speed is why the Local Opportunity Radar pilot runs on it today.

Why it fits now

Reusable scrapers, browser automation, structured datasets, and API delivery in one platform.

Where it ends

Apify accelerates us but never becomes the system of record — our warehouse, history, and scores do.

Actor and dataset documentation: apify.com → docs.apify.com/actors · docs.apify.com/storage/dataset

Sources, in order of preference

We work down this list and stop at the highest tier that answers the question.

Tier 1

Authoritative APIs and open data

Government registries, regulatory filings, corporate disclosures, official market feeds. Highest authority, predictable schemas, easiest provenance — always our first choice.

Tier 2

Licensed datasets and feeds

Commercial feeds where data is essential, hard to collect, or legally sensitive — with contracts that permit storage, derived scores, and multilingual delivery.

Tier 3

Customer-authorized data

Your CRM, inventory, or private documents — logically separated by tenant, never used for shared training without explicit permission.

Tier 4

Permitted public-web collection

Targeted, documented, proportionate. Every source reviewed for terms, access signals, and personal-data presence before activation.

Tier 5

Derived and synthetic data

Wherever the question allows it, we keep scores and aggregates — not dossiers. Company-level opportunity scores instead of employee profiles.

Where we go as sources scale

Each provider below is a future adapter behind our standard observation format — never a new system of record.

Bright Data

Enterprise-scale collection, proxies, managed acquisition

Fallback when a source outgrows prototyping in reliability or volume.

Oxylabs

Proxies, unblocking, headless browsing across international sources

Resilient alternative for difficult, dynamically rendered pages.

Zyte

Managed extraction and developer-friendly collection APIs

Structured extraction pipelines as source count grows.

Firecrawl

Websites converted to clean, AI-ready structured content

Source-document preparation for retrieval and knowledge bases.

Diffbot

Automatic entity, product, and relationship extraction

Enrichment input and external knowledge-graph cross-checks.

Browserbase

Cloud browsers for agents that complete permitted workflows

Later execution layer — moving from recommending an action to completing it.

ScrapingBee · ScraperAPI

Simple API page retrieval with proxy rotation

Commodity adapters and fallbacks for straightforward sources.

Browse AI · Octoparse

No-code extraction and monitoring

Fast business-user prototypes before engineering a source properly.

Crawlee · Crawl4AI

Self-hosted, code-controlled crawling

Long-term control for sources that prove stable and valuable.

Advantage goes to the most trusted supply chain — not the biggest crawl.

Source registry, provenance on every signal, suppression honored everywhere. Read our privacy position.

Back to the universe