Netstate’s business data methodology is a source-to-record process. We collect 128 million US business records from primary sources — 66 in all, spanning every US state, DC, Puerto Rico and Canada plus 16 federal datasets — normalize them into a common schema, resolve filings to the correct legal entity, and attribute and timestamp every field. Names, legal status, registration numbers, NAICS codes, registered agents, addresses, officers and directors, UCC liens, litigation, SBA loans and beneficial-owner links all trace back to the public record they came from. This page documents the collection pipeline, source coverage and cadence, entity resolution, data quality controls, the limits of public records, and how to correct what we publish.
01 Business data collection pipeline
Our collection methods are the same for every record: five steps from the issuing agency to the profile you read. Nothing skips a step, and each step is logged.
- Ingest from the primary sourceData collection starts at the agency that issues each filing — Secretary-of-State registries, federal regulators, courts and bulk data programs. We collect data via official bulk files, APIs and FOIA-released datasets: only publicly available data, no second-hand or resold aggregations.
- Normalize & validateRaw data is parsed into a common schema. Names, addresses, dates and identifiers are standardized; malformed or duplicate filings are flagged for data quality before they reach a profile.
- Resolve to a legal entityProbabilistic and rule-based matching links filings that belong to the same entity across states and agencies, keyed on registration numbers, EIN/TIN, officers and addresses. Low-confidence matches are held for review, not auto-merged.
- Attribute & timestampEvery field is tagged with its source dataset, the issuing agency, and the date it was retrieved. This is what powers the "source" and "verified" labels on each company profile.
- Refresh & reconcileSources are re-pulled on a fixed cadence (see catalog below). When a source changes a record, the profile updates and the prior value is retained as historical data in the entity's history.
02 Data sources & update cadence
Netstate draws on 66 primary sources: all 50 Secretary-of-State registries plus 16 federal datasets, covering both listed and private companies. The cadence column is each register’s own publishing schedule — we do not claim data is fresher than its source, and a very recent filing may not yet appear.
| # | Source | Source cadence |
|---|---|---|
| 00 | Secretary of State (×50) State business registries | Weekly |
| 01 | U.S. Treasury | Daily |
| 02 | Federal Licenses FCC · FAA · DEA | Weekly |
| 03 | Small Business Administration | Monthly |
| 04 | Dept. of Labor | Quarterly |
| 05 | USCIS | Monthly |
| 06 | Dept. of Transportation | Weekly |
| 07 | Securities & Exchange Comm. | Daily |
| 08 | Patent & Trademark Office | Weekly |
| 09 | Patent & Trademark Office | Weekly |
| 10 | Library of Congress | Monthly |
| 11 | CMS | Monthly |
| 12 | Federal Election Comm. | Weekly |
| 13 | CBP | Daily |
| 14 | Internal Revenue Service | Monthly |
| 15 | Dept. of Labor / EBSA | Quarterly |
| 16 | UCC & Court records State filing offices & courts | Weekly |
Each profile shows the specific retrieval date per field. A “verified” label means the field was reconfirmed against its source within the last refresh window.
03 Entity resolution
The hard part of public-records search is not the data collection — it is deciding which filings describe the same company. A single business, including small private companies, can appear under slightly different names across fifty states, with separate registration numbers, multiple registered agents and decades of historical data on address changes.
Netstate links records using a combination of deterministic keys (EIN/TIN, state filing numbers, DUNS) and probabilistic signals (normalized name similarity, shared officers, co-located addresses, filing timelines). Matches above our confidence threshold are merged into one profile; everything below it is queued for manual review rather than guessed. When a merge is later found to be wrong, the entities are split and the change is logged in the profile’s history.
04 Data quality & limitations
Data quality and management are explicit: we say what the data can and cannot tell you, and we do not publish accuracy percentages we have not audited.
What we commit to. Every field is traceable to its source dataset and retrieval date. Low-confidence matches are queued for review rather than auto-merged. Correction requests are re-verified against the issuing source before any change is published, and we aim to respond within three business days.
Known limitations. Source agencies publish on their own schedules, so a very recent filing may not yet appear. Some states redact officer or address details. Absence of a record (for example, “no bankruptcies on file”) means none was found in the searched sources — not a guarantee that none exists. Netstate reports what the public record says; it does not adjudicate or score the underlying facts.
Data management. Our business data collection is built on publicly available data only — no scraped social profiles or purchased consumer data. Every field carries its source and retrieval date, so the provenance behind our data quality is auditable end to end.