United States Methodology

Methodology

How the dataset is built — from scattered public records to one clean, standardized file.

Sources

The dataset is aggregated from publicly available U.S. property records and listings. We combine authoritative public-record fields (parcel, characteristics, sale history) with current market listing status, then reconcile them into a single record per property.

Standardization

Raw property data is famously inconsistent — every jurisdiction names and formats things differently. We normalize everything into one schema: a canonical address, a fixed set of property-type categories, consistent units (square feet, dollars), and a single status field (active, sold, off-market). See the full column list in the data dictionary.

Deduplication & identity

Each property is resolved to one stable record and assigned a durable public id (stp_…). That id is also a live API endpoint — GET /property/{id} on Straply — so any row in a downloaded file can be looked up for its current state.

Coverage & freshness

Coverage spans all 50 states plus DC and territories — roughly 146 million properties. The dataset is compiled monthly and released aged ~30–60 days behind live — each release is a point-in-time snapshot, stamped with its “as of” month on every page and in the download filename. Data decays: prices change, listings turn over. The free file is a photograph, taken last month; the API is the live video. Need it fresher than the monthly file? That's exactly what the API is for.

What we exclude

We publish facts about properties, not people. No owner names, no contact information, no skip-trace or people-search fields. Just the physical and transactional facts of each home.

Want the always-fresh version?
See the API →

RELATED · MORE

See all more →