Boroughly › Data & sources

Data & sources

What Boroughly's numbers actually are, where each one comes from, and which are still estimates. Updated whenever the data pipeline runs.

Read this before relying on a figure. Prices are median sale prices from HM Land Registry records — real transactions. Scores are still indicative composites. An area median is not a valuation of any particular home: two streets five minutes apart can differ by hundreds of thousands of pounds.

Status by dataset

CityDatasetStatusWhat that meansSource
LondonpricesSourcedMedian sale price per postcode district from HM Land Registry Price Paid Data, covering 278 of 278 areas. Areas with fewer than 30 recorded sales are left as estimates and flagged individually.HM Land Registry Price Paid Data
LondoncrimeEstimatedSafety scores are composite judgements, not derived from police.uk recorded crime counts.
LondonschoolsEstimatedSchool scores are composite judgements, not derived from DfE performance tables.
LondontransportSourcedLondon journey times are fetched live from the TfL Journey Planner API at request time.Transport for London Unified API
ManchesterpricesNo dataNo data. Area names and coordinates only.
ManchestercrimeNo dataNo data.
ManchesterschoolsNo dataNo data.
ManchestertransportNo dataNo data.
BirminghampricesNo dataNo data. Area names and coordinates only.
BirminghamcrimeNo dataNo data.
BirminghamschoolsNo dataNo data.
BirminghamtransportNo dataNo data.

What the statuses mean

Sourced

Derived directly from a named primary source by an automated pipeline, with the fetch date recorded. Reproducible — run the script and you get the same number.

Estimated

An indicative figure entered into the dataset rather than derived from records. Directionally useful for comparing areas, but not a measurement.

No data

The area exists in the system with a name and location only. Nothing is published for it.

Where the sourced data comes from

As each dataset moves from estimated to sourced, it is pulled from these, all of which are free and openly licensed:

How the scores are built

Where a score is sourced, it's a decile rank across every area we cover: a 10 for safety means the area sits in the lowest-crime tenth of our dataset, not that it is objectively safe. We publish the transformation rather than a judgement, so you can disagree with the weighting and still trust the underlying number.

What the data can't tell you

Averages hide enormous street-level variation — two roads five minutes apart can differ by hundreds of thousands of pounds. Crime statistics reflect what was reported, not what happened. School catchments move annually. Community sentiment scores are opinion, collected imperfectly. Every one of these is a reason to visit an area at different times of day before deciding, and none of them is a reason to skip that step.

Corrections

If a figure looks wrong, it may well be. Tell us and we'll check it against the source and publish a correction.