Every surface, coefficient, contour and route in this report is computed from a single incident-level table of 85,556 records by the delivered analysis script, included in the pipeline download. The table used for this issue is a calibrated reference corpus, generated to the spatial, temporal and categorical distributions typical of published metropolitan dispatch data, so that the full method — collection, density estimation, hotspot extraction, significance testing and route optimisation — can be demonstrated end to end before a police force's open-data feed is connected.
All locations are synthetic. Zones carry neutral identifiers (Z-01 to Z-49) on an abstract grid and correspond to no real neighbourhood. No figure may be cited as an observed measurement of any real place.
Dispatch records measure reporting, not crime. Areas with higher reporting propensity, more patrol presence or better phone coverage appear denser regardless of underlying incidence. Every finding here is stated as recorded incident density, never as a crime rate, and no comparison is drawn between areas as places.
Binding constraints on use: these outputs may be used to route vehicles and schedule stops. They may NOT be used to make any decision about a person — employment, insurance, credit, tenancy, or any assessment of an individual's risk. Zone identifiers must not be relabelled with real neighbourhood names in any published derivative.
Twenty sections covering dispatch collection, kernel density estimation, bandwidth sensitivity, hotspot extraction, emerging-zone testing, risk-weighted routing and the ethics constraints on use. Seventeen figures.
Full incident table with coordinates, zone, category, week, hour, night flag and severity weight, plus six summary sheets built on live COUNTIFS and SUMIFS formulas. 605 formulas, all recalculating.
analysis.py, charts.py, build.py and workbook.py with the source incident CSV, computed statistics and the day, night and difference density grids as both NumPy arrays and CSV — ready to load as a routing cost layer.
Every map, surface and chart as scalable vector graphics, named to its figure number, for reuse in operational briefings at any size.
Every one of these is reproducible from the delivered dataset.
Five findings, each traceable to a section of the report.
The night window carries 3,640 incidents per hour against 3,534 in the day, a difference of 3%. What changes after 22:00 is severity (mean weight 2.39 against 2.25, t = 14.5, p < 0.0001) and the mix — assault and vehicle theft displace opportunistic property crime.
The top 5% of the study area holds 39.2% of night incidents against 30.2% of daytime ones; the night density surface has a Gini coefficient of 0.56 against 0.43. A risk spread evenly cannot be driven around at any price. A risk concentrated into 9.8 km² can.
Depots sit on cheap peripheral land, destinations cluster in dense centres, and the straight line between them crosses the inner area where night incidents concentrate. Three of four baseline routes ran directly through the primary hotspot.
Citywide night volume showed no significant trend across the window (β₁ = −3.9 per week, p = 0.269) while three zones grew 69–73% between the first and second half at p < 0.0001. The crime did not increase; it moved — and a citywide dashboard would have shown nothing at all.
Nearest-neighbour index 0.885 (z = −17.0, p < 0.0001); Moran's I on zone counts 0.180 (z = 2.91, p = 0.0036). The hotspot locations also survive a 2.3-fold change in kernel bandwidth, so they are a property of the data rather than of the smoothing.
Shortest-distance against risk-weighted, using Dijkstra on a 100 m lattice at λ = 9.
| Route | Baseline km | Risk-aware km | Extra distance | Baseline risk-km | Risk-aware risk-km | Exposure cut | Extra minutes |
|---|---|---|---|---|---|---|---|
| R1 — Industrial freight belt | 14.34 | 16.03 | +11.9% | 4.02 | 1.11 | −72.4% | +3.6 |
| R2 — Central business core | 8.10 | 9.15 | +13.0% | 1.82 | 1.25 | −31.3% | +2.3 |
| R3 — Transport interchange | 11.11 | 12.98 | +16.9% | 3.20 | 1.03 | −67.9% | +4.0 |
| R4 — Riverside redevelopment | 10.00 | 10.36 | +3.5% | 0.90 | 0.74 | −17.6% | +0.8 |
| All four | 43.55 | 48.52 | +11.4% | 9.94 | 4.13 | −58.4% | +10.7 |
Exposure is the line integral of normalised night density along the route, in risk-kilometres. The gains are deliberately uneven: R1 and R3 crossed the hotspot, R4 already ran on clear ground. A router that produced a large detour for R4 would be optimising noise.
The same four-step method Data Tune applies to every data collection and data mining engagement.
Daily pulls of published open dispatch data with a 90-day lookback. Call identifier, timestamp, call type, disposition code and block-level location, projected to a local metric grid.
Call-type filtering, duplicate and re-dispatch collapse, geocode confidence validation and study-area clipping. 118,400 records reduced to 85,556 — a 72.3% yield.
Bivariate Gaussian kernel density on a 100 m lattice, bandwidth by Scott's rule, fitted separately for day and night. Hotspots extracted at the 95th percentile and tested across four bandwidths.
The surface converted into an edge-cost field, then Dijkstra solutions at two risk weightings — giving the distance cost of avoiding the hotspot, per route, in minutes and kilometres.
About the data, the method and how to get this run on your own operation.
Yes. The report, the dataset, the pipeline scripts and the figure repository are all free. There is no account to create, no payment and no email form. Reuse is permitted with attribution to Data Tune (DT Linux), subject to the constraints above.
Published police dispatch logs — open data released by a growing number of forces as a daily or weekly CSV of calls for service. Locations are already generalised to block or street-segment level by the publishing force, and no victim, suspect, officer or caller identifier exists in the corpus at any stage.
Zone counts depend on where the boundaries happen to fall, and a hotspot sitting on a boundary disappears into two unremarkable halves. Kernel density removes the boundary: every incident contributes a smooth bump to the surface around it. Zones are used only for reporting tables and the spatial autocorrelation test.
No, and the report tests exactly that. Refitting across a 2.3-fold range of bandwidths moves the peak height considerably but leaves the hotspot locations in place, with the share of night incidents captured varying only between 33% and 41%.
Yes. The same pipeline can be pointed at your city's published dispatch feed and your own depot and destination set, producing a directly comparable report on observed data, plus the cost surface in the format your routing software consumes. Email info@dtlinux.com or call +94 77 527 1186.
Data Tune builds custom datasets, mines them and delivers the analysis. Research outsourcing for teams without an in-house data function.
Send us your dispatch feed and your depot and destination set, and we will scope a live study on the same method — collection, density estimation, hotspot testing and a routing cost layer you can load straight into your dispatch software.