GeoInsightsSite Selection
Riyadh · M3 alpha
RiyadhTier 1Weighted MCDAMeasured 2026-08-31 03:26 UTC

How the Riyadh data is built

Written for the analyst who has to defend a site decision to someone else. Every figure on this page is measured from the live baseline at the moment you loaded it.

Hexagons measured
42,682
resolution 9 — about a city block each
Named sources
12
3 carry share-alike terms
Factors you can weight
14
9 are withheld — see §6
Population vs. census
+0.35%
governorate scope only
Section 1

Where the data comes from

Every source below is named, licensed, and credited. This list is generated from the same registry the pipeline stamps onto every row of data, so it cannot describe a source differently from the way the data records it. Retrieval dates are read from the ingest ledger — a date recorded here is a record of a run, not a claim about one.

SourceUsed forLicenceRetrieved
Foursquare OS Places
dt=2025-02-06
© Foursquare Labs, Inc. Foursquare OS Places, licensed under the Apache License 2.0.
Competition tier — POI inventory (DATA_STRATEGY §3).Apache-2.02026-08-15 04:33 UTC
OpenStreetMap
Overpass API live query; Geofabrik gcc-states-latest.osm.pbf
© OpenStreetMap contributors, ODbL 1.0
Competition tier — POIs; accessibility tier — road class mix, junction density, traffic exposure and transit stops (SS-2-04). The ODbL anchor for §9 trap #2.ODbL-1.0
share-alike
2026-08-15 04:24 UTC
Overture Maps Foundation — places
release 2026-07-22.0, theme=places
© Overture Maps Foundation
Competition tier — POIs.CDLA-Permissive-2.02026-08-15 04:30 UTC
Overture Maps Foundation — buildings
release 2026-07-22.0, theme=buildings
© Overture Maps Foundation; contains data © OpenStreetMap contributors (ODbL)
Supply tier — building_count and footprint_area_sqm.ODbL-1.0
share-alike
2026-08-15 04:36 UTC
WorldPop
Global 2000-2020, SAU, 2020, 100 m (unconstrained)
WorldPop (www.worldpop.org), School of Geography and Environmental Science, University of Southampton — Global High Resolution Population Denominators Project. Licensed CC BY 4.0.
Demand tier — age/sex structure.CC-BY-4.02026-08-15 05:44 UTC
Kontur Population
kontur_population_SA_20231101
Kontur Population Dataset, © Kontur (kontur.io), CC BY 4.0
Demand tier — H3 population, input to dasymetric redistribution.CC-BY-4.02026-08-15 06:21 UTC
GHSL — Global Human Settlement Layer
R2023A, GHS-BUILT-S, epoch 2020, 100 m, Mollweide (ESRI:54009)
European Commission, Joint Research Centre (JRC): Global Human Settlement Layer GHS-BUILT-S and GHS-POP, release R2023A (epoch 2020, 100 m, Mollweide)
Supply tier — built-up surface, the dasymetric weighting surface.EC-JRC-Free2026-08-15 04:35 UTC
geoBoundaries — SAU ADM2 (administrative boundaries)
gbOpen SAU ADM2 @ 9469f09
Boundaries © geoBoundaries (gbOpen, SAU ADM2), derived from OpenStreetMap contributors — © OpenStreetMap contributors, ODbL 1.0
Base geography — governorate polygons carrying the census join.ODbL-1.0
share-alike
2026-08-15 06:22 UTC
GASTAT — Saudi General Authority for Statistics
Saudi Census 2022 (via RCRC open-data portal)
General Authority for Statistics (GASTAT), Saudi Census 2022; republished by the Royal Commission for Riyadh City open-data portal under the KSA Open Data Licence
Demand tier — governorate population and household counts.Saudi-Open-Data2026-08-15 06:22 UTC
SREM — Saudi Real Estate Market (Ministry of Justice)
Dashboard/GetAreaInfo transactions extract
Ministry of Justice — Saudi Real Estate Market (البورصة العقارية), Saudi Open Data Licence
Affluence tier — real-estate transaction values.Saudi-Open-Data2026-08-15 06:02 UTC
KSA district boundaries (user-supplied, provenance unverified)
user-supplied GeoJSON, 3,732 features, no version identifier
District boundaries: source unverified. Supplied without attribution or licence; provenance not established. Not cleared for redistribution or client-facing display.
Base geography — district polygons carrying the SREM land-value join. NEEDS-REVIEW: now composed into land_value_sqm (SS-1-14) and, through it, into the modelled affluence_index (SS-1-08), with the licence question still open.unknown2026-08-16 11:28 UTC
H3 — Uber hierarchical hexagonal grid (h3-pg)
h3 4.5.0 / h3_postgis 4.5.0
H3 © Uber Technologies, Apache License 2.0
Base geography — the H3 res-8/res-9 cell grid itself.Apache-2.02026-08-16 21:49 UTC

3 of these carry share-alike terms: publishing a derivative database built on them carries an obligation to license the result the same way. Contributions from those sources are tagged at row level so they stay identifiable and separable.

Section 2

Does the population add up

Three independent gridded population products and one census, over the same governorate. They disagree — they always do, because each models the same people from different inputs — and the size of that disagreement is the honest error bar on any population figure in this product. The baseline is built on the one that lands closest to the census.

EstimatePopulationvs. census 
Kontur Population Dataset (H3-native, 2023-11-01 release)
the baseline is built on this one
7,033,492+0.35%
GHSL GHS-POP R2023A (100 m raster, epoch 2020)
8,216,348+17.22%
WorldPop 2020 unconstrained age/sex (SAU)
6,437,777-8.15%
GASTAT 2022 census — Ar Riyadh governorate
reference
7,009,120

Redistribution neither creates nor destroys people

Redistribution moves people around inside a source zone; it must never create or destroy any. This compares what the baseline holds, at each resolution, against what the population source published.

  • Resolution 8 — baseline holds 7,033,492 against the source’s 7,033,492 (0.0e+0% difference)
  • Resolution 9 — baseline holds 7,033,492 against the source’s 7,033,492 (3.8e-8% difference)

⚠️ The census reference is a governorate total: it tests how many people the model has, not whether it has them in the right places. The district-level test that would answer the second question is blocked — see the validation section.

Section 3

Where the people actually are

The population source publishes on the larger resolution-8 grid, which spreads residents evenly across land nobody lives on. Each cell's total is kept — it is the best local measurement available — but where inside the cell those people sit is decided from building evidence: each building is counted, and credited with at most a residential-sized floorplate. That cap is what stops one airport terminal or warehouse outweighing a street of houses, and it is measured from this region's own classified residential buildings rather than assumed.

Floorplate cap
559.3 m²
pooled mean footprint of Overture-classified residential buildings in this region at res 9
Measured from
59,626
classified residential buildings in this region — not a literature value
Split on building evidence
99.9%
of the region's population; the rest falls back to built-up surface or an even split

Worked example — King Khalid International Airport

Nobody lives on an airport apron. A gridded population product that reports residents here is showing you its disaggregation error, and every trade area that overlaps the cell inherits it.

The source grid puts 6,279 residents in the single large cell covering it. That total is kept. Inside it, the people are placed on the 7 smaller cells according to what is actually built there — and 1 of them, with no mapped building, comes out at exactly zero.

Sub-cellBuildingsFootprint (m²)Residents placed
618461236018020351392,5541,712.5
618461236016971775233,9671,141.6
618461236017758207247,1941,141.6
618461236017233919225,5781,141.6
618461236018282495122,133570.8
6184612360174960631785570.8
618461236018544639000
Total6,279

Note the third column against the fourth: the cell with a 92,554 m² terminal roof does not get proportionally more people than one with a small building, because each building is credited with at most a residential-sized floorplate. That cap is the whole mechanism.

The larger source cell keeps its total. Inside it, the people move onto the sub-cells that carry building evidence, and a sub-cell with no mapped building goes to exactly zero — which is a result, not missing data. Read the resolution-9 row for placement; the resolution-8 row is still the source's own figure, smear included.

The limit of this method. Redistribution works strictly inside a source zone. It never moves people between resolution-8 cells, because doing so would discard the only real population measurement available and substitute our building model for it. The resolution-8 layer therefore still carries the source's own smear, by design — which is why every trade area reads resolution 9.

Section 4

What can be scored on

Every column the scoring engine can see, and the state it is actually in. Only the columns marked usable can carry weight in a score: a column with no value would otherwise be scored as a zero, and a column with no spatial signal would rank locations on a constant.

A blank is not a zero

A blank is not a zero. Where no source covers a cell the value is left empty, and every rollup in the product is written to know the difference: a zero would state that nobody lives there, which is a measurement, and nothing downstream could tell it apart from one.

In this region that applies to 9,272 of 42,682 cells for population. Kontur publishes only populated hexagons. No Kontur cell covers this grid cell, so nothing is known here. SS-1-02 deliberately declined to write a 0 for such cells because 'Kontur reported nothing here' is information a 0 would destroy; this composition honours that. Aggregations must COALESCE or filter explicitly.

In a score, a factor with no value for a location is excluded and the remaining weights are renormalised, so scores stay comparable — and the score reports what share of its intended weight it was actually computed from.

VariableStateCoverageOrigin
population
Residents per hex, Kontur redistributed onto building footprints.
usable
modeled
78.3% of cellsKontur Population Dataset (SA extract, release 20231101)
derived
affluence_index
MODELLED affluence, 0–100 (DATA_STRATEGY §2.3). Land value + POI affluence mix + plot spaciousness + dwelling scale. NOT income and NOT a measurement — there is no open income dataset for KSA at usable granularity.
usable
modeled
29.6% of cellssiteselect affluence model (modelled composite)
derived
poi_counts
{category: count} of open, countable POIs.
usable100.0% of cellsstaging.poi_unified (SS-1-05) + staging.poi_brand (SS-1-06)
ODbL-1.0
brand_presence
{brand_id: count} after Arabic/English brand normalization.
usable100.0% of cellsstaging.poi_unified (SS-1-05) + staging.poi_brand (SS-1-06)
ODbL-1.0
road_class_mix
{osm highway class: metres in the hex}, from OSM ways clipped to the cell polygon.
usable100.0% of cellsOpenStreetMap (Geofabrik gcc-states extract)
ODbL-1.0
junction_density
Street intersections per km² — OSM nodes with ≥3 incident edges on the public street network, over the true cell area.
usable100.0% of cellsOpenStreetMap (Geofabrik gcc-states extract)
ODbL-1.0
traffic_exposure
MODELED. Σ(clipped length × class weight × carriageway factor), in primary-arterial-equivalent metres. See baseline.road_classes.
usable
modeled
100.0% of cellsOpenStreetMap (Geofabrik gcc-states extract)
ODbL-1.0
transit_stops
Public transport stop places whose point falls in the hex.
usable100.0% of cellsOpenStreetMap (Geofabrik gcc-states extract)
ODbL-1.0
footfall_value
MODELED, UNCALIBRATED. 0–100 footfall proxy index (DATA_STRATEGY §5 option A): category-weighted POI density, traffic exposure, junction density, built-up intensity and anchor proximity. Res 9 only — the index is normalised against the cell population it is computed over, so a res-8 value would not be comparable. Not a visitor count. See baseline.footfall_weights.
usable
modeled
100.0% of cellssiteselect footfall proxy (proxy)
ODbL-1.0
footfall_confidence
Weight of evidence behind footfall_value, not probability of correctness. Capped at 0.5 while footfall_source='proxy' so an uncalibrated model cannot be mistaken for a purchased panel.
usable
modeled
100.0% of cellssiteselect.compose_footfall
ODbL-1.0
land_value_sqm
SAR/m² land value from srem.moj.gov.sa — the median of a district's monthly transaction prices over a trailing 12-month window, inherited by every res-9 cell in that district. ⚠️ DISTRICT granularity stored per hex: within-district variation is absent, and 46.4% of res-9 cells are NULL because no SREM transaction in the window reached them. Uncovered cells are NOT imputed. See baseline.compose_land_value.
usable46.4% of cellsSREM — Saudi Real Estate Market (Ministry of Justice)
Saudi-Open-Data
built_up_pct
Percent of sampled GHS-BUILT-S pixel area that is built surface.
usable100.0% of cellsGHSL GHS-BUILT-S R2023A (E2020, 100 m)
EC-JRC-Free
building_count
Overture building footprints whose centroid falls in the hex.
usable100.0% of cellsOverture Maps — buildings
ODbL-1.0
footprint_area_sqm
Total building footprint area in the hex.
usable100.0% of cellsOverture Maps — buildings
ODbL-1.0
pop_age_0_14
Residents aged 0–14, population × WorldPop band share.
the composer flagged this column as spatially constant across the region: a real quantity, but identical in every location, so it cannot rank one against another
no spatial signal
modeled
78.3% of cellsWorldPop 2020 unconstrained age/sex structure (SAU) share × dasymetric population
CC-BY-4.0
pop_age_15_29
Residents aged 15–29, population × WorldPop band share.
the composer flagged this column as spatially constant across the region: a real quantity, but identical in every location, so it cannot rank one against another
no spatial signal
modeled
78.3% of cellsWorldPop 2020 unconstrained age/sex structure (SAU) share × dasymetric population
CC-BY-4.0
pop_age_30_44
Residents aged 30–44, population × WorldPop band share.
the composer flagged this column as spatially constant across the region: a real quantity, but identical in every location, so it cannot rank one against another
no spatial signal
modeled
78.3% of cellsWorldPop 2020 unconstrained age/sex structure (SAU) share × dasymetric population
CC-BY-4.0
pop_age_45_64
Residents aged 45–64, population × WorldPop band share.
the composer flagged this column as spatially constant across the region: a real quantity, but identical in every location, so it cannot rank one against another
no spatial signal
modeled
78.3% of cellsWorldPop 2020 unconstrained age/sex structure (SAU) share × dasymetric population
CC-BY-4.0
pop_age_65_plus
Residents aged 65+, population × WorldPop band share.
the composer flagged this column as spatially constant across the region: a real quantity, but identical in every location, so it cannot rank one against another
no spatial signal
modeled
78.3% of cellsWorldPop 2020 unconstrained age/sex structure (SAU) share × dasymetric population
CC-BY-4.0
households
Households per hex.
The only household source available is a GASTAT governorate total. Distributing it by a constant households-per-person ratio yields a column that is a scalar multiple of `population` — no spatial signal, no information, but it looks like a measurement and would be weighted as one. A NULL column is honest; a constant-ratio one is not.
unavailable
daytime_pop_index
Daytime vs residential population (DATA_STRATEGY §2.2).
DATA_STRATEGY §2.2 needs five inputs. SS-2-04 delivered two (transit_stops, traffic_exposure); retail GLA is unavailable (SS-1-15) and school/university capacity is not ingested.
unavailable
spend_potential
SAR/year spend potential = population × age propensity × affluence × HES basket.
Two of its four factors are still unavailable. The GASTAT Household Expenditure Survey category basket was not obtained, so there is no basket to multiply by; and the age propensity term would multiply through a spatially constant age structure (SS-1-12 — exactly one distinct 65+ share across all 55,230 hexes). What is left is population × affluence × a constant, published in SAR — a currency-denominated restatement of two columns that already exist, and the currency is what would make it read as a measurement.
unavailable
retail_gla_sqm
Gross leasable retail area.
Requires building height or storey count. 169 of 756,675 Overture buildings carry a height and 397 carry explicit levels — 0.06% of the stock. Extrapolating GLA from that is fiction, not sparsity (SS-1-15).
unavailable
footfall_source
'proxy' | 'telco' | 'card' | 'sdk'. Which of DATA_STRATEGY §5's options produced footfall_value; 'proxy' until a panel is purchased.
this column records where a value came from rather than measuring the place; it is not a factor
provenance label100.0% of cellssiteselect.compose_footfall
ODbL-1.0
Section 5

What the score is

Four models sit on a ladder, and each rung needs data the one below it does not. The platform never hides which rung produced a number.

A Tier 1 score is an index, not a forecast. It ranks locations against each other on the factors you weighted; it does not predict revenue and is never denominated in SAR. Revenue forecasting is Tier 4 and requires your own stores' sales history.

TierModelNeedsGives youStatus
1Weighted MCDAnothing beyond the regional baseline0–100 suitability indexrunning now
2Huff / gravity — distance-decay market sharea competitor set and an attractiveness proxy% share captured per competing siteunlocks on your data
3Analog — nearest neighbour to top-performing storesroughly 20–30 of the tenant's own stores“this site resembles your Store 12 and Store 27”unlocks on your data
4ML regression on trade-area featuresroughly 50+ stores with sales historyrevenue forecast with a confidence bandunlocks on your data

What the running engine reports

Output
suitability index, 0-100
Is it a forecast?
no
Currency units
none — it is an index
Factors available
16
Weights calibrated?
no — PRD §4.4 placeholder profiles — direction and rough ranking per vertical, not fitted to observed store performance

This is a relative suitability index, not a revenue forecast. It has no monetary interpretation and must never be labelled in SAR. Revenue estimation requires Tier 2 (Huff) for share and Tier 4 (ML) for value, neither of which is in this number.

What the scoring API refuses to do, and why

Section 6

What this does not do

Read this section before you rely on anything above it. Each statement below is one a careful analyst would find on their own within an afternoon; each is stated here first, with the measurement behind it, because a limitation discovered after signing costs more than every claim on this page earns.

Footfall is a modeled proxy, not a measurement

read before relying on it

The footfall figure is an index this platform computes from POI density, traffic exposure, junction density, built-up surface and proximity to anchors. It is not a count of people, it does not come from mobile-device data, and it has not been calibrated against any observed visit counts. It is useful for comparing one location with another and must not be read as a visitor volume.

footfall_valuefootfall_confidence
unit
relative index, 0-100, region-normalised (NOT a visitor count)
method
weighted sum of five normalised terms, published on 0-100. poi_density, traffic_exposure and junction_density are ln(1+x) divided by the region's q=0.999 of that transform and clipped to 1; built_up_pct is divided by 100; anchor_proximity is a weight-summed exp(-d/d0) to the nearest anchor of each type and is already in [0, 1]. The weights are renormalised per cell over the terms that cell has an input for, so a missing input does not act as a zero.
calibrated
no
source_labels
label: proxy · cells: 42,682
inputs
region_baseline_h3.poi_counts, region_baseline_h3.traffic_exposure, region_baseline_h3.junction_density, region_baseline_h3.built_up_pct, staging.poi_unified_open, staging.stg_osm_transit_stops, app.siteselect.baseline.footfall_weights

What would change this: calibration against a tenant's own transaction or door-count data, or a purchased mobility dataset

Affluence is a modelled index — it is not income

read before relying on it

There is no open income dataset for Saudi Arabia at any sub-governorate granularity, so nothing in this platform has observed a household income. The affluence figure is an index computed from real-estate transaction prices, the mix of venues in the surrounding neighbourhood, and building typology. It orders neighbourhoods within one city; it is not a currency, not a percentile of income, and not comparable to another city's index. Its largest single input is a district-level land price, so it is coarser than the hexagon it is stored on.

affluence_index
unit
relative index, 0-100, region-normalised. NOT a currency, NOT a percentile of income, NOT comparable to another region's index until both are rescaled together.
method
weighted sum of four normalised terms, renormalised per cell over the terms that cell has an input for, published on 0-100. land_value is ln(1+SAR/m²) min-max scaled between the region's q0.02 and q0.98 of that transform; poi_affluence_mix is 0.5 + 3·(affluent − deprivation)/all open countable venues in an h3_grid_disk of radius k=2, clipped to [0,1]; plot_spaciousness is 1 − min(footprint coverage / 0.4, 1); dwelling_scale is a log-space trapezoid on mean footprint per building with corners [60.0, 220.0, 900.0, 2600.0] m².
calibrated
no
coverage_pct
29.6
term_weights
land_value: 0.45 · dwelling_scale: 0.1 · plot_spaciousness: 0.15 · poi_affluence_mix: 0.3
excluded_inputs
age_structure: WorldPop's unconstrained age/sex product has spatially constant shares in Riyadh — exactly one distinct 65+ share across all 55,230 hexes (SS-1-12). Nothing derived from it can vary in space. · nationality_mix: DATA_STRATEGY §2.3 lists it; it is available only as a governorate share. A district composition approximated from one region-wide number is region-constant, and a constant term multiplied through a weighted index reconciles perfectly while distinguishing nothing (the SS-1-12 failure mode). NOT approximated. · household_income: no open income dataset exists for Saudi Arabia at any sub-governorate granularity, and Meta's Relative Wealth Index does not cover KSA (DATA_STRATEGY §2.3). This absence is why the column is modelled. · population_density: available and deliberately unused. SS-1-03 redistributes Kontur by capped building footprint, so population is a transform of the same footprint data plot_spaciousness and dwelling_scale read; a fifth term on it would count one input three times. · building_height_or_storeys: 169 of 756,675 Overture buildings in Riyadh carry a height (SS-1-15). Apartment-vs-villa is exactly what height would settle and 0.02% coverage cannot settle it.
granularity_warning
the land_value term (45% of the weight) is DISTRICT-CONSTANT — SS-1-14 stores one district price on every res-9 cell in that district. Within-district variation in this index comes entirely from the POI mix and the two building terms.
null_means
not composed, for one of four reasons recorded per cell in source_flags._affluence_null.reason: no_land_value (the required term is missing), insufficient_inputs (a priced cell with no buildings and no nearby venues), outside_aoi, or resolution_not_composed (res 8). There is no default of 50, no zero-fill and no imputation — a NULL here means the model declined to guess, not that the area is poor.
inputs
region_baseline_h3.land_value_sqm, region_baseline_h3.building_count, region_baseline_h3.footprint_area_sqm, staging.poi_unified_open, app.siteselect.baseline.affluence_weights

What would change this: a household income or wealth survey published below governorate level, which would make the index calibratable rather than only arguable

Land value resolves to a district, not to a hexagon

read before relying on it

Every square metre price on this platform is the median of one district's monthly transaction prices, and every hexagon inside that district carries the same number. The prices themselves are measured — they are real recorded sales, not a model — but the geography is coarse: a corner plot on an arterial and an interior street two blocks away are indistinguishable here. Districts with no transaction in the window are left empty and are not filled in from their neighbours, which is why a large part of the surface has no price at all.

land_value_sqmaffluence_index
granularity
district
coverage_pct
46.43
districts_priced
131
recency_window
end: 2026-07-01 · start: 2025-08-01 · months: 12 · rationale: staged transactions span 2006-2026 and Riyadh land prices moved by 2-4x over that span; widening to 3 years would add 203 res-9 cells (0.48% of the grid) at the cost of mixing 2024 prices into a current-value column · anchored_to: latest staged monthly period, not wall-clock now()
estimator
median_of_monthly_prices
estimator_rejected
volume-weighted mean — rejected because two single deals of ~7.1bn SAR in Oct 2025 put King Faisal at 83,752 and King Abdulaziz at 68,752 SAR/m², which is one transaction, not a district land value. Per-district volume-weighted means are kept in staging.stg_district_land_value for comparison.
within_district_modulation
none — deliberately not modulated by building density or POI mix, which would make a measured price partly modelled and would feed the affluence model a reflection of its own inputs (SS-1-08)
null_means
no SREM transaction in the recency window reached any district polygon overlapping this cell. NOT imputed from neighbouring districts.

What would change this: parcel-level or street-level transaction geography, which SREM does not publish

The district boundaries have no established licence or attribution

blocks use as a factor

Land value and the affluence index are placed on the map by a district boundary layer that arrived without a source, a licence, or an attribution line. Its schema resembles Saudi National Address / SPL lineage, and that resemblance is not evidence — this platform does not assert a licence it cannot produce. The layer is recorded as needs-review, every cell composed through it is tagged in the data so the dependency can be found and the columns withdrawn wholesale, and neither the boundaries themselves nor anything derived from them is cleared for redistribution until the question is resolved. The transaction prices carried on that geography are SREM's under the Saudi Open Data Licence; that is a separate and settled question.

land_value_sqmaffluence_index
  • source
    KSA district boundaries (user-supplied, provenance unverified)
    licence
    unknown
    provenance_verified
    no
    credit_line
    District boundaries: source unverified. Supplied without attribution or licence; provenance not established. Not cleared for redistribution or client-facing display.
    staging_tables
    staging.stg_ksa_districts, staging.stg_district_h3, staging.stg_district_srem_bridge
    compliance_state
    needs-review

What would change this: an identified publisher and licence for the boundary layer, or a replacement layer whose licence permits commercial use

Age and sex structure does not vary within the city

blocks use as a factor

The age bands are real counts and they sum correctly to the total population, but the shares behind them come from an unconstrained national product: every hexagon in the region carries the same age profile. The counts are usable as a magnitude. The shares cannot distinguish one location from another, so the scoring engine refuses to place weight on them — a weight there would rank candidates on a constant.

pop_age_0_14pop_age_15_29pop_age_30_44pop_age_45_64pop_age_65_plus
distinct_65_plus_shares
1
measured_over_cells
55,230
stddev_65_plus_share
1.11e-10
shares
pop_age_0_14: 0.24180048 · pop_age_15_29: 0.2190024 · pop_age_30_44: 0.30597054 · pop_age_45_64: 0.20383618 · pop_age_65_plus: 0.0293904
safe_use
counts are usable as a magnitude (they sum to `population`); age *shares* are identical in every hex and must never be used as an MCDA factor or presented as a differentiator between candidate locations
ticket
SS-1-12

What would change this: district-level census tables, which need a Riyadh district polygon layer

Having district boundaries is not having a district census — the ±5 % population check still cannot be run

blocks use as a factor

The milestone's exit gate asks whether the hexagon populations sum to within ±5 % of the official population of each district. Answering it needs two things: district boundaries, and an official population per district. The boundaries now exist. The per-district counts do not — the boundary file carries an id, a city, a region, two names and a centroid, and no population field of any kind. Nothing has been substituted for the missing counts and no version of this check has been run and passed. The population surface remains a redistribution of Kontur's own totals, validated against the governorate total it was built from, which is a weaker statement than the gate asks for. For the same reason the age and sex structure still carries no within-city signal, and the banking and clinic weight profiles still refuse to produce a score.

populationpop_age_0_14pop_age_15_29pop_age_30_44pop_age_45_64pop_age_65_plus
district_polygons_staged
3,731
district_polygons_reaching_this_region
173
population_field_in_the_boundary_layer
none — 0 of the population-like column names exist on staging.stg_ksa_districts, checked against the catalogue
boundary_layer_columns
district_id, city_id, region_id, name_ar, name_en, X, Y
profiles_still_refusing_to_score
banking, clinic
note
the exit gate's own state is measured by the SS-1-11 harness and reported under `validation.exit_gate`; it is not restated here

What would change this: GASTAT population counts published per district, which would make the gate evaluable and would give the age columns real spatial variation

`households`, `daytime_pop_index`, `spend_potential` and `retail_gla_sqm` are unavailable — not estimated

blocks use as a factor

These columns are empty. They could each have been filled with a plausible-looking number derived from the data that does exist, and they were not, because every available method produces a rescaling of a column we already have wearing a different label. An empty column is visibly empty; a fabricated one is not, and it would be weighted as a measurement in every score built on top of it. The scoring engine drops these factors and renormalises the remaining weights, and every score says what share of its intended weight was dropped.

householdsdaytime_pop_indexspend_potentialretail_gla_sqm
  • column
    households
    why
    The only household source available is a GASTAT governorate total. Distributing it by a constant households-per-person ratio yields a column that is a scalar multiple of `population` — no spatial signal, no information, but it looks like a measurement and would be weighted as one. A NULL column is honest; a constant-ratio one is not.
    unblocked_by
    SS-1-04 (district calibration), blocked on SS-1-13
    cells_with_a_value
    0
  • column
    daytime_pop_index
    why
    DATA_STRATEGY §2.2 needs five inputs. SS-2-04 delivered two (transit_stops, traffic_exposure); retail GLA is unavailable (SS-1-15) and school/university capacity is not ingested.
    unblocked_by
    a workplace/education capacity source, plus SS-1-15 for GLA
    cells_with_a_value
    0
  • column
    spend_potential
    why
    Two of its four factors are still unavailable. The GASTAT Household Expenditure Survey category basket was not obtained, so there is no basket to multiply by; and the age propensity term would multiply through a spatially constant age structure (SS-1-12 — exactly one distinct 65+ share across all 55,230 hexes). What is left is population × affluence × a constant, published in SAR — a currency-denominated restatement of two columns that already exist, and the currency is what would make it read as a measurement.
    unblocked_by
    a GASTAT HES category basket, and SS-1-04 district age structure
    cells_with_a_value
    0
  • column
    retail_gla_sqm
    why
    Requires building height or storey count. 169 of 756,675 Overture buildings carry a height and 397 carry explicit levels — 0.06% of the stock. Extrapolating GLA from that is fiction, not sparsity (SS-1-15).
    unblocked_by
    SS-1-15 — a height/levels source, or a commercial GLA dataset
    cells_with_a_value
    0

What would change this: SS-1-04 (district calibration), blocked on SS-1-13; SS-1-15 — a height/levels source, or a commercial GLA dataset; a GASTAT HES category basket, and SS-1-04 district age structure; a workplace/education capacity source, plus SS-1-15 for GLA

The score is a suitability index, not a revenue forecast

read before relying on it

A Tier 1 score is an index, not a forecast. It ranks locations against each other on the factors you weighted; it does not predict revenue and is never denominated in SAR. Revenue forecasting is Tier 4 and requires your own stores' sales history.

score
  • tier
    1
    model
    Weighted MCDA
    requires
    nothing beyond the regional baseline
    output
    0–100 suitability index
    available
    yes
  • tier
    2
    model
    Huff / gravity — distance-decay market share
    requires
    a competitor set and an attractiveness proxy
    output
    % share captured per competing site
    available
    no
  • tier
    3
    model
    Analog — nearest neighbour to top-performing stores
    requires
    roughly 20–30 of the tenant's own stores
    output
    “this site resembles your Store 12 and Store 27”
    available
    no
  • tier
    4
    model
    ML regression on trade-area features
    requires
    roughly 50+ stores with sales history
    output
    revenue forecast with a confidence band
    available
    no

What would change this: your own store network and sales history — Tiers 2 to 4 unlock on data you already hold

Section 7

How we check ourselves

The baseline is measured by a re-runnable harness, not signed off by hand. It profiles every column, re-derives the population totals against their sources, and looks for columns that vary too little to carry information. Findings below are everything that is not a pass.

FAIL 5INFO 6N/A 1PASS 74WARN 1measured 2026-08-31 02:40 UTC

The test we cannot run

The milestone’s own acceptance test is DEVELOPMENT_PLAN §4 M1 — hex population sums within ±5% of GASTAT *district* totals. It cannot be evaluated here. Blocker: SS-1-13 — no per-district population (polygons exist, census does not). A governorate-level comparison is reported in Section 2 and is a weaker test — it says how many people the model has, not whether it has them in the right places. It is reported as what it is rather than substituted for the test that was specified.

Everything that is not a pass (7 of 87 checks)

FAILspatial_signalpop_age_0_14

`pop_age_0_14` is a fixed multiple of `population` (ratio CV 2.94e-08 < 1e-04). Its values vary between hexes, but every bit of that variation comes from `population` — as a differentiator it carries no information of its own.

Never present `pop_age_0_14` as a spatial differentiator, and never give it MCDA weight alongside `population` — that double-counts `population` under a second name.

FAILspatial_signalpop_age_15_29

`pop_age_15_29` is a fixed multiple of `population` (ratio CV 3.01e-08 < 1e-04). Its values vary between hexes, but every bit of that variation comes from `population` — as a differentiator it carries no information of its own.

Never present `pop_age_15_29` as a spatial differentiator, and never give it MCDA weight alongside `population` — that double-counts `population` under a second name.

FAILspatial_signalpop_age_30_44

`pop_age_30_44` is a fixed multiple of `population` (ratio CV 4.77e-08 < 1e-04). Its values vary between hexes, but every bit of that variation comes from `population` — as a differentiator it carries no information of its own.

Never present `pop_age_30_44` as a spatial differentiator, and never give it MCDA weight alongside `population` — that double-counts `population` under a second name.

FAILspatial_signalpop_age_45_64

`pop_age_45_64` is a fixed multiple of `population` (ratio CV 3.57e-08 < 1e-04). Its values vary between hexes, but every bit of that variation comes from `population` — as a differentiator it carries no information of its own.

Never present `pop_age_45_64` as a spatial differentiator, and never give it MCDA weight alongside `population` — that double-counts `population` under a second name.

FAILspatial_signalpop_age_65_plus

`pop_age_65_plus` is a fixed multiple of `population` (ratio CV 3.13e-08 < 1e-04). Its values vary between hexes, but every bit of that variation comes from `population` — as a differentiator it carries no information of its own.

Never present `pop_age_65_plus` as a spatial differentiator, and never give it MCDA weight alongside `population` — that double-counts `population` under a second name.

WARNshare_alike_separability

18 populated column(s) rest on ODbL-1.0 data; 8 of them carry the `odbl_derived` flag that makes the share-alike subset selectable in one query.

Add `odbl_derived: true` to population, pop_age_0_14, pop_age_15_29, pop_age_30_44, pop_age_45_64, pop_age_65_plus, road_class_mix, junction_density, traffic_exposure, transit_stops — ODbL reaches them through an input, and separability has to be provable by query, not by reading prose (DATA_STRATEGY §9).

N/Aexit_gatepopulation

The M1 exit gate — hex population sums within ±5% of GASTAT DISTRICT totals — CANNOT BE EVALUATED. It is neither passed nor failed.

Obtain per-district GASTAT population for Riyadh (RCRC opendata.rcrc.gov.sa publishes districtcode/district/districtar attributes; likely needs a KSA network path), join it to staging.stg_ksa_districts, then run SS-1-04 calibration and re-run this harness. A polygon layer alone does NOT unblock this gate.

74 checks passed and are not listed. The harness is re-runnable against any region: python -m app.siteselect.baseline.validate --region Riyadh