Evidence map for part II: central findings, supporting sources, evidence strength and jurisdictional depth
| Finding | Supporting sources | Evidence strength | Jurisdictional depth |
|---|---|---|---|
| Out-of-sample validation and backtesting are prerequisites for defensible AVM deployment | Schulz et al. (2014), Steurer et al. (2021), Krause et al. (2020) | Established | Broad |
| Prediction intervals, calibration and distributional or conformal methods should accompany every AVM output; in line with supervisory expectations | Rigby and Stasinopoulos (2005), Wheeler et al. (2014), Angelopoulos and Bates (2023), Granna et al. (2025), European Union (2024), Federal Register (2024) | Established (methods); emerging (regulatory codification) | EU, US, UK |
| Data quality, density and granularity set the upper bound of AVM performance more fundamentally than model complexity | Openshaw (1984), Granna et al. (2022), Krause et al. (2020), Steurer et al. (2021) | Established | Broad |
| Concept drift and temporal instability materially affect AVM reliability and require systematic monitoring, not ad hoc maintenance | Gama et al. (2014), Gudigantala and Mehrotra (2024) | Established | Primarily US illustration; drift theory jurisdiction-agnostic |
| Alternative and multimodal data can enrich AVM inputs, but added value is context-dependent and often proprietary | Koch et al. (2019), Despotovic and Brunauer (2024) | Emerging | Broad |
| ESG signals influence property values (green premiums, brown discounts, stranding risk), but ESG data infrastructure and standardisation lag behind regulatory demand | Eichholtz et al. (2010), Cajias et al. (2019), Hirsch et al. (2019), RICS (2024), European Union (2020, 2022) | Emerging | EU-focused |
| Synthetic data can support prototyping and stress testing but are not a substitute for real transaction data; risks include model collapse | Goodfellow et al. (2014), Kingma and Welling (2014), Shumailov et al. (2024), Alemohammad et al. (2023), Bidanset (2025) | Emerging | Broad |
| AVMs can reproduce and amplify historical bias; disparities across racial and socioeconomic groups are empirically documented in US markets | Howell and Korver-Glenn (2021), Freddie Mac (2021), Zhu et al. (2024), Neal et al. (2020), CFPB (2023) | Established (US); Under-researched (elsewhere) | US-focused |
| Formal fairness criteria cannot be jointly satisfied; mitigation is a governance choice, not a purely technical one | Kleinberg et al. (2017), Barocas et al. (2019) | Established (theoretical) | Broad |
| Hybrid human AVM architectures are the most resilient operating model under current regulatory expectations; practitioner adoption depends on institutional trust and clarity of responsibility | Kasim et al. (2026), Renigier-Biłozor et al. (2022), Glumac and Des Rosiers (2021) | Emerging (practitioner evidence); Conceptual (framework) | Broad (with EU emphasis) |
| Finding | Supporting sources | Evidence strength | Jurisdictional depth |
|---|---|---|---|
| Out-of-sample validation and backtesting are prerequisites for defensible AVM deployment | Established | Broad | |
| Prediction intervals, calibration and distributional or conformal methods should accompany every AVM output; in line with supervisory expectations | Established (methods); emerging (regulatory codification) | EU, US, UK | |
| Data quality, density and granularity set the upper bound of AVM performance more fundamentally than model complexity | Established | Broad | |
| Concept drift and temporal instability materially affect AVM reliability and require systematic monitoring, not ad hoc maintenance | Established | Primarily US illustration; drift theory jurisdiction-agnostic | |
| Alternative and multimodal data can enrich AVM inputs, but added value is context-dependent and often proprietary | Emerging | Broad | |
| ESG signals influence property values (green premiums, brown discounts, stranding risk), but ESG data infrastructure and standardisation lag behind regulatory demand | Emerging | EU-focused | |
| Synthetic data can support prototyping and stress testing but are not a substitute for real transaction data; risks include model collapse | Emerging | Broad | |
| AVMs can reproduce and amplify historical bias; disparities across racial and socioeconomic groups are empirically documented in US markets | Established (US); Under-researched (elsewhere) | US-focused | |
| Formal fairness criteria cannot be jointly satisfied; mitigation is a governance choice, not a purely technical one | Established (theoretical) | Broad | |
| Hybrid human AVM architectures are the most resilient operating model under current regulatory expectations; practitioner adoption depends on institutional trust and clarity of responsibility | Emerging (practitioner evidence); Conceptual (framework) | Broad (with EU emphasis) |
Note(s): Evidence strength for commercial systems is based on public disclosures and secondary literature; internal validation data are not accessible. “Emerging” reflects findings supported by fewer than three independent studies or restricted to single jurisdictions. “Established” reflects findings replicated across multiple studies and jurisdictions
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.