This study aims to develop and empirically validate an integrated framework bridging performance audit with policy evaluation, addressing a methodological gap identified since the 1990s.
Drawing on contribution analysis (Mayne, 2012) and methodological triangulation (Denzin, 2017), the framework articulates four interdependent components: theory of change as a testable analytical structure, expanded evaluation criteria integrating the International Standards of Supreme Audit Institutions 300 standards and the Organisation for Economic Co-operation and Development (Organisation Economic Co-operation and Development (OECD)) dimensions, structured multi-actor participation and holistic governance assessment. Three performance audits conducted by a regional SAI in Southern Europe (Sindicatura de Comptes de Catalunya) constitute the empirical basis: social benefits (EUR 978m), extrahospital renal care (378,350 sessions) and hospital service compensation (54.2% of care expenditure).
The theory of change validation exposed broken causal links between policy design and implementation, revealing accountability failures invisible to standard economy, efficiency and effectiveness (3E) criteria: EUR 167.56m in improper payments, outsourced services exceeding hospital provision costs by 12.4% and pricing mechanisms lacking documented economic justification. A cross-case analysis identified six recurrent governance dysfunctions across the three sectors.
The single SAI institutional context limits generalisability. Longitudinal studies are required to confirm the sustainability of identified improvements.
SAIs can implement the framework sequentially without prior legislative changes, progressively building methodological capacity across audit planning, execution and reporting phases.
The framework strengthens public governance accountability by enabling audit mechanisms to diagnose causal determinants of policy failure rather than merely documenting performance shortfalls.
This is the first study to empirically validate an integrated framework combining theory of change validation, expanded OECD evaluation criteria and structured stakeholder participation within an institutional audit context, demonstrating performance audit’s capacity for double-loop institutional learning.
Introduction
Performance audit has expanded substantially over the past four decades as a central instrument for public sector accountability, addressing increasingly complex questions about the value of public policies and programmes (Barrados and Lonsdale, 2020; Overman and Schillemans, 2022). Yet, a persistent methodological tension remains: international standards centre on economy, efficiency and effectiveness (3E), whilst contemporary governance challenges demand analytical capacity to diagnose the causal mechanisms underlying policy failure – mechanisms that 3E criteria cannot systematically address (Power, 1997; Leeuw, 1996; Rana et al., 2022). This study addresses that gap by developing and empirically validating an integrated framework that bridges performance audit with policy evaluation within the institutional context of supreme audit institutions (SAIs). Standards and guidelines for the implementation of performance audits have evolved through successive iterations of International Organisation of Supreme Audit Institutions (INTOSAI) standards. They were first endorsed as field standards in government auditing in 2001 and subsequently reformulated as Fundamental Principles of Performance Auditing in 2013 to establish coherent conceptual foundations and were most recently updated in 2019 within the INTOSAI Framework of Professional Pronouncements (INTOSAI, 2019a, b). According to INTOSAI, the main objective of performance auditing is to constructively promote economical, effective and efficient governance and contribute to accountability and transparency by assisting those with governance and oversight responsibilities to improve performance, thereby contributing to public accountability (INTOSAI, 2019a; Grossi et al., 2023). However, standard performance audit, which focuses exclusively on economy, efficiency and effectiveness, has substantial limitations in assessing public policy outcomes in dynamic governance contexts (Rana et al., 2022). Contemporary governance challenges require an understanding of policy relevance, internal and external coherence, differential impacts on diverse populations and long-term sustainability – dimensions that standard performance audit criteria cannot adequately capture (Organisation for Economic Co-operation and Development [OECD], 2021). These gaps become particularly evident when examining intersectoral and multi-level interventions addressing interconnected social challenges, in which traditional audit frameworks struggle to diagnose systemic coordination failures and access equity issues emerging from fragmented policy implementations. Integrating evaluation methodologies addresses these constraints by enabling auditors to select approaches based on the questions at hand and articulating the different strengths and purposes of both practices in a unified analytical framework (Barrados and Lonsdale, 2020).
Various SAIs have developed integrated methodologies that combine audit frameworks with public policy evaluation approaches; however, this development remains fragmented without offering a systematically integrated framework. This study addresses both academic and practitioner audiences, advancing accountability theory whilst providing an operationally tested framework that SAIs can adapt to their specific institutional and legal mandates. Based on lessons from advanced SAIs, this study develops and empirically validates a system that combines four interdependent components. First, it adopts the theory of change as a conceptual foundation to understand the causal logic of public interventions. Second, it integrates international evaluation criteria as a normative complement to the International Standards of Supreme Audit Institutions (ISSAI) 300 standards to integrate the dimensions of relevance, coherence, impact and sustainability. Third, it incorporates the structured participation strategies of interested parties, beneficiaries, managing entities, service providers, sector experts and civil society organisations throughout the audit process to enrich the analysis with contextual knowledge about the real functioning of public policies.
Three research questions guide this study:
Does integrating theory of change validation, expanded evaluation criteria and structured stakeholder participation generate diagnostic capacity that exceeds standard 3E audit approaches?
What structural governance dysfunctions does such a framework reveal that a conventional performance audit would not systematically detect?
Under what institutional conditions is this integrated framework viable within an SAI context?
The implementation of three performance audits with an evaluative approach by the Sindicatura de Comptes de Catalunya in the domains of social rights (economic benefits of subjective rights, Sindicatura de Comptes de Catalunya, 2025a) and health (provision of extrahospital service for renal insufficiency care, Sindicatura de Comptes de Catalunya, 2024; and consideration system for acute hospital and specialised care services, Sindicatura de Comptes de Catalunya, 2025b) in a Southern European country demonstrates the viability of the hybrid approach and its adaptability to various public management contexts. A comparative analysis of these cases demonstrates that the proposed system improves the quality and applicability of recommendations and identifies systemic inefficiencies inadequately diagnosed by standard performance audits.
Theoretical background and literature review
Theoretical foundations of contemporary performance audit
Academic scholarship has identified several theoretical perspectives through which performance audits have been analysed. New public management (Hood, 1991) reinforced the economy, efficiency and effectiveness (3E) criteria, rooted in pre-NPM public audit traditions, towards results-based assessment (Rana et al., 2022). Public value theory (Moore, 1995) expanded the evaluative horizon to encompass legitimacy, trust and social cohesion, whilst accountability scholarship (Bovens, 2010; Overman and Schillemans, 2022) reconceptualised audit as enabling multiple accountability forms. Such expanded criteria are not without precedent: consulting users and staff already featured in value-for-money work at the UK National Audit Office – though understood there as effectiveness (Midgley et al., 2024) – a strict reading of the 3Es would not readily accommodate them. Critical perspectives have further questioned whether SAIs primarily serve democratic accountability rather than public sector improvement (Pallot, 2003; Ferry and Midgley, 2022) and whether performance criteria are objective or methodologically constructed (Radcliffe, 1998). These traditions differ less in emphasis than in their view of what performance audit is for, from democratic accountability to efficiency. The exchange between Radcliffe (2008, 2011) and Funnell (2011) shows that what an auditor may examine reflects the institutional limits of its mandate; the present study therefore offers a technique deployable within any mandate, not a redefinition of it. However, these perspectives share a common limitation: they advance what a performance audit should assess, without fundamentally addressing how auditors can investigate the causal mechanisms underlying the policy outcomes.
This methodological gap was identified in the 1990s. Power (1997) characterised a performance audit as a “ritual of verification” whose institutional expansion reflected legitimation demands rather than substantive improvement capacity. Leeuw (1996) documented a fundamental tension, i.e. performance auditing adopted evaluative questions about programme effectiveness whilst retaining methodologies ill-suited to answer them. Pollitt et al. (1999) demonstrated how performance audits oscillated between compliance verification and performance assessment through a five-country comparison. These critiques established that the convergence of audit and evaluation constituted not only a technical challenge but also an epistemological question about how accountability institutions generate actionable knowledge about policy effectiveness (Pollitt, 2003; Ferry et al., 2022).
Subsequent studies have developed operational bridges between these two practices. Mayne’s (2012) contribution analysis provided a framework for establishing plausible causal links between interventions and outcomes without experimental designs that are directly applicable to audit contexts in which randomisation is infeasible. Evidence from the Canadian federal audit system illustrates how the theory of change operationalises this bridge in practice, enabling auditors to move from verifying whether resources were used efficiently to examining why policies succeed or fail (Mayne, 2001). Barrados and Lonsdale (2020) synthesised these developments, confirming the viability of the audit–evaluation crossover, whilst documenting that no institution had achieved systematic integration of all dimensions. The present study directly addresses this documented gap, providing the first empirically validated framework that combines theory of change validation, expanded OECD evaluation criteria, and structured stakeholder participation within an operational SAI context. Empirical evidence from SAI studies confirmed that performance audit contributions to improvement depend substantially on stakeholder perceptions, contextual factors and institutional dynamics (Reichborn-Kjennerud and Vabo, 2017; Johnsen, 2019; Michael et al., 2025).
ISSAI 300 (INTOSAI, 2019a) provides the principal internationally recognised normative reference for performance audit within the INTOSAI framework, defining it as an independent examination oriented to determine whether public resources are managed according to the principles of economy, efficiency and effectiveness (par. 9). Whilst the standard acknowledges the importance of stakeholder interaction, it conceptualises participation primarily as unidirectional consultation rather than knowledge co-construction, limiting stakeholder involvement to providing contextual information rather than shaping methodology or conclusions (par. 29). This configuration reflects a structural limitation with direct theoretical implications: 3E criteria are designed to verify whether established objectives were achieved efficiently, not to interrogate whether those objectives are adequately defined or whether the causal assumptions embedded in policy design are valid – a distinction that Argyris and Schön (1996) would characterise as the difference between single-loop and double-loop learning.
The integration of GUID 9020 (INTOSAI, 2019c) and OECD evaluation criteria (OECD, 2021) into performance audit practice addresses this structural limitation. GUID 9020 establishes a framework for structured stakeholder participation as a systematic audit component, specifying that policy evaluation “does not only consist in correcting administrative dysfunctions but rather in improving a policy” (par. 4.2) whilst protecting auditor independence through explicit safeguards: the evaluating entity maintains “the final word throughout the process of evaluation” (par. 4.2), and the advisory body “shall not under any circumstances make a decision on the methodology and the conclusions” (par. 5.2). This approach positions performance audit as a contributor to institutional change rather than a mere verification mechanism (Reichborn-Kjennerud and Vabo, 2017).
Comparative international evidence
Various SAIs have individually incorporated elements of public policy evaluation into performance audit, although the development remains fragmented. To identify which methodological components have been adopted and their potential for systematic integration, this study examines nine audit institutions selected to represent diverse legal-administrative contexts (Pollitt and Bouckaert, 2017): parliamentary audit systems (Auditor General of Canada, Australian National Audit Office and Riksrevisjonen of Norway); collegial court models (Cour des Comptes of France and Tribunal de Cuentas of Spain); supranational governance (European Court of Auditors); federal executive models (Government Accountability Office of the United States of America and Tribunal de Contas da União of Brazil) and sub-national audit bodies operating within devolved frameworks (Audit Scotland and Cámara de Cuentas of Andalusia). These categories are inherently imprecise: some institutions display features of more than one and sit across their boundaries. Selection prioritises documented empirical evidence and analytical relevance over comprehensive coverage (Peters, 2021; Ferry et al., 2023).
Canada and Norway have developed the most systematic evidence on performance audit effects. Morin (2014) differentiated instrumental effects (direct recommendation implementation), conceptual effects (renewal of professional perspectives) and strategic effects (influence on decision-making), with a longitudinal analysis of Canadian federal audits (2001–2011) revealing a significantly higher conceptual impact (4.48 on a 7-point Likert scale) than instrumental or strategic effects, a differentiation subsequently replicated in Belgium and the Netherlands (Desmedt et al., 2017). Reichborn-Kjennerud (2013) surveyed 353 Norwegian public officials and demonstrated that audit usefulness depends on auditee acceptance of criteria and institutional credibility. Reichborn-Kjennerud (2014) identified four institutional response strategies: capitulating, defying, copying and ignoring, revealing that the impact depends primarily on agreement with audit conclusions. However, Nordrum and Vabo (2024) documented that Norwegian practice prioritises regulatory compliance, with state economic regulations appearing as audit criteria in approximately 80% of reports. Both institutions maintain the exclusive application of standard 3E criteria, exclude external stakeholders from audit processes and provide no explicit theory of change linking audit recommendations to institutional improvement.
France and Scotland have pioneered distinct participation models. Since 2022, the Cour des Comptes has institutionalised a digital platform for citizen participation that has generated 58 citizen-initiative audit reports, demonstrating significant public engagement (Cour des Comptes, 2024). However, participation operates exclusively in topic selection, without involvement in methodological design or execution, and audit methodology maintains standard performance audit criteria. Audit Scotland implements a complementary model through advisory groups with external specialised members for its annual programme of approximately 14 performance audit reports, integrating expert participation throughout audit execution in its “Best Value” statutory framework (Audit Scotland and Accounts Commission, 2023). Further, the European Court of Auditors has demonstrated that cooperative audit methodologies function across multi-level governance contexts with diverse administrative traditions. Despite these advances, participation in all cases remains bounded: citizen engagement addresses topic selection but not methodology, expert advisory groups concentrate primarily on local governments and none of these participation models are integrated with expanded evaluation criteria beyond economy, efficiency and effectiveness.
Spain represents the most developed integration of evaluative dimensions in operational auditing. Coordinated training with the Institute of Fiscal Studies and the National Institute of Public Administration since 2019 has prepared audit personnel for evaluation methodologies. Additionally, operational audits of the Demographic Challenge Plan (2021–2023) assessed equity and social impact along with economy, efficiency and effectiveness. Garde Roca et al. (2023) produced the first systematic methodological guide for applying an “evaluative approach” to performance audits, whilst the Cámara de Cuentas of Andalusia implemented this framework on social benefits and healthcare programmes (Orellana Hidalgo and Salguero Carretero, 2023). The Australian National Audit Office has separately formalised ethics as a fourth criterion through three audit scenarios aligned with INTOSAI Standard ISSAI 3000 (ANAO, 2023, 2024), and the Government Accountability Office has incorporated equity assessments into its evaluation standards. Equity and ethics, however, are far from uncontroversial criteria: subjective and politically contested, they demand particular prudence and transparency about the criteria applied. Nevertheless, practical implementation reveals persistent challenges in counterfactual analysis, indicator definition and long-term impact measurement (Orellana Hidalgo and Salguero Carretero, 2023), and none of these institutions have integrated expanded criteria with structured stakeholder participation throughout the audit cycle.
Thus, international comparative analysis reveals three individually demonstrated innovations: rigorous impact measurement systems (Canada and Norway), structured participation mechanisms (France and Audit Scotland) and expanded evaluation criteria (Spain and ANAO). Additional innovations in multilateral audit coordination (the Tribunal de Contas da União of Brazil’s governance evaluation scale applied across 11 Latin American countries for 2030 Agenda implementation) confirmed the institutional capacity for methodological experimentation in diverse governance contexts. However, no examined models had systematically integrated all three dimensions into a single coherent framework. Three factors explain this persistent fragmentation. First, audit and evaluation maintain distinct professional cultures, training pathways and epistemological assumptions that create institutional silos even within the same organisation (Leeuw, 1996; Barrados and Lonsdale, 2020). Second, legal mandates in most jurisdictions assign audit and evaluation to separate organisational units, discouraging methodological crossover. Third, comprehensive integration requires substantial investment in interdisciplinary capacity that few SAIs have prioritised.
This integration gap, i.e., combining theory of change as an analytical structure, evaluation criteria beyond the standard “3 Es” and structured stakeholder participation throughout the audit cycle, constitutes the strategic opportunity that the proposed framework addresses.
Integrated framework: performance audit with an evaluative approach
Conceptual foundations
Before describing the components of the framework, it is necessary to specify the meaning of methodological integration in this context. Integration is distinct from juxtaposition (applying audit and evaluation tools simultaneously without mutual interaction) and subordination (one tradition absorbing the other). Based on Denzin’s (2017) triangulation framework and Greene’s (2007) typology for combining methods, this study defines integration as the systematic articulation of analytical tools from audit and evaluation traditions, whereby each component informs the application of others, producing a diagnostic capacity that is not generated independently. Greene (2007) distinguished five purposes for combining methods. This framework serves primarily the purposes of complementarity (evaluation criteria reveal dimensions that audit verification alone cannot examine) and initiation (theory of change generates analytical questions absent from standard audit criteria).
However, this definition has operational consequences, as integration requires that the theory of change reconstruction shapes criterion selection, stakeholder participation informs both the causal model and evidence strategy, and governance assessment synthesises findings across all sources. Thus, the four components function as an interdependent system, rather than as a sequential checklist.
The theory of change provides a conceptual foundation for understanding the causal chains through which public interventions generate outcomes (Weiss, 1995; Chen, 1990). Its application in this framework operates through two distinct analytical moments. The reconstructive moment maps the assumed causal logic of how resources enable activities to produce outputs that generate outcomes. This identifies the explicit and implicit assumptions underlying policy design, which are frequently undocumented and sometimes internally contradictory (Funnell and Rogers, 2011). The validating moment is analytically decisive, as it subjects reconstructed assumptions to empirical scrutiny, examining whether the hypothesised causal links hold in practice. Following Mayne’s (2001, 2012) contribution analysis, validation assesses plausible causal contribution by examining the evidence strength for each link, identifying where the links are confirmed, weakened or broken. Thus, the theory of change serves not as static programme logic but as a testable analytical framework whose empirical confrontation reveals critical disconnection points between policy design and implementation. This distinction between reconstruction and validation differentiates this approach from using the theory of change as a purely descriptive planning tool.
Standard performance audit criteria, namely 3E, focus on resource management and objective achievement. However, contemporary governance challenges require the assessment of additional dimensions beyond these three criteria. Additionally, the OECD (2021) proposed six evaluation criteria – relevance, coherence, effectiveness, efficiency, impact and sustainability – that function as “multiple lenses” interrelating to offer a holistic understanding of public intervention. Furthermore, OECD (2019) guidance advocates that the application must be “reflexive and context-adapted” rather than “mechanical,” meaning criteria selection and prioritisation occur according to intervention complexity and audit objectives, enabling deep investigation of how governance inadequacies manifest across multiple dimensions.
GUID 9020 establishes a framework for structured stakeholder participation as a systematic audit component. The guide specifies that policy evaluation “does not only consist in correcting administrative dysfunctions but rather in improving a policy” (INTOSAI, 2019c, par. 4.2). This improvement orientation must be read against each institution’s mandate: what a policy ought to be is an inherently political question that, in many jurisdictions, lies outside the auditor’s remit (Funnell, 2011). It engages actors involved in policy implementation and affected by policy outcomes. Critically, GUID 9020 protects auditor independence through explicit safeguards:
The evaluating entity maintains “the final word throughout the process of evaluation” (INTOSAI, 2019c, par. 4.2);
The advisory body “shall not under any circumstances make a decision on the methodology and the conclusions of the evaluation because these issues are the exclusive responsibility of the independent evaluator” (INTOSAI, 2019c, par. 5.2);
The entity in charge of evaluation “is solely responsible for the decision to undertake the evaluation and shall refuse it when the criteria on the object and the requirements on the process are not met” (INTOSAI, 2019c, par. 5.1).
The guide indicates that stakeholders, implementing administrative entities, locally elected representatives, private entities (NGOs, companies, professional organisations, and unions) and representatives of beneficiaries may “be involved in the choice of the specific object of the evaluation,” “be active participants in the evaluation,” “benefit from interim or final reports” and “have a role in the post-evaluation decision-making process” (INTOSAI, 2019c, par. 4.2). Nevertheless, this participation operates under the requirement that “the entity conducting the evaluation should give an independent opinion on its own on the findings, analyses, conclusions, and recommendations of the public policy evaluation” (INTOSAI, 2019c, par. 6.3). Thus, the evaluating entity maintains exclusive control over methodological determinations and conclusions.
Unlike the ISSAI 300, which confines stakeholder participation to unidirectional, time-limited external consultations, GUID 9020 proposes co-construction (collaborative) models that integrate stakeholders into a structural component. These models safeguard methodological autonomy whilst identifying governance inefficiencies and determining improvement pathways.
Performance audit with an evaluative approach combines performance audit rigour with an evaluative methodology. This integrated approach represents an evolution that enhances policy impact whilst maintaining the professional rigour and independence that define the discipline.
Components and methodology
The integrated methodological framework proposed in this study connects four interconnected components that operationalise the conceptual foundations established above into concrete analytical procedures that are systematically applied throughout the complete audit cycle (Figure 1).
A flowchart illustrating an integrated framework for performance audit with an evaluative approach. The diagram consists of four interconnected pillars leading to improved public governance. Pillar 1, labeled Theory of Change, describes a causal chain mapping from resources to activities, outcomes, and impacts. Pillar 2, labeled Integrated Criteria, combines ISSAI 300 principles of economy, efficiency, and effectiveness with OECD principles of relevance, coherence, impact, and sustainability, applied contextually based on intervention type. Pillar 3, labeled Structured Stakeholder Engagement, involves planning through indirect consultation, execution through testimonial evidence, and reporting through validation, while maintaining auditor independence. Pillar 4, labeled Holistic Governance Assessment, evaluates six dimensions: planning, resources, coordination, outcomes, verification, and follow-up, providing a comprehensive view of the governance system. Integrated framework for performance audit with an evaluative approach. Source: Authors’ elaboration
A flowchart illustrating an integrated framework for performance audit with an evaluative approach. The diagram consists of four interconnected pillars leading to improved public governance. Pillar 1, labeled Theory of Change, describes a causal chain mapping from resources to activities, outcomes, and impacts. Pillar 2, labeled Integrated Criteria, combines ISSAI 300 principles of economy, efficiency, and effectiveness with OECD principles of relevance, coherence, impact, and sustainability, applied contextually based on intervention type. Pillar 3, labeled Structured Stakeholder Engagement, involves planning through indirect consultation, execution through testimonial evidence, and reporting through validation, while maintaining auditor independence. Pillar 4, labeled Holistic Governance Assessment, evaluates six dimensions: planning, resources, coordination, outcomes, verification, and follow-up, providing a comprehensive view of the governance system. Integrated framework for performance audit with an evaluative approach. Source: Authors’ elaboration
Practical implementation of the theory of change involves systematic reconstruction of the causal chain during the planning phase and mapping how assigned resources enable specific activities that produce measurable results. This reconstruction detects critical junctures where the chain may break, resources fail to enable activities, activities generate unintended consequences, or immediate results do not produce the expected long-term impacts. Additionally, auditors require three specific competencies: (1) analysis of complex systems with multiple institutional levels and sectoral actors, (2) understanding of public policy design in the context of uncertainty and incomplete information and (3) the capacity to triangulate from partial documentation, such as administrative records, stakeholder testimony and contextual evidence.
The second component integrates the evaluation criteria that combine the ISSAI 300 standard audit principles with OECD evaluation criteria. These criteria are reflexively applied, contextualised and operationalised through the formulation of hierarchical audit questions that transform generic questions into specific and relevant ones. The application of criteria is prioritised according to the specific objectives of each evaluation, adapting to the context and particular characteristics of the audited intervention.
Stakeholder participation constitutes the third component, applying specific mechanisms of GUID 9020 to integrate multiple perspectives whilst preserving methodological integrity. The degree of involvement varies according to stage: during planning, stakeholder needs are considered without compromising autonomy in topic selection; during execution, collaboration intensifies to obtain testimonial evidence from multiple perspectives; and during report elaboration, contradictory validation is maximised, focused on substantiating findings and evidence rather than validating methodological choices, whilst preserving the decision-making authority of the auditing entity. Recent studies on SAIs recognise that systematic engagement with stakeholders contributes to increasing the pressure to implement audit recommendations and strengthen the accountability system (OECD, 2023; INTOSAI-P 12, 2019).
The fourth component develops a holistic governance assessment, i.e., arrangements that are put in place to ensure that the intended outcomes for stakeholders are defined and achieved (CIPFA/IFAC, 2014). This assessment examines six interconnected dimensions: capacity to strategically plan and establish coherent objectives, adequacy in allocation and management of available resources, effectiveness of interinstitutional and cross-level coordination mechanisms, capacity to produce expected results, robustness of performance monitoring and control systems and the existence of continuous improvement and institutional adaptation processes.
The practical implementation follows the structured methodological sequence established by ISSAI 300 for performance audit, integrating four components throughout the three main phases of the process (INTOSAI, 2019a, par. 35). During planning, auditors construct a theory of change, identify contextual criteria adapted to a specific environment and design appropriate stakeholder participation. In the execution phase, evidence collection from multiple sources is combined with rigorous methodological triangulation to guarantee sufficiency and adequacy. Elaboration integrates the systematic analysis of criteria with a contradictory validation phase involving interested parties (INTOSAI, 2019a, par. 29) and develops recommendations oriented to strengthening institutional capacities (INTOSAI, 2019a, paras. 36–40).
Finally, applying a performance audit with an evaluative approach requires analysing the specific criteria that establish its viability. According to GUID 9020, an application is appropriate when three fundamental conditions are met (INTOSAI, 2019c, paras. 4.1–5.1). The first is policy importance, which is determined by budgetary volume, number of actors involved and scope of potential effects. The second is the measurability of effects, allowing distinction between immediate results and long-term impacts and the establishment of causal relationships. The third is the sufficient time elapsed since inception, a minimum of two to three years, to generate evaluable effects beyond immediate ones. These criteria guided the selection of regional cases to illustrate the practical viability and transformative results of the proposed framework.
Methods
The Sindicatura de Comptes de Catalunya is the external audit institution of the regional public sector, created by regional legislation. Although sub-national, the Sindicatura shares the defining features of an SAI: a statutory basis in the Statute of Autonomy of Catalonia (articles 80–81), Act 18/2010, and an ISSAI-aligned mandate. It is a founding member of EURORAI. Its institutional characteristics position it favourably for methodological innovation: organic dependence on Parliament, with full organisational, functional and budgetary autonomy, guarantees independence from audited governments; the broad scope of competence that embraces the regional government, local entities and the entire public sector allows transversal methodological experimentation; and the capacity to elaborate its own activity programme facilitates autonomous strategic decisions. In this institutional context, the Sindicatura’s current strategic plan orients the organisation towards evaluative methodologies, reinforcing this trajectory by incorporating new specialised professional profiles and specific training programmes for technical staff. This favourable institutional context has allowed its application in three audits that exemplify its practical viability.
A methodological caveat warrants acknowledgement. The researchers participated in the design and application of the framework being evaluated, creating a potential reflexivity concern. Three safeguards mitigate this risk. First, all empirical findings derive from administrative data verifiable through public records and official databases, not from researcher judgement. Second, the structured contradictory validation phase subjects preliminary findings to scrutiny by audited entities and sectoral experts, who can challenge interpretations. Third, theory of change validation follows Mayne’s (2012) contribution analysis protocol, which requires explicit evidence assessment for each causal link rather than holistic researcher judgement. These safeguards do not eliminate reflexivity but constrain it within verifiable boundaries.
The study selected the three audits conducted by the Sindicatura de Comptes de Catalunya according to the criteria established by GUID 9020 (INTOSAI, 2019c, pars. 14–18): (1) substantive budgetary relevance; (2) availability of administrative data for quantitative analysis; and (3) sufficient implementation duration (≥3 years) to identify systemic effects: economic benefits of subjective right (EUR 978 million, 16 social benefit programmes managed across fragmented information systems and more than 3 years of implementation), provision of extrahospital service for renal insufficiency care (EUR 98m, 378,350 annual sessions, territorial variability and contracts established in 2014) and consideration system for acute hospital and specialised care services (54.20% of care expenditure excluding COVID-19 funding, complex price mechanism and more than 10 years of trajectory) (Sindicatura de Comptes de Catalunya, 2024, 2025a, 2025b).
These three audits integrate performance audit elements encompassing regulatory compliance and financial verification using an expanded evaluative approach. The compliance dimension examines the adequacy of the contracting system against relevant regulatory requirements and identifies irregularities, such as null and void contracts due to untimely signature or extensions exceeding legal deadlines. The financial dimension verifies price and tariff adequacy, detects billing inconsistencies, and evaluates whether economic studies support service costs.
The methodological implementation has combined qualitative techniques with quantitative analyses. The quantitative strategy combines statistical analysis of administrative data to identify territorial distribution patterns and service variations; population weighting techniques to extrapolate sample findings to broader populations; sensitivity testing to verify result robustness across different assumptions; and comparative analysis across provider types and territorial units to identify systematic cost differentials and service access variations. Qualitative analysis, closely linked to stakeholder participation mechanisms, takes the form of semi-structured interviews and systematic consultations with identified actors that provide testimonial evidence from multiple perspectives, validating quantitative results and discovering dimensions often absent in administrative data.
Each audit combines administrative data analysis with structured stakeholder consultations through a triangulation strategy (Denzin, 2017) designed to cross-validate quantitative patterns against qualitative contextual evidence.
Social benefits: The Sindicatura analysed payment data from programmes managed by eight separate information technology applications operated by four different providers. The fragmentation of information systems, a finding in itself, requires separate data requests to each provider and additional validation checks to verify data integrity. Improper payments were identified through systematic cross-referencing of payment records against eligibility registers held by the Social Security Treasury and regional databases, detecting cases where beneficiaries simultaneously received incompatible benefits or payments continued after eligibility conditions ceased. A budget gap analysis compared allocated resources against estimated need using 2022 data from the Living Conditions Survey and Household Panel. Access coverage was estimated using administrative records of approved applications against target populations derived from official poverty statistics.
Renal care: The Sindicatura examined billing records from contracted extrahospital centres and hospital cost data. The annual cost per extrahospital haemodialysis patient was calculated by aggregating session fees, transport costs and material costs billed to the public health service, which were derived from hospital activity and cost data for equivalent treatments. Billing irregularities were detected through patient-day matching of billing records, identifying instances in which different centres invoked dialysis sessions for the same patient on the same day. Territorial access analysis mapped patient distribution against the distance threshold established in the contractual specifications. Stakeholder consultations included actors identified in the planning phase (see stakeholder identification, above).
Hospital compensation. The Sindicatura analysed pricing and contractual data using the minimum basic dataset (CMBD) hospital discharge registry and contractual addendums. Price comparisons across fiscal years and provider categories examined activity in six categories: hospitalisation, outpatient consultations, emergency services, specific procedures, high-complexity care and result-based compensation. Contractual incentive effectiveness was assessed by comparing the stated objectives with measured outcomes using provider-reported data and hospital activity records.
Across all three cases, the triangulation strategy operated at three levels: quantitative findings from administrative data analysis identified patterns; stakeholder consultations provided contextual explanations and identified dimensions not captured in the quantitative data; and documentary analysis of legislation, contracts and planning documents established a normative framework against which findings were assessed. This integration of evidence sources enabled the theory of change validation described in the results section.
Structured procedures separated consultative stakeholder input from audit decisions to maintain audit independence. The auditors conducted interviews and systematic consultations with the identified stakeholders, audited entities, sectoral experts, health professionals, scientific societies and representative associations in each sectoral domain. The planning phase identified implementation actors, the execution phase integrated stakeholder consultations and the validation phase presented preliminary findings for accuracy verification whilst preserving auditor control over conclusions and audit methodology.
The proposed framework translates its components into hierarchical audit questions (see Table 1) that systematically link the reconstructed theory of change with criteria selected for each specific context. This formulation transforms generic questions into specific interrogatives that examine the underlying processes of evaluated policies and identify critical implementation points, allowing the detection of where and why public interventions may generate unexpected results or fail to achieve the intended objectives.
Audit question formulation process
| Theory of change component | International criteria | Audit question examples (from actual reports) | Focus |
|---|---|---|---|
| Objectives | Relevance | Social Benefits: Were they planned based on identified needs? | Systemic |
| Renal Care: Were resources planned according to identified needs? | |||
| Hospital Services: Are pricing objectives coherent with the Health Plan? | |||
| Objectives and resources | Coherence | Social Benefits: Are they coherent among themselves and with complementary social and sectoral policies? | Systemic |
| Renal Care: Is territorial planning coherent and does it address coordination gaps between care levels? | |||
| Hospital Services: Do services and programmes align with strategic health plan axes? | |||
| Resources | Economy | Social Benefits: Are resources and amounts sufficient to cover needs? | Systemic |
| Renal Care: Was service cost determined using economy criteria? | |||
| Hospital Services: Does contracting adhere to uniform, objective criteria respecting economy principles? | |||
| Resources and activities | Efficiency | Social Benefits: Do procedures comply with timelines? What controls have been implemented? | Process |
| Renal Care: Are resources sufficient and is their relationship with services efficient? | |||
| Hospital Services: Does compensation appropriately reflect structural differences and case complexity among providers? | |||
| Outcomes and impacts | Effectiveness | Social Benefits: Has access been guaranteed? What outcomes have been achieved? | Results |
| Renal Care: What mechanisms monitor user satisfaction and clinical outcomes? Are they systematically applied? | |||
| Hospital Services: Have economic incentives achieved their intended effects? | |||
| Outcomes | Equity | Social Benefits: Does the eligible population correspond to identified social needs, particularly among vulnerable and territorial groups? | Results |
| Renal Care: Do policies address vulnerable groups’ needs? | |||
| Governance (cross-cutting) | Transparency and accountability | Social Benefits: Are there transparency and oversight mechanisms for improper payments? | Systemic |
| Renal Care: Is there monitoring to ensure contract compliance and quality? | |||
| Hospital Services: Does contracting apply objective criteria and maintain documented economic justification for pricing? |
| Theory of change component | International criteria | Audit question examples (from actual reports) | Focus |
|---|---|---|---|
| Objectives | Relevance | Social Benefits: Were they planned based on identified needs? | Systemic |
| Renal Care: Were resources planned according to identified needs? | |||
| Hospital Services: Are pricing objectives coherent with the Health Plan? | |||
| Objectives and resources | Coherence | Social Benefits: Are they coherent among themselves and with complementary social and sectoral policies? | Systemic |
| Renal Care: Is territorial planning coherent and does it address coordination gaps between care levels? | |||
| Hospital Services: Do services and programmes align with strategic health plan axes? | |||
| Resources | Economy | Social Benefits: Are resources and amounts sufficient to cover needs? | Systemic |
| Renal Care: Was service cost determined using economy criteria? | |||
| Hospital Services: Does contracting adhere to uniform, objective criteria respecting economy principles? | |||
| Resources and activities | Efficiency | Social Benefits: Do procedures comply with timelines? What controls have been implemented? | Process |
| Renal Care: Are resources sufficient and is their relationship with services efficient? | |||
| Hospital Services: Does compensation appropriately reflect structural differences and case complexity among providers? | |||
| Outcomes and impacts | Effectiveness | Social Benefits: Has access been guaranteed? What outcomes have been achieved? | Results |
| Renal Care: What mechanisms monitor user satisfaction and clinical outcomes? Are they systematically applied? | |||
| Hospital Services: Have economic incentives achieved their intended effects? | |||
| Outcomes | Equity | Social Benefits: Does the eligible population correspond to identified social needs, particularly among vulnerable and territorial groups? | Results |
| Renal Care: Do policies address vulnerable groups’ needs? | |||
| Governance (cross-cutting) | Transparency and accountability | Social Benefits: Are there transparency and oversight mechanisms for improper payments? | Systemic |
| Renal Care: Is there monitoring to ensure contract compliance and quality? | |||
| Hospital Services: Does contracting apply objective criteria and maintain documented economic justification for pricing? |
Cross-case analysis: from theory to results
These three audits revealed recurring institutional dysfunctions that transcend sectoral particularities and confirmed the utility of the integrated approach. Three critical dysfunction dimensions recurred across cases:
First, strategic planning had significant deficiencies that affected multiple dimensions of the public policy cycle. In social benefits, inadequate planning was evident in the absence of updated needs studies quantifying the severe poverty gap estimated at EUR 1,183.29 million annually, 121% of assigned resources (EUR 978.29m in recognised obligations in 2022). In renal care, obsolete planning was evident in a strategic plan dating from 2008 without subsequent updates, generating territorial inequalities with 9.79% of patients living more than 20 km from the nearest dialysis centre, exceeding the established maximum threshold. In the hospital consideration system, planning deficiencies appeared in the absence of periodic evaluation of the alignment between pricing objectives and health-plan priorities, allowing certain incentives to become anachronistic.
Second, problems with supervision and control generated measurable economic impacts and revealed structural deficiencies in governance systems. In social benefits, the audit detected EUR 167.56m in improper payments during the 2016–2024 period, mainly caused by insufficient automatic controls to detect incompatibilities between benefits and other income, lack of effective integration between administrative databases and weaknesses in verification systems that allow persistence of irregularities for prolonged periods without detection. In renal care, the audit identified billing irregularities worth EUR 4.33 million, specifically in 31,574 records of duplicate invoices, revealing the absence of a robust automated billing validation system and a lack of periodic audits of patient care sheets. In hospital service consideration, control deficiencies manifested as serious problems of legal non-compliance: provision of services by the regional public health provider on behalf of the regional public health purchaser was not done under the corresponding contract programme, an instrument provided for in legislation; untimely formalisation of additional clauses of agreements with contracted centres was observed; and the absence of effective verification mechanisms for the achievement of established contractual objectives. These deficiencies reflect a broader problem: lack of integration between information systems, insufficient result orientation of control mechanisms and absence of systematic monitoring protocols allowing early detection of irregularities.
Third, the absence of cost studies that justify public policy decisions. In terms of social benefits, the absence of cost-benefit evaluation studies prevented determining the efficiency of monetary transfers compared with other alternative policies to reduce poverty, as evidenced by the disproportion between the real poverty gap (EUR 1,183.29m) and assigned resources (EUR 978.29m). In renal care, the estimated annual cost per patient for extra-hospital haemodialysis (EUR 47,744) significantly exceeded hospital costs (EUR 42,461) and peritoneal dialysis costs (EUR 29,724) without documented cost analysis or methodological justification. In hospital service consideration, the absence of cost studies manifested in multiple dimensions: lack of comparative analyses between public and contracted sector costs, absence of studies justifying the adequacy of established prices compared to real service costs and lack of cost-effectiveness evaluations allowing the determination of whether incentive mechanisms generated expected care results in relation to the investment made.
Beyond these common patterns, each case revealed specific sectoral problems. In renal care, the audit identified facts that could constitute indications of collusive behaviour between winning companies, with the geographical market distribution maintaining the status quo without real competition. In social benefits, the audit detected the risk of conflict of interest in the management and supervision of some benefits outsourced to non-profit collaborating entities, for which a framework of regularity and transparency protecting the general interest in public policy implementation had not been guaranteed. Access inequalities were manifested in both social benefits and healthcare. In social benefits, the analysis revealed that resources are quantitatively insufficient for large households with dependent children, as benefit amounts failed to lift them above the poverty threshold despite receiving support. In renal care, 29.3% of patients with university education received home dialysis treatment, whereas only 6.8% of those with primary education accessed dialysis, a governance gap that requires further investigation and contradicts the principles of health equity. Additionally, peritoneal dialysis, internationally recognised as cost-effective, was underutilised, representing only 8.7% of dialysis patients, well below the European referents.
The identified deficiencies that systematically interconnect all three cases reveal six critical governance dimensions (Table 2), revealing broader structural problems of the regional public system.
Governance dimensions affected by sector
| Governance dimensions | Social benefits | Renal care | Hospital service compensation |
|---|---|---|---|
| Strategic planning | No needs-based eligibility framework; benefit map fragmented | Strategic plan obsolete; care level definitions and territorial criteria undefined | System objectives misaligned with health policy directives |
| Resource management | Allocated resources insufficient; no adjustment mechanisms | Extrahospital pricing exceeded hospital costs; insufficient home-based treatment funding discouraged cost-effective modalities | Price determination lacked economic justification; providers grouped arbitrarily |
| Control systems | Monitoring absent; improper payments widespread. Externalised management lacked oversight frameworks | Billing validation inadequate: duplicate sessions detected for identical patients | Contract program not formalised; annual clauses signed after service delivery began |
| Intersectoral coordination | No coordination with local authorities; fragmented coverage | Lack of coordinated care pathways and shared protocols between hospital and outpatient nephrology services limited care continuity | Payment incentives conflicted with emergency care policy; multiple funding mechanisms created asset overfunding risks |
| Monitoring and evaluation | No systematic evaluation of policy impact or beneficiary satisfaction | Clinical outcomes unmonitored; user satisfaction not formalised | Provider payment timing inadequate; financial incentives ineffective |
| Transparency and accountability | Conflict of interest management inadequate; no transparency frameworks for partner entities | Procurement lacked documented justification; collusive indicators revealed; territorial criteria undisclosed | Service assignments made without documented or transparent criteria |
| Governance dimensions | Social benefits | Renal care | Hospital service compensation |
|---|---|---|---|
| Strategic planning | No needs-based eligibility framework; benefit map fragmented | Strategic plan obsolete; care level definitions and territorial criteria undefined | System objectives misaligned with health policy directives |
| Resource management | Allocated resources insufficient; no adjustment mechanisms | Extrahospital pricing exceeded hospital costs; insufficient home-based treatment funding discouraged cost-effective modalities | Price determination lacked economic justification; providers grouped arbitrarily |
| Control systems | Monitoring absent; improper payments widespread. Externalised management lacked oversight frameworks | Billing validation inadequate: duplicate sessions detected for identical patients | Contract program not formalised; annual clauses signed after service delivery began |
| Intersectoral coordination | No coordination with local authorities; fragmented coverage | Lack of coordinated care pathways and shared protocols between hospital and outpatient nephrology services limited care continuity | Payment incentives conflicted with emergency care policy; multiple funding mechanisms created asset overfunding risks |
| Monitoring and evaluation | No systematic evaluation of policy impact or beneficiary satisfaction | Clinical outcomes unmonitored; user satisfaction not formalised | Provider payment timing inadequate; financial incentives ineffective |
| Transparency and accountability | Conflict of interest management inadequate; no transparency frameworks for partner entities | Procurement lacked documented justification; collusive indicators revealed; territorial criteria undisclosed | Service assignments made without documented or transparent criteria |
Applying the framework to three regional cases empirically validated the effectiveness of the four constitutive pillars of the performance audit framework using an evaluative approach, demonstrating how each component contributes to insight generation.
The theory of change validation followed the two-stage process described above: the reconstruction of assumed causal logic, followed by empirical testing of each critical link (Mayne, 2012; Funnell and Rogers, 2011). In social benefits, the reconstructed causal chain assumed the following: legislative recognition of subjective right → adequate budget allocation → administrative processing → monetary transfer → poverty reduction. Validation revealed that the resource link was structurally underfunded (EUR 978.29m allocated covered 82.7% of the estimated EUR 1,183.29m poverty gap); the processing link was broken by fragmented information systems across 16 programmes, which generated EUR 167.56m in improper payments between 2016 and 2024; the access link was critically compromised, with approximately 50% of entitled individuals failing to access benefits; and the final poverty reduction link could not be validated as no outcome monitoring systems existed. In renal care, the assumed logic posited that outsourcing to extrahospital centres would expand territorial access and improve efficiency. Validation showed that territorial access was partially confirmed but unevenly distributed (9.79% of patients beyond the 20 km threshold), whilst the efficiency assumption was refuted: extrahospital haemodialysis cost per patient (EUR 47,744) exceeded hospital provision (EUR 42,461) by 12.4% and a finding triangulated with billing data revealed EUR 4.33 million in duplicate invoicing across 31,574 records. In hospital compensation, the reconstructed logic assumed that contractual incentives would improve care quality and efficiency. Validation revealed unequal effectiveness: surgical activity incentives proved effective, readmission penalties showed no measurable impact and outpatient surgery incentives generated inefficient resource use when targets were already met without documented cost-effectiveness justification.
The reflexive application of the expanded criteria identified problems not systematically addressed by the mechanical application of 3E principles. The incorporation of relevance, coherence and equity criteria revealed the structural problems of intersectoral coordination and access inequalities that affect intervention effectiveness.
The differentiated participation of the involved parties has enriched the analysis, providing perspectives often absent in standard performance audits. Consultations with scientific societies on renal care have revealed misalignments between clinical practices and contractual incentives. Interviews with social workers identified barriers to access that were not detected in the quantitative analysis. In hospital consideration, consultations with hospital managers and providers substantiated the audit evidence of inadequate objective criteria in resource allocation and misalignment between incentive systems and care objectives.
These structural problems allowed the formulation of recommendations that go beyond operational adjustments to address systemic causes (Table 3).
Structural recommendations identified in the three case studies
| Audit | Structural recommendations with transformative potential |
|---|---|
| Social benefits (EUR 978 million) | Define benefit objectives with targets and indicators enabling accountability and evidence-based decisions |
| Integrate information systems enabling real-time cross-database verification and transparent data traceability | |
| Reduce provision complexity; eliminate separation between state pensions and guaranteed citizenship income | |
| Implement automated beneficiary verification and database integration preventing structural payment failures | |
| Renal care (378,350 sessions) | Approve updated strategic plan renewing prevalence projections and care level definitions, update haemodialysis centre location criteria |
| Estimate the cost of different dialysis services and update tariffs paid to various service providers | |
| Re-evaluate peritoneal and home haemodialysis tariffs to facilitate hospital provision of home-based techniques | |
| Extend maximum wait-time guarantees to vascular access interventions; accelerate arteriovenous fistula placement | |
| Hospital service compensation (54.20% of care expenditure, excluding COVID-19 funding) | Establish uniform provider unit criteria: each unit must correspond to one acute hospital facility |
| Promote modification whereby hospital discharge pricing is grounded in a case-complexity-centred payment model using case-mix weights estimated from actual costs, with necessary adjustments for centre structure factors | |
| Develop uniform analytical accounting framework for all regional hospital network acute hospitals to enable reliable estimations of actual regional hospital costs, replacing current US-based cost estimates for accurate case-mix weighting |
| Audit | Structural recommendations with transformative potential |
|---|---|
| Social benefits (EUR 978 million) | Define benefit objectives with targets and indicators enabling accountability and evidence-based decisions |
| Integrate information systems enabling real-time cross-database verification and transparent data traceability | |
| Reduce provision complexity; eliminate separation between state pensions and guaranteed citizenship income | |
| Implement automated beneficiary verification and database integration preventing structural payment failures | |
| Renal care (378,350 sessions) | Approve updated strategic plan renewing prevalence projections and care level definitions, update haemodialysis centre location criteria |
| Estimate the cost of different dialysis services and update tariffs paid to various service providers | |
| Re-evaluate peritoneal and home haemodialysis tariffs to facilitate hospital provision of home-based techniques | |
| Extend maximum wait-time guarantees to vascular access interventions; accelerate arteriovenous fistula placement | |
| Hospital service compensation (54.20% of care expenditure, excluding COVID-19 funding) | Establish uniform provider unit criteria: each unit must correspond to one acute hospital facility |
| Promote modification whereby hospital discharge pricing is grounded in a case-complexity-centred payment model using case-mix weights estimated from actual costs, with necessary adjustments for centre structure factors | |
| Develop uniform analytical accounting framework for all regional hospital network acute hospitals to enable reliable estimations of actual regional hospital costs, replacing current US-based cost estimates for accurate case-mix weighting |
A cross-case analysis of the three regional cases demonstrated the three key contributions of the integrated framework: (1) identification of structural inefficiencies not systematically addressed by standard performance audits, (2) generation of recommendations with transformative potential and (3) understanding of the underlying causes facilitating organisational learning.
Discussion
These findings illuminate a theoretical mechanism underlying the framework’s diagnostic capacity. Standard 3E criteria operate within what Argyris and Schön (1996) term as single-loop learning: they detect deviations from established objectives but cannot question whether objectives themselves are adequate. The integrated framework enables double-loop learning by using theory of change to surface and test the assumptions embedded in policy design. When validation revealed broken causal links, as in the social benefits case where the assumed poverty reduction pathway was structurally disconnected, the analytical focus shifted from whether resources were used efficiently to why the policy failed to achieve its stated purpose. This shift from compliance verification to causal explanation constitutes the framework’s core theoretical contribution to public sector accountability literature.
The empirical analysis revealed that structural dysfunctions were not systematically addressed by standard performance audits that focused exclusively on economy, efficiency, and effectiveness. Analysis of the studied cases generated structural recommendations: simplification of fragmented benefit maps, integration of dispersed information systems, intersectoral coordination protocols and objective resource allocation criteria that address causal mechanisms instead of procedural adjustments.
This evidence directly challenged the theoretical foundation of public accountability (Overman and Schillemans, 2022; Brandsma and Schillemans, 2013). Most studies focus on formal – reports, appearances and sanctions – without examining whether they generate tangible institutional improvements. The studied cases showed that focusing exclusively on formal mechanisms without verifying their real impact generates gaps in detecting organisational dysfunctions, limiting the effectiveness of SAIs in identifying improvements. Within the public sector improvement tradition, this suggests that accountability mechanisms oriented towards governance enhancement benefit from demonstrating that decisions generate socially valuable results, rather than merely documenting them and accepting the consequences. In this context, the participation of social agents served as a key element in overcoming these limitations.
The structured participation of social agents constituted evidence of real policy functioning beyond democratic legitimation (Bryson et al., 2013; Baldwin, 2019). Beneficiaries, providers, professionals and organisations contributed contextual knowledge about how policies function in practice, identifying unanticipated effects. However, stakeholder engagement operated at the institutional level, managing entities, professionals and scientific societies, without direct participation of individual beneficiaries. This reflects the operational constraints inherent to SAI audit processes, but it also means that experiential knowledge about access barriers and service quality is mediated through institutional actors rather than directly captured. Following Arnstein’s (1969) participation typology, the model operates at the consultation and placation levels, not at the co-construction level. Expanding audit scope also redistributes analytical authority, raising questions about who defines governance failure and whose knowledge constitutes evidence (Radcliffe, 1998), a dimension future studies should examine alongside mechanisms for direct beneficiary participation (Cousins and Whitmore, 1998; Dassen and Lavin, 2024).
The analysis revealed limitations in the conditioning generalisation. The cases benefit from a favourable context: high constitutional autonomy, collaborative administrative culture, specialised technical resources and an openness to innovative approaches. Not all SAIs operate in similar contexts. A related validity concern warrants consideration: whether the framework itself generates the documented diagnostic capacity, or whether competent auditors would have reached similar findings without it. Two observations address this. First, the specific dysfunctions identified: broken causal links, refuted efficiency assumptions and structural planning disconnects – emerged directly from the reconstructive-validating sequence that the theory of change provides, a procedure absent from standard audit methodology. Second, the cross-case pattern of recurring governance failures across three unrelated sectors suggests systematic analytical capacity rather than case-specific institutional insight. Comparative literature documents the difficulties of transferring administrative innovations between diverse contexts (Rana et al., 2022). Administrative traditions, political cultures, technical capacities and social expectations can profoundly affect the viability and effectiveness of integrated approaches. The proposed framework requires contextual adaptation rather than mechanical application. In jurisdictions where constitutional or legislative mandates restrict SAIs to standard 3E criteria, implementation requires prior engagement with Parliament, the primary institutional actor authorising SAI mandate expansion (Ferry and Midgley, 2022; Pallot, 2003), before methodological integration can proceed. Where mandates already permit evaluative dimensions, the sequential implementation described above offers a viable path.
The case-based design also limits definitive causal inference, since other contextual factors may have shaped the observed results.
Conclusion
This study makes three interrelated theoretical contributions to the public sector accounting and accountability literature. First, it operationalises the audit–evaluation convergence – identified but not resolved since Power (1997) and Leeuw (1996) – into an empirically validated methodological framework applicable within institutional audit contexts. Second, it advances accountability theory by demonstrating that performance audit mechanisms can enable double-loop institutional learning, shifting the analytical focus from documenting performance shortfalls to diagnosing the causal determinants of policy failure. Third, it extends contribution analysis (Mayne, 2001, 2012) from evaluation practice into SAI audit contexts, confirming its viability without experimental designs. This integration is methodological rather than institutional: performance audit and policy evaluation retain distinct mandates and accountability functions. The framework does not propose that SAIs become evaluation agencies but that they selectively incorporate evaluative tools where intervention complexity warrants deeper causal analysis, thereby generating added value beyond standard verification and diagnoses of why policies fail, not merely whether they do.
The empirical evidence from three performance audits demonstrates that the framework identifies accountability failures invisible to standard 3E criteria: EUR 167.56m in improper payments, outsourced services exceeding hospital provision costs by 12.4% and pricing mechanisms lacking documented economic justification.
The empirical evidence demonstrates that incorporating evaluative elements strengthens, rather than compromises, the fundamental value of public audits. The structured participation of actors involved in audited policies, beneficiaries, service providers, sector professionals and social organisations improved analytical quality through contextual knowledge about the real functioning of policies, whereas the theory of change provided methodological rigour to understand the underlying causal mechanisms. This synthesis challenged the dichotomy between control verification and institutional improvement by repositioning performance audit with an evaluative approach as an institutional learning mechanism oriented towards public governance transformation.
For SAIs operating under mandates that permit evaluative dimensions, the gradual implementation of the integrated framework offers a viable path without immediate legislative changes; where mandates are more restrictive, the framework identifies the methodological aspiration whilst acknowledging the institutional prerequisites. A practical implementation sequence emerges from the empirical evidence: SAIs can begin by incorporating theory of change reconstruction into existing audit planning processes, subsequently expanding criteria selection based on intervention characteristics and finally integrating structured stakeholder consultation as institutional capacity develops. This sequencing allows progressive capability building without disrupting established audit programmes. Evidence from Canada and Norway confirms that investment in methodological training generates measurable returns in recommendation implementation rates (Reichborn-Kjennerud and Vabo, 2017; Morin, 2014). Whilst the three cases confirm operational viability and enhanced diagnostic capacity, questions of scalability and generalisability across diverse administrative traditions remain open.
Implementing this framework requires investment beyond standard performance audit practice, with additional resource concentration in the planning phase. Auditor training in evaluation methodology and the possible incorporation of public policy analysts represent a strategic human capital investment that strengthens long-term institutional capacity whilst generating diagnostic returns and identification of structural governance failures and causal disconnections, which standard 3E approaches do not systematically produce.
Future studies should address three priorities: a longitudinal analysis of impacts to establish the sustainability of changes; adaptation of the framework to diverse administrative contexts to determine transferability limits; and exploration of synergies with other democratic governance instruments. These lines will refine the framework and provide precise guidance for its implementation in various institutional contexts.
Artificial intelligence assistance disclosure
The authors used Claude (Anthropic) for English language editing to meet journal formatting requirements. All conceptual development, methodological design, data analysis and substantive conclusions are the authors’ own.

