Purpose

This paper contributes to the field of public services’ performance measurement systems by proposing a benchmarking-based methodology that improves the effective use of big and open data in analyzing and evaluating efficiency, for supporting internal decision-making processes of public entities.

Design/methodology/approach

The proposed methodology uses data envelopment analysis in combination with a multivariate outlier detection algorithm—local outlier factor—to ensure the proper exploitation of the data available for efficiency evaluation in the presence of the multidimensional datasets with anomalous values that often characterize big and open data. An empirical implementation of the proposed methodology was conducted on waste management services provided in Italy.

Findings

The paper addresses the problem of misleading targets for entities that are erroneously deemed inefficient when applying data envelopment analysis to real-life datasets containing outliers. The proposed approach makes big and open data useful in evaluating relative efficiency, and it supports the development of performance-based strategies and policies by public entities from a data-driven public sector perspective.

Originality/value

Few empirical studies have explored how to make the use of big and open data more feasible for performance measurement systems in the public sector, addressing the challenges related to data quality and the need for analytical tools readily usable from a managerial perspective, given the poor diffusion of technical skills in public organizations. The paper fills this research gap by proposing a methodology that allows for exploiting the opportunities offered by big and open data for supporting internal decision-making processes within the public services context.

The social and economic challenges that governments face put increasing pressure on public entities to improve services for citizens with efficient solutions while avoiding higher spending. Consequently, governments have introduced several business-inspired practices and tools, such as performance measurement and management (Bouckaert and Halligan, 2008; Barbato and Turri, 2017; Mwita, 2000), for a smarter government that is more efficient and closer to community needs (Criado and Gil-Garcia, 2019; Twizeyimana and Andersson, 2019). In this context, the efficient use of resources plays a crucial role in the sustainable development of public entities by directing available resources toward improving the quantity and quality of services to guarantee effectiveness and equity in meeting citizens’ needs (Andrews and Entwistle, 2014; Pollitt and Bouckaert, 2011).

A performance measurement system (PMS) is defined as the processes, tools, and mechanisms used to identify objectives and support strategic processes and ongoing management through analysis, planning, measurement, and control of performance (Ferreira and Otley, 2009). Many scholars have focused on the measurement and evaluation of the efficiency of public services and the effects that business-like management practices and instruments can generate (Pawsey et al., 2018), analyzing specific services, public service networks, and public entities as a whole (Agostino and Arnaboldi, 2018; Speklé and Verbeeten, 2014).

In setting up PMSs, problems concerning the nature of public services must be addressed, particularly when the lack of a competitive market makes it difficult to assess the efficiency and effectiveness of public entities, due to the absence of useful parameters to identify challenging targets. In this case, data benchmarking can play a relevant role, acting as a useful tool for comparing entities within the same field of activity, supporting the identification of challenging efficiency targets to inform feedback mechanisms, and creating a new culture of comparison as a source of learning (De Witte and Geys, 2011; Dorsch and Yasin, 1998).

Big and open data offer interesting opportunities to improve the identification of targets in contexts such as the aforementioned, as the vast amount of open information can be used to compare outputs, activities, and resources, to enable benchmarking with similar entities, and to identify the most efficient ones (Maciejewski, 2017; Rogge et al., 2017). The expression “big and open data” generally refers to big data produced and released by organizations and collected as publicly available data (Roth et al., 2020; Weerakkody et al., 2017). These are huge amounts of information characterized by volume, variety, openness, and interoperability that can enhance the PMS and internal decision-making processes (Chen et al., 2012; Kwon et al., 2014; Badia and Donato, 2022).

Big and open data can also hide some challenges for their effective use. First, there is uncertainty over the quality of such datasets. Many big and open datasets have duplicate, inconsistent, and missing data, making it more difficult to generate value from their use. Accuracy problems related to data errors in (bad manual) collection and measurement processes are often reported (Sadiq and Indulska, 2017). This data quality uncertainty is a threat to the effective use of big and open data within PMSs for the improvement of efficiency processes.

Another significant challenge for big and open data in the context of public services is the lack of skills in Information and Communication Technologies (ICT) to manage the complexity of such data (Valle-Cruz, 2019; Manyika et al., 2011). Analytical tools that yield readily interpretable information are, therefore, needed. These should allow managers to identify the best performers and set benchmarking targets, favoring a scientific approach that overcomes the weaknesses of a purely subjective and discretionary approach (Donthu et al., 2005; Coupet et al., 2021).

This paper’s aim is to propose a benchmarking-based methodology that improves the feasibility of using big and open data in the efficiency evaluation of organizations providing public services, to exploit the potential of these data in supporting internal decision-making processes and to overcome the critical issues mentioned above. The proposed methodology can support policymakers and managers in defining challenging targets and assessing their performance. In particular, the study employs data envelopment analysis (DEA) in combination with a multivariate outlier detection technique—the local outlier factor (LOF) algorithm—to ensure the effective use of available information in the presence of multidimensional datasets with anomalous data. An empirical implementation of the methodology is conducted on waste management services in Italy, for which big and open data pertaining to municipalities are available. Notwithstanding, the methodology is replicable in other kinds of public services.

The remainder of the paper is structured as follows: Section 2 reviews the literature on the topics of interest, and the study’s data sources and methodology are described in Section 3. Section 4 presents and discusses the main results, and Section 5 draws conclusions and possible further research development.

The adoption and implementation of PMSs has been a leading topic related to public sector reforms undertaken in recent decades around the world (Hood, 1995; Humphrey et al., 2005). New public management (NPM) has promoted PMSs because their principles and tools support economic rationality and a focus on results (Hood, 1991), although a strand of the literature highlights that NPM effects seem to be controversial, as they encourage business-inspired practices not aligned with public values, such as equity and impartiality (Broadbent and Laughlin, 1998; Brignall and Modell, 2000; Bejerot and Hasselbladh, 2013). However, there is unquestionably a need to measure the performance of public sector organizations to provide reliable and comparable information about their operations and activities, to support decision-making processes, and to inform citizens and other stakeholders (de Kool and Bekkers, 2016; Steccolini et al., 2020; Melo and Mota, 2020).

In particular, the topic of efficiency measurement and evaluation related to the provision of public services has been widely explored in the literature due to the growing demand for quantity and quality, together with the scarcity of financial resources. In fact, the amount of dedicated resources can hardly be increased, as raising taxes and public debt beyond a certain level is neither appropriate nor politically convenient. This represents one of the more relevant reasons for the great attention paid to the proper functioning of the PMSs of entities that provide public services (Benito et al., 2019; Lo Storto, 2016). This has entailed the growing use of quantitative techniques to determine the efficiency of public utilities, the premise of services’ delivery improvement, and thus public value creation (Bannister and Connolly, 2014; Elston et al., 2018; Worthington, 2000).

Big and open data can improve the efficiency measurement and performance management of public organizations. Although the literature on the potential that big and open data can offer to public value creation is quite rich (Ruijer et al., 2023; Zuiderwijk and Janssen, 2014), it essentially focuses on open government issues. Scholars have examined aspects such as public entities’ transparency and accountability toward the community, as well as citizens’ involvement in providing new services to better respond to their needs (Schmidthuber et al., 2019; Hitz-Gamper et al., 2019), thus focusing on the external impact of the use of such data. Conversely, the topic of how big and open data can influence internal decision-making processes in public entities, thereby sustaining the definition of public strategies and programs, is very scarcely explored. According to the OECD (Van Ooijen et al., 2019), this is an important area of application and development of big and open data that can become an essential source of information for supporting the performance measurement of public organizations to pursue efficiency and effectiveness (Rogge et al., 2017). These data can be employed to support data benchmarking, increasing both the number of entities compared and the information on the inputs and outputs of their processes (Song et al., 2017).

Big and open data also present limitations and challenges (Boyd and Crawford, 2012; Picazo-Vela et al., 2012). Some scholars identify two broad categories of critical issues: “privacy-related problems” and “technical difficulties” (Desouza and Jacob, 2017). In the context of the public sector, the latter are exacerbated by the poor diffusion of technical capabilities that are necessary to fully exploit big and open data (Grundke et al., 2018). For the same reason, the ability to process vast amounts of data is a critical challenge for many local entities (Mergel et al., 2016). Therefore, analytical tools that are understandable from a managerial view should be available to public managers to help them extract value from data and improve the PMS (OECD, 2019). Thus, the literature proposes several applications of DEA in the big and open data context to evaluate the relative efficiency of a vast group of entities called decision-making units (DMUs) in the case of multiple input and multiple output processes (Chen and Jia, 2017; Chu et al., 2018; Badiezadeh et al., 2018). Some of these applications also occur in the public sector, particularly in the field of environmental issues (Zhu, 2022).

The use of big and open data poses a further issue related to the quality of the dataset. Data quality dimensions, such as accuracy, completeness, and consistency, are a fundamental notion in performance measurement processes based on big and open data. As the issue of performance measurement is often not considered when big and open data are originally produced and collected (Sadiq and Indulska, 2017), the quality of huge amounts of data flowing across different data sources can become a problem. Users deal with unexplored and large datasets that potentially contain incorrect, incomplete, and inconsistent data or those lacking integrity. These can generate misleading performance evaluations and, thus, decisions with undesirable impacts or negative consequences for the community (OECD, 2015). This issue is especially critical for data-driven benchmarking, such as DEA benchmarking, because the presence of outliers in the reference set can compromise the identification of inefficient entities and the subsequent assignment of improvement targets (Johnson and McGinnis, 2008). In other words, only if the data quality condition is satisfied can DEA be correctly used to define inefficient operating entities and calculate the efficiency gap, enabling policymakers to develop performance-based public policies aimed at improving the results of both single public organizations and the public sector as a whole (Guerrini et al., 2015).

The literature provides a wide range of techniques for assessing and improving the quality of data. Most of these are analyzed in the context of mathematical-statistical studies and aim to identify possible data quality errors by detecting the records of datasets that can be considered outliers (Batini et al., 2009). Outlier detection is the focus of several review papers that highlight the advantages and limitations of various techniques (Markou and Singh, 2003a, b; Hodge and Austin, 2004). These are distinguished by several criteria, such as the dimension of the feature space (one or multiple), the data distribution (known or unknown), the range of surrounding data points (global or local), and the data labels (present or absent). Among these techniques, local outlier algorithms, such as the LOF, are especially suitable for real-world and multidimensional datasets characterized by variable densities, no a priori knowledge of the data distribution, and no labeled data, all features that are very frequent in the context of big and open data (Smiti, 2020; Alghushairy et al., 2021). To the best of our knowledge, no empirical studies have applied these algorithms in the context of big and open data.

In summary, the relevant literature highlights an important gap that this paper aims to address. The opportunities of big and open data for PMSs in the public sector are still largely unrealized (Jensen et al., 2023; European Commission, 2022; Van Ooijen et al., 2019), and empirical research proposing practices on how public organizations can handle big and open data to support internal decision-making processes is needed. In particular, very few empirical studies have explored how to make the use of big and open data more feasible for this purpose (Rogge et al., 2017; Di Vaio et al., 2022; Abuljadail et al., 2023), overcoming the constraints related to data quality and the need for analytical tools that are readily usable from a managerial perspective, given the poor diffusion of technical skills in public organizations (Sadiq and Indulska, 2017; Weerakkody et al., 2017; OECD, 2019).

This paper aims to contribute to filling this gap within the strand of the literature on improving the technical features of PMSs in the public sector (Garengo and Sardi, 2021; Fryer et al., 2009).

To achieve the research aim, a methodology that combines DEA with the LOF method was employed on a specific public utility—waste management—in the context of the separated waste collection services provided in Italy, for which big and open data are made available from the Italian Institute for Environmental Protection and Research (ISPRA).

The research protocol can be summarized in the following three consequential steps: (1) the entities are grouped into homogeneous subgroups to control for environmental heterogeneity (on the basis of the available data, the most suitable exogenous variables are chosen); (2) taking into account the peculiarities of big and open data, the LOF algorithm is applied within each subgroup, to identify and remove local outlier; and (3) within each subgroup, DEA is applied for the remaining entities to calculate the efficiency scores and the targets to be assigned. To evaluate the advantage obtainable from this procedure, in terms of more achievable targets, steps 1 and 3 can be applied again without step 2 (identifying and removing the outliers), comparing the targets thus obtained with those previously calculated with the described procedure.

To realize the first step, the data retrieved from the ISPRA database related to DMUs (the municipalities) were grouped into homogeneous subgroups to control for environmental heterogeneity (Charnes et al., 1981). In many real-life applications, non-homogeneity is common among DMUs due to diverse environmental variables, so efficiency evaluation must deal with these misleading differences to avoid bias (Dyson et al., 2001). Splitting the set of DMUs into multiple groups allows each DMU to be evaluated against only true peers, that is, those whose environmental contexts are similar to its own (Cook et al., 2015). This solution is very intuitive compared to others proposed for controlling environmental heterogeneity (Sarra et al., 2017; Fried et al., 2002), and it is therefore effective from a managerial point of view.

The second step of the methodology was to apply an outlier detection technique to each subgroup of municipalities. Although outlier detection is not a negligible step in data analysis, in many articles on DEA applications in a real “big and open data” environment, the issue of outliers or techniques used to identify them often does not emerge explicitly (Chen and Jia, 2017; Chu et al., 2018; Zhu et al., 2017). Given the practical significance of treating outliers in DEA-based benchmarking, the proposed methodology emphasizes the use of a technique specifically suitable for certain characteristics of datasets in a “big and open data” context, such as variable densities, missing data, or data errors (Manyika et al., 2011; Badiezadeh et al., 2018), and the impracticality of pre-labeling data as outliers due to the huge volume of datasets. For this reason, the LOF algorithm, an unsupervised outlier detection technique based on a local approach, was used to identify and remove outliers in each subgroup of municipalities. This algorithm does not require any assumptions about the distribution of data.

The LOF algorithm evaluates the degree of outlyingness based on the level of isolation of a data point with regard to the surrounding neighborhood. The basic idea is that the density around an outlier differs significantly from the density around its neighbors. Therefore, a data point can be considered an outlier, even if it is a short distance from an extremely dense group of neighbors. Consequently, for datasets with variable densities, the algorithm gives better results than the global approach, which may not consider such a data point an outlier. To understand how to calculate the LOF, the following three key concepts are needed:

k-distance (A): The distance between point A and its k-th nearest neighbor. The k-distance neighborhood, denoted by Nk(A), includes a set of points that lie in or on the circle of radius k-distance, centered on point A. The size of Nk(A) is always equal to or greater than k (if two or more neighbors are at the same distance from A).

Reachability distance RD(A, Xj): The maximum between the k-distance of point A and the distance between A and another point Xj. If point Xj lies within the k-neighborhood of A, the reachability distance will be the k-distance of A; otherwise, it will be the distance between A and Xj.

Local reachability density LRD(A): inverse of the average reachability distance of A from its neighbors (Xi belonging to Nk(A) with i = 1, …, k).

Based on these concepts, the LOF can be obtained using the following equation:

where

  • k = number of nearest neighbors of point A (defined by the analyst)

  • Nk(A) = set of nearest neighbors of point A (Xi, i = 1, …, k)

  • |Nk(A)| = size of the local neighborhood of point A

  • LRDk(Xi) = local reachability density of point Xi from its k-neighbors

  • LRDk(A) = local reachability density of point A from its k-neighbors

LOFk(A) expresses the degree to which point A can be considered a local outlier. A value nearly equal to 1 indicates that point A has a density similar to that of its neighbors and thus is not an anomalous point, whereas values significantly larger than 1 indicate a higher likelihood of outlyingness. The empirical cumulative distribution function of LOF values can be used to suggest a cutoff value above which a DMU is deemed an outlier. In fact, as noted in the literature (Coles et al., 2001; Li et al., 2022), the LOF values with the highest cumulative probabilities at which the increase in the frequency of individual values levels off can be considered “rare events” in the dataset (Pokrajac et al., 2007). Hence, DMUs with LOF values at which the empirical cumulative distribution function flattens out (generally corresponding to a cumulative probability above 95%) can be removed as outliers.

In the third step, efficiency measurement and the consequent assigning of targets were conducted by applying DEA for the remaining DMUs in each subgroup. Regarding the efficiency concept, a classic definition of economic efficiency was employed, that is, the relationship between the costs of the inputs used and the outputs obtained (Pollitt and Bouckaert, 2011; Andrews and Entwistle, 2014). Since the efficiency objective is to minimize the input for a given level of output, an input-oriented DEA model was used to calculate the scores (Coelli et al., 2005). This model, assuming variable returns to scale (VRS), can be expressed in a dual form as follows:

subject to

where

  • xij = quantity of input i consumed by the j-th DMU

  • yrj = quantity of output r produced by the j-th DMU

  • λj = weights of outputs and inputs of the j-th DMU

  • si = input slacks

  • sr = output slacks

  • ɛ = non-Archimedean value (smaller than any positive real number and greater than 0)

DMU k is efficient if and only if θk = 1 and all slacks are zero. Due to the removal of outliers in the previous step, the efficiency scores generated by the model are not influenced by anomalous data and enable the obtaining of achievable targets for inefficient units. To highlight the usefulness of applying a local outlier detection technique in combination with DEA to define adequate targets for DMUs, the values resulting from this last step were compared with those obtained with the application of DEA without removing outliers.

As specified above, the proposed methodology was employed on separated waste collection services provided in Italy, for which big and open data are available. In Italy, municipalities are responsible for services and activities connected to waste management (collection, transportation, and treatment). They may manage these services directly or outsource them to companies often owned by the same municipalities; however, the service remains public, as the overall responsibility and objectives pursued lie with the municipalities.

Data were retrieved from the ISPRA website, which collects and organizes information acquired and processed by entities involved in waste management services that must answer a questionnaire and fill out a specific form composed of various columns and sections. Therefore, missing data or entry errors may occur, meaning that the ISPRA dataset, which pertains to all Italian municipalities and is made available for anyone interested, is user-dependent (Price and Shanks, 2005).

The dataset retrieved from ISPRA comprised 3,205 municipalities, all for which disaggregated data were available on separated waste collection costs (SWCC), treatment and recycling costs (TRC), and the volume of separated waste collection (SWC). The municipalities were grouped according to two segmentation variables: (1) geographical location, defined with respect to the three macro-areas of Italy (north, center, and south), and (2) population density, adjusted to account for incoming tourist flow. For each municipality, data on the arrival of nonresidents were collected from the website of the Italian National Institute of Statistics (ISTAT). These two items, often cited in the literature as exogenous variables that may influence the efficiency of waste management services (Bosch et al., 2000; Benito et al., 2019), effectively control for environmental heterogeneity in the Italian context, enabling the detection of major differences in local operational conditions between municipalities. By crossing the two segmentation variables, 12 clusters of municipalities were obtained (3 macro-areas × 4 quartiles of adjusted population density), as shown in Table 1.

Table 1

Segmentation variables

NorthValle D'Aosta, Piedmont, Liguria, Lombardy, Emilia-Romagna, Trentino-Alto Adige, Veneto, Friuli-Venezia Giulia
Geographical location: macro-areasCenterTuscany, Umbria, Marche, Lazio
SouthAbruzzo, Molise, Campania, Apulia, Basilicata, Calabria, Sicily, Sardinia
Adjusted population density: quartiles (inhabitants per km2)1<86
2(86–226]
3(226–664]
4>664

Source(s): Authors own work

SWCC and TRC were considered inputs, and the total volume of SWC was considered output (Worthington and Dollery, 2001; De Jaeger et al., 2011; Sarra et al., 2017; Romano and Molinos-Senante, 2020). The data were normalized to take into account the number of inhabitants in each municipality, thus eliminating the size effect.

To demonstrate the relevance of the proposed methodology, DEA was first applied to each subgroup of municipalities without removing outliers. The input-oriented DEA model based on VRS was solved using DEA Frontier™ software. Having already normalized the cost values by the number of inhabitants and having used the population density as a variable for grouping the municipalities, the assumption of VRS allows to implicitly take into account additional factors that influence the scale efficiency (e.g. the size of the municipality’s area) and to concentrate on technical efficiency. Table 2 reports the DEA results; notably, the percentage of municipalities with low efficiency scores was very high in each cluster (Zhu, 2000), as highlighted in the last line of the table.

Table 2

Interval distribution of efficiency scores for each cluster

Efficiency score% of municipalities
NorthCenterSouth
123412341234
[0–0.1)0.90.20.00.00.00.80.00.02.11.00.00.0
[0.1–0.2)9.94.84.71.55.44.21.20.022.425.210.920.0
[0.2–0.3)27.433.626.718.214.723.519.87.429.431.133.930.3
[0.3–0.4)25.126.928.633.114.016.814.816.716.415.519.420.0
[0.4–0.5)14.615.118.619.314.716.011.116.78.86.810.99.0
[0.5–0.6)6.47.410.311.014.015.113.611.16.15.88.54.5
[0.6–0.7)4.73.63.85.69.34.214.813.03.93.93.64.5
[0.7–0.8)2.93.62.32.98.54.26.29.33.61.94.81.9
[0.8–0.9)2.01.92.32.73.13.48.65.61.51.50.61.9
[0.9–1)1.21.30.52.44.72.50.01.92.11.50.60.6
15.01.72.23.411.69.29.918.53.65.86.77.1
Total100100100100100100100100100100100100
<0.577.880.778.672.148.861.346.940.779.179.675.279.4

Source(s): Authors own work

If the measurement and control system took these scores into account to identify future objectives, many municipalities would be assigned very high goals, which could be excessively demanding. Table 3 shows the percentage of DMUs in each cluster that should achieve an improvement in their performance through the reduction of both inputs (SWCC and TRC) by more than 50%. For example, in cluster North-2, almost all municipalities should at least halve their costs: 87.8% for SWCC and 83.6% for TRC. In absolute terms, this means that, on average, each of these municipalities should reduce SWCC by 110,429 euros per year (on an average current value of around 150,000 euros) and transport costs by 46,470 euros per year (on an average current value of around 62,000 euros).

Table 3

Input reductions greater than 50% for SWCC and TRC

Cluster% of municipalities
SWCCTRC
North-180.277.8
North-287.883.6
North-378.778.7
North-473.172.1
Center-151.248.8
Center-261.363.0
Center-350.653.1
Center-442.640.7
South-182.780.0
South-281.681.1
South-377.075.2
South-479.479.4

Source(s): Authors own work

The high percentages may result from the existence of outliers that can be visualized with a 3D scatterplot in which the x and y axes represent the SWCC and TRC inputs, respectively, while the z axis represents the SWC output. In Figure 1, the North-2 cluster (geographical area north, population density 2nd quartile) is presented as an example. Some municipalities have completely anomalous values very far from the point cloud (red circle), whereas others are located in a region of space with low density compared to that of their neighbors (blue circle). Nevertheless, these points may not be considered outliers by the global approach.

Figure 1

3D scatterplot, cluster North-2

Figure 1

3D scatterplot, cluster North-2

Close Figure 1

To detect outliers, the value of k in the LOF algorithm was set to 10, that is, the minimum value to remove unwanted statistical fluctuations in the results (Breunig et al., 2000), and the LOF value corresponding to 95% of the empirical cumulative distribution was considered the threshold. Table 4 shows the number of DMUs indicated as benchmarks in each cluster that may be considered outliers, presenting anomalous values.

Table 4

Outliers as benchmarks in each cluster

ClusterBenchmark DMUsOutliers as benchmarks%
North-1171058.8
North-28787.5
North-312866.7
North-4201470.0
Center-115746.7
Center-211545.5
Center-38337.5
Center-410550.0
South-112866.7
South-212866.7
South-311763.6
South-411872.7

Source(s): Authors own work

Removing these outliers before calculating the efficiency scores allowed the municipalities within each cluster to use more reliable benchmarks to support the identification of targets. With reference to the North-2 cluster, Table 5 shows the changes in efficiency scores in the interval distribution before and after the removal of outliers.

Table 5

Interval distribution of efficiency score before and after outlier removal for cluster “North-2”

Efficiency score% of municipalities
With outlierWithout outlier
[0–0.1)0.20.0
[0.1–0.2)4.80.0
[0.2–0.3)33.61.0
[0.3–0.4)26.915.6
[0.4–0.5)15.117.1
[0.5–0.6)7.423.2
[0.6–0.7)3.619.4
[0.7–0.8)3.611.5
[0.8–0.9)1.95.4
[0.9–1)1.33.8
11.73.1

Source(s): Authors own work

As shown, the percentage of municipalities with efficiency scores higher than 0.5 was more significant after the removal of outliers, rising from 19.3% to 66.3%. This means that some of the entities considered benchmarks in the first application of the DEA, having obtained an efficiency score equal to 1, were actually outliers. As already stated, the presence of these municipalities may distort the efficiency evaluation of all the others, potentially leading to incorrect performance-based political choices.

Comparing the results before and after the removal of outliers enabled evaluating the reduction in the efficiency gap, as shown in Table 6, with reference to the North-2 cluster. After the outliers were removed, the improvement targets underwent a significant decrease, as shown in Table 6. The percentage of municipalities with input reduction targets of more than 50% fell from 87.8% to 34.5% for SWCC and from 83.6% to 33.7% for TRC. Not only were these municipalities significantly reduced in number, but their targets also became more achievable, albeit still challenging, going from 110,429 to 60,800, on average, for SWCC, and from 46,470 to 23,947 for TRC.

Table 6

Interval distribution of input reduction percentage for cluster North-2

Inputs reduction (%)% municipalities
With outlierWithout outlier
SWCCTRCSWCCTRC
01.71.73.13.1
(0–0.1)0.41.13.13.8
[0.1–0.2)1.11.15.15.4
[0.2–0.3)1.53.211.211.2
[0.3–0.4)1.72.919.619.6
[0.4–0.5)5.96.523.023.2
[0.5–0.6)12.215.517.617.1
[0.6–0.7)27.526.516.315.6
[0.7–0.8)41.035.51.00.8
[0.8–0.9)6.95.90.00.3
[0.9–1)0.20.20.00.0
10.00.00.00.0

Source(s): Authors own work

The Kolmogorov–Smirnov test was performed for all clusters to verify the significant differences in the distribution of input reduction before and after outlier removal.

As shown in Table 7, the distances (D) between the two empirical distribution functions of input reduction before and after the removal of outliers were statistically significant for the north and south clusters. This indicates that the DMUs considered outliers were “influential observations” (Wilson, 1995), the removal of which produced relevant changes in efficiency measures. By contrast, the differences for the center clusters were not significant from a statistical point of view. Indeed, the latter clusters had higher average efficiency scores (Tables 2 and 3) and fewer outliers as benchmarks (Table 4). This means that in each center cluster, the “inlier” municipalities in the reference set were considered benchmarks by the majority of the units; therefore, the distribution of efficiency scores was not greatly affected by the removal of outliers on the efficiency frontier.

Table 7

Kolmogorov-Smirnov tests

ClusterSWCCTRC
Dp-valueDp-value
North-10.210<0.00010.193<0.0001
North-20.625<0.00010.549<0.0001
North-30.241<0.00010.237<0.0001
North-40.404<0.00010.398<0.0001
Center-10.1140.4410.0790.864
Center-20.0960.7490.0610.992
Center-30.1500.4100.1600.329
Center-40.1890.5100.2230.303
South-10.277<0.00010.244<0.0001
South-20.382<0.00010.382<0.0001
South-30.1750.0220.1720.025
South-40.1470.1110.1760.031

Source(s): Authors own work

The removed outliers must be analyzed to determine whether the entities actually achieved an extraordinary performance or if it resulted from data-entry errors. Given the high cost of data checking, particularly in the case of huge amounts of data, the proposed methodology is useful in defining a prioritization for further investigation, focusing first on the outliers considered benchmarks, followed by those that have the highest LOF values. This examination provides a better understanding of the outliers without increasing costs and work, a crucial aspect, considering the often lacking ICT skills and resources needed to manage the complexity of these data in the context of public services.

As above mentioned, the removal of the anomalous values allows to improve the benchmarking and overall efficiency measurement processes, obtaining information that helps to identify more achievable targets. If the outlier analysis reveals that the anomalous values are due to measurement or transcription errors, such data will no longer be considered. Only if this analysis shows that such values are due to real virtuous cases, these extraordinary entities should be thoroughly analyzed to understand which management decisions and conditions have favored the achievement of excellent results, in order to evaluate the opportunity to import such good practices. This could support the setting of challenging medium and long-term goals that have the important advantage of not being determined subjectively, representing the arrival point of a path that probably entails improvements in organizational and management processes.

Our results highlight how this methodology can produce a significant improvement in public service PMSs. Applying DEA in combination with the LOF algorithm can improve the feasibility of using big and open data to support the internal decision-making processes of public organizations, overcoming the critical issues related to both data quality uncertainty and the need for analytical tools, yielding readily interpretable information from a managerial perspective.

As demonstrated by the results, the proposed methodology effectively addresses the problem of unrealistic targets assigned to entities that may have been erroneously considered inefficient due to the presence of outliers, as frequently encountered in real-world datasets. This problem hinders the exploitation of the opportunities offered by big and open data for the relative efficiency evaluation and benchmarking processes in the public sector, making such opportunities essentially unrealizable (Van Ooijen et al., 2019). The described procedure represents a practical method for managing data quality uncertainty and therefore taking advantage of the potential of big and open data in allowing the comparison of efficiency results among a very large number of public organizations, both at the central and local levels (Rogge et al., 2017). Further, it offers managers and local policymakers an intuitive and cheap methodology—a simple procedure even for small organizations—that allows them to meet the challenge concerning the need for reliable and easily usable tools for measurement and management purposes, given the poor diffusion of ICT capabilities in the public sector (Manyika et al., 2011). For these reasons, the proposed methodology contributes to defining good practices on how public organizations can handle big and open data for improving PMS and can be used to carry out further empirical studies within the analyzed topic that remain scarce, as the literature shows (Garengo and Sardi, 2021).

This paper proposes a methodology to improve the PMSs of entities that provide public services by exploiting the potentialities offered by big and open data, with a particular focus on the proper target setting of economic efficiency. Combining DEA with LOF, the procedure described allows for addressing some of the most relevant challenges that big and open data present when used for internal decision-making processes, thereby improving the feasibility of their use within PMSs. This paper thus aims to contribute to that strand of the literature focused on the technical and operational aspects of PMSs in the public sector (Agasisti et al., 2020), responding to the call of previous articles that highlighted the need to propose solutions to overcome technical problems related to PMSs concerning this context, with particular reference to data quality (Fryer et al., 2009).

The managerial implications of using the proposed methodology concern two different perspectives. From a micro perspective, local policymakers and managers responsible for a single entity can compare their results with those of the best-performing entities in a peer group, thus detecting the relative efficiency of their organizations. They can identify, in simple, cheap, and non-discretionary ways, the more challenging efficiency targets and, consequently, put into effect coherent new programs to achieve a more efficient allocation of resources and improve performance. This is particularly important for public services provided in a non-competitive market and can help avoid self-referentiality in target settings (Deilmann et al., 2016; McAfee et al., 2012). Furthermore, this is relevant because some kinds of public services are often provided by very small entities (municipalities or companies) characterized by a scarcity of financial and human resources to dedicate to such decision and evaluation processes. Applying this methodology, entities providing public services can go beyond the traditional perspective of observing results, which looks inward and backward, and they can adopt an outward and forward looking perspective. This offers a comparison with external contexts, suggesting objectives that could be attained in the future by adopting efficient management solutions that other outstanding entities have effectively implemented (Kouzmin et al., 1999; Magd and Curry, 2003). Furthermore, it is useful to recall that errors in setting targets and, thus, objectives can have several negative consequences in relation to programming activities and allocating scarce available resources. The pursuit of goals that are not actually achievable is counterproductive and puts excessive strain on an organization and its human resources.

From a macro perspective, central policymakers (national and supranational) can use the proposed methodology to improve decision processes regarding specific kinds of public services based on real data referring to the state of the art. The objectives and related regulations can take into account particular issues affecting the public service and related management questions. Furthermore, using benchmarking in the public sector can make programs and policies more accountable to citizens, as entities can inform, explain, and justify their performance and practices to citizens and other stakeholders, making the whole policy process more transparent and democratic, thus strengthening their legitimacy (Boyne et al., 2009).

The proposed methodology was applied in a specific context: Italian waste management services. Future research should address the analyzed topic, applying the methodology over time and in different public sector contexts to empirically assess its potential benefits in terms of supporting managerial decisions and improving public services PMSs through the exploitation of big and open data, which will be increasingly widespread and available in the future, including for public sector entities.

This work has been funded by the European Union - NextGenerationEU under the Italian Ministry of University and Research (MUR) National Innovation Ecosystem grant ECS00000041 - VITALITY - CUP D83C22000710005.

Disclaimer: Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union or the European Commission. Neither the European Union nor the European Commission can be held responsible for them.

Abuljadail
,
M.
,
Khalil
,
A.
,
Talwar
,
S.
and
Kaur
,
P.
(
2023
), “
Big data analytics and e-governance: actors, opportunities, tensions, and applications
”,
Technological Forecasting and Social Change
, Vol. 
193
, 122612, doi: .
Agasisti
,
T.
,
Agostino
,
D.
and
Soncin
,
M.
(
2020
), “
Implementing performance measurement systems in local government: moving from the ‘how’ to the ‘why’
”,
Public Performance and Management Review
, Vol. 
43
No. 
5
, pp. 
1100
-
1128
, doi: .
Agostino
,
D.
and
Arnaboldi
,
M.
(
2018
), “
Performance measurement systems in public service networks. The what, who, and how of control
”,
Financial Accountability and Management
, Vol. 
34
No. 
2
, pp. 
103
-
116
, doi: .
Alghushairy
,
O.
,
Alsini
,
R.
,
Soule
,
T.
and
Ma
,
X.
(
2021
), “
A review of local outlier factor algorithms for outlier detection in big data streams
”,
Big Data and Cognitive Computing
, Vol. 
5
No. 
1
, pp. 
1
-
24
, doi: .
Andrews
,
R.
and
Entwistle
,
T.
(
2014
),
Public Service Efficiency: Reframing the Debate
,
Routledge
,
London
.
Badia
,
F.
and
Donato
,
F.
(
2022
), “
Opportunities and risks in using big data to support management control systems: a multiple case study
”,
Management Control
, Vol. 
12
No. 
3
, pp. 
39
-
63
, doi: .
Badiezadeh
,
T.
,
Saen
,
R.F.
and
Samavati
,
T.
(
2018
), “
Assessing sustainability of supply chains by double Frontier network DEA: a big data approach
”,
Computers and Operations Research
, Vol. 
98
, pp. 
284
-
290
, doi: .
Bannister
,
F.
and
Connolly
,
R.
(
2014
), “
ICT, public values and transformative government: a framework and programme for research
”,
Government Information Quarterly
, Vol. 
31
No. 
1
, pp. 
119
-
128
, doi: .
Barbato
,
G.
and
Turri
,
M.
(
2017
), “
Understanding public performance measurement through theoretical pluralism
”,
International Journal of Public Sector Management
, Vol. 
30
No. 
1
, pp. 
15
-
30
, doi: .
Batini
,
C.
,
Cappiello
,
C.
,
Francalanci
,
C.
and
Maurino
,
A.
(
2009
), “
Methodologies for data quality assessment and improvement
”,
ACM Computing Surveys
, Vol. 
41
No. 
3
, pp. 
1
-
52
, doi: .
Bejerot
,
E.
and
Hasselbladh
,
H.
(
2013
), “
Forms of intervention in public sector organizations: generic traits in public sector reforms
”,
Organization Studies
, Vol. 
34
No. 
9
, pp. 
1357
-
1380
, doi: .
Benito
,
B.
,
Faura
,
U.
,
Guillamón
,
M.D.
and
Ríos
,
A.M.
(
2019
), “
The efficiency of public services in small municipalities: the case of drinking water supply
”,
Cities
, Vol. 
93
, pp. 
95
-
103
, doi: .
Bosch
,
N.
,
Pedraja
,
F.
and
Suárez-Pandiello
,
J.
(
2000
), “
Measuring the efficiency of Spanish municipal refuse collection services
”,
Local Government Studies
, Vol. 
26
No. 
3
, pp. 
71
-
90
, doi: .
Bouckaert
,
G.
and
Halligan
,
J.
(
2008
),
Managing Performance: International Comparisons
,
Routledge
,
New York
.
Boyd
,
D.
and
Crawford
,
K.
(
2012
), “
Critical questions for big data: provocations for a cultural, technological, and scholarly phenomenon
”,
Information, Communication and Society
, Vol. 
15
No. 
5
, pp. 
662
-
679
, doi: .
Boyne
,
G.A.
,
James
,
O.
,
John
,
P.
and
Petrovsky
,
N.
(
2009
), “
Democracy and government performance: holding incumbents accountable in English local governments
”,
The Journal of Politics
, Vol. 
71
No. 
4
, pp. 
1273
-
1284
, doi: .
Breunig
,
M.M.
,
Kriegel
,
H.P.
,
Ng
,
R.T.
and
Sander
,
J.
(
2000
), “
LOF: identifying density-based local outliers
”,
ACM SIGMOD Record
, Vol. 
29
No. 
2
, pp. 
93
-
104
, doi: .
Brignall
,
S.
and
Modell
,
S.
(
2000
), “
An institutional perspective on performance measurement and management in the ‘new’ public sector
”,
Management Accounting Research
, Vol. 
11
No. 
3
, pp. 
281
-
306
, doi: .
Broadbent
,
J.
and
Laughlin
,
R.
(
1998
), “
Resisting the ‘new public management’: absorption and absorbing groups in schools and GP practices in the UK
”,
Accounting, Auditing & Accountability Journal
, Vol. 
11
No. 
4
, pp. 
403
-
435
, doi: .
Charnes
,
A.
,
Cooper
,
W.W.
and
Rhodes
,
E.
(
1981
), “
Evaluating program and managerial efficiency: an application of data envelopment analysis to program follow-through
”,
Management Science
, Vol. 
27
No. 
6
, pp. 
668
-
697
, doi: .
Chen
,
L.
and
Jia
,
G.
(
2017
), “
Environmental efficiency analysis of China's regional industry: a data envelopment analysis (DEA) based approach
”,
Journal of Cleaner Production
, Vol. 
142
, pp. 
846
-
853
, doi: .
Chen
,
H.
,
Chiang
,
R.H.L.
and
Storey
,
V.C.
(
2012
), “
Business intelligence and analytics: from big data to big impact
”,
MIS Quarterly
, Vol. 
36
No. 
4
, pp. 
1165
-
1188
, doi: .
Chu
,
J.F.
,
Wu
,
J.
and
Song
,
M.L.
(
2018
), “
An SBM-DEA model with parallel computing design for environmental efficiency evaluation in the big data context: a transportation system application
”,
Annals of Operations Research
, Vol. 
270
No. 
1
, pp. 
105
-
124
, doi: .
Coelli
,
T.J.
,
Rao
,
D.S.P.
,
O'Donnell
,
C.J.
and
Battese
,
G.E.
(
2005
),
An Introduction to Efficiency and Productivity Analysis
,
Springer Science & Business Media
,
New York
.
Coles
,
S.
,
Bawa
,
J.
,
Trenner
,
L.
and
Dorazio
,
P.
(
2001
),
An Introduction to Statistical Modeling of Extreme Values
,
Springer
,
London
.
Cook
,
W.D.
,
Harrison
,
J.
,
Imanirad
,
R.
,
Rouse
,
P.
and
Zhu
,
J.
(
2015
), “Data envelopment analysis with non-homogeneous DMUs”, in
Zhu
,
J.
(Ed.),
Data Envelopment Analysis. International Series in Operations Research and Management Science
,
Springer
,
Boston
, pp. 
309
-
340
, doi: .
Coupet
,
J.
,
Berrett
,
J.
,
Broussard
,
P.
and
Johnson
,
B.
(
2021
), “
Nonprofit benchmarking with data envelopment analysis
”,
Nonprofit and Voluntary Sector Quarterly
, Vol. 
50
No. 
3
, pp. 
647
-
661
, doi: .
Criado
,
J.I.
and
Gil-Garcia
,
J.R.
(
2019
), “
Creating public value through smart technologies and strategies: from digital services to artificial intelligence and beyond
”,
International Journal of Public Sector Management
, Vol. 
32
No. 
5
, pp. 
438
-
450
, doi: .
De Jaeger
,
S.
,
Eyckmans
,
J.
,
Rogge
,
N.
and
Van Puyenbroeck
,
T.
(
2011
), “
Wasteful waste-reducing policies? The impact of waste reduction policy instruments on collection and processing costs of municipal solid waste
”,
Waste Management
, Vol. 
31
No. 
7
, pp. 
1429
-
1440
, doi: .
de Kool
,
D.
and
Bekkers
,
V.
(
2016
), “
The Perceived Value-relevance of open data in the parents' choice of Dutch primary schools
”,
International Journal of Public Sector Management
, Vol. 
29
No. 
3
, pp. 
271
-
287
, doi: .
De Witte
,
K.
and
Geys
,
B.
(
2011
), “
Evaluating efficient public good provision: theory and evidence from a generalised conditional efficiency model for public libraries
”,
Journal of Urban Economics
, Vol. 
69
No. 
3
, pp. 
319
-
327
, doi: .
Deilmann
,
C.
,
Lehmann
,
I.
,
Reißmann
,
D.
and
Hennersdorf
,
J.
(
2016
), “
Data envelopment analysis of cities. Investigation of the ecological and economic efficiency of cities using a benchmarking concept from production management
”,
Ecological Indicators
, Vol. 
67
, pp. 
798
-
806
, doi: .
Desouza
,
K.C.
and
Jacob
,
B.
(
2017
), “
Big data in the public sector: lessons for practitioners and scholars
”,
Administration and Society
, Vol. 
49
No. 
7
, pp. 
1043
-
1064
, doi: .
Di Vaio
,
A.
,
Hassan
,
R.
and
Alavoine
,
C.
(
2022
), “
Data intelligence and analytics: a bibliometric analysis of human - artificial intelligence in public sector decision-making effectiveness
”,
Technological Forecasting and Social Change
, Vol. 
174
, 121201, pp. 
1
-
17
, doi: .
Donthu
,
N.
,
Hershberger
,
E.K.
and
Osmonbekov
,
T.
(
2005
), “
Benchmarking marketing productivity using data envelopment analysis
”,
Journal of Business Research
, Vol. 
58
No. 
11
, pp. 
1474
-
1482
, doi: .
Dorsch
,
J.J.
and
Yasin
,
M.M.
(
1998
), “
A framework for benchmarking in the public sector: literature review and directions for future research
”,
International Journal of Public Sector Management
, Vol. 
11
Nos
2/3
, pp. 
91
-
115
, doi: .
Dyson
,
R.G.
,
Allen
,
R.
,
Camanho
,
A.S.
,
Podinovski
,
V.V.
,
Sarrico
,
C.S.
and
Shale
,
E.A.
(
2001
), “
Pitfalls and protocols in DEA
”,
European Journal of Operational Research
, Vol. 
132
No. 
2
, pp. 
245
-
259
, doi: .
Elston
,
T.
,
MacCarthaigh
,
M.
and
Verhoest
,
K.
(
2018
), “
Collaborative cost cutting: productive efficiency as an interdependency between public organizations
”,
Public Management Review
, Vol. 
20
No. 
12
, pp. 
1815
-
1835
, doi: .
European Commission
(
2022
), “
The digital economy and society index (DESI)
”,
available at:
https://digital-strategy.ec.europa.eu/en/policies/desi_on_06.01.2023
Ferreira
,
A.
and
Otley
,
D.
(
2009
), “
The design and use of performance management systems: an extended framework for analysis
”,
Management Accounting Research
, Vol. 
20
No. 
4
, pp. 
263
-
282
, doi: .
Fried
,
H.O.
,
Lovell
,
C.K.
,
Schmidt
,
S.S.
and
Yaisawarng
,
S.
(
2002
), “
Accounting for environmental effects and statistical noise in data envelopment analysis
”,
Journal of Productivity Analysis
, Vol. 
17
Nos
1/2
, pp. 
157
-
174
, doi: .
Fryer
,
K.
,
Antony
,
J.
and
Ogden
,
S.
(
2009
), “
Performance management in the public sector
”,
International Journal of Public Sector Management
, Vol. 
22
No. 
6
, pp. 
478
-
498
, doi: .
Garengo
,
P.
and
Sardi
,
A.
(
2021
), “
Performance measurement and management in the public sector: state of the art and research opportunities
”,
International Journal of Productivity and Performance Management
, Vol. 
70
No. 
7
, pp. 
1629
-
1654
, doi: .
Grundke
,
R.
,
Marcolin
,
L.
,
Nguyen
,
T.L.B.
and
Squicciarini
,
M.
(
2018
), “
Which skills for the digital era? Returns to skills analysis
”,
Technology and Industry Working Papers
, No. 
9
, doi: .
Guerrini
,
A.
,
Romano
,
G.
,
Leardini
,
C.
and
Martini
,
M.
(
2015
), “
Measuring the efficiency of wastewater services through data envelopment analysis
”,
Water Science and Technology
, Vol. 
71
No. 
12
, pp. 
1845
-
1851
, doi: .
Hitz-Gamper
,
B.S.
,
Neumann
,
O.
and
Stürmer
,
M.
(
2019
), “
Balancing control, usability and visibility of linked open government data to create public value
”,
International Journal of Public Sector Management
, Vol. 
32
No. 
5
, pp. 
451
-
466
, doi: .
Hodge
,
V.
and
Austin
,
J.
(
2004
), “
A survey of outlier detection methodologies
”,
Artificial Intelligence Review
, Vol. 
22
No. 
2
, pp. 
85
-
126
, doi: .
Hood
,
C.
(
1991
), “
A public management for all seasons?
”,
Public Administration
, Vol. 
69
No. 
1
, pp. 
3
-
19
, doi: .
Hood
,
C.
(
1995
), “
The ‘new’ public management in the 1980s: variations on the theme
”,
Accounting, Organizations and Society
, Vol. 
20
Nos
2/3
, pp. 
93
-
109
, doi: .
Humphrey
,
C.
,
Guthrie
,
J.
,
Jones
,
L.R.
and
Olson
,
O.
(
2005
), “The dynamics of public financial management change in an international context”, in
Guthrie
,
J.
,
Humphrey
,
C.
,
Jones
,
L.R.
and
Olson
,
O.
(Eds),
International Public Financial Management Reform: Progress, Contradictions, and Challenges
,
Information Age Publishing
,
Greenwich
, pp. 
1
-
22
.
Jensen
,
M.H.
,
Persson
,
J.S.
and
Nielsen
,
P.A.
(
2023
), “
Measuring benefits from big data analytics projects: an action research study
”,
Information Systems and E-Business Management
, Vol. 
21
No. 
2
, pp. 
323
-
352
, doi: .
Johnson
,
A.L.
and
McGinnis
,
L.F.
(
2008
), “
Outlier detection in two-stage semiparametric DEA models
”,
European Journal of Operational Research
, Vol. 
187
No. 
2
, pp. 
629
-
635
, doi: .
Kouzmin
,
A.
,
Löffler
,
E.
,
Klages
,
H.
and
Korac‐Kakabadse
,
N.
(
1999
), “
Benchmarking and performance measurement in public sectors: towards learning for agency effectiveness
”,
International Journal of Public Sector Management
, Vol. 
12
No. 
2
, pp. 
121
-
144
, doi: .
Kwon
,
O.
,
Lee
,
N.
and
Shin
,
B.
(
2014
), “
Data quality management, data usage experience and acquisition intention of big data analytics
”,
International Journal of Information Management
, Vol. 
34
No. 
3
, pp. 
387
-
394
, doi: .
Li
,
Z.
,
Zhao
,
Y.
,
Hu
,
X.
,
Botta
,
N.
,
Ionescu
,
C.
and
Chen
,
G.
(
2022
), “
ECOD: unsupervised outlier detection using empirical cumulative distribution functions
”,
IEEE Transactions on Knowledge and Data Engineering
, Vol. 
35
No. 
12
, pp. 
12181
-
12193
, doi: .
Lo Storto
,
C.
(
2016
), “
The trade-off between cost efficiency and public service quality: a non-parametric Frontier analysis of Italian major municipalities
”,
Cities
, Vol. 
51
, pp. 
52
-
63
, doi: .
Maciejewski
,
M.
(
2017
), “
To do more, better, faster and more cheaply: using big data in public administration
”,
International Review of Administrative Sciences
, Vol. 
83
No. 
1_suppl
, pp. 
120
-
135
, doi: .
Magd
,
H.
and
Curry
,
A.
(
2003
), “
Benchmarking: achieving best value in public‐sector organisations
”,
Benchmarking: An International Journal
, Vol. 
10
No. 
3
, pp. 
261
-
286
, doi: .
Manyika
,
J.
,
Chui
,
M.
,
Brown
,
B.
,
Bughin
,
J.
,
Dobbs
,
R.
,
Roxburgh
,
C.
and
Hung Byers
,
A.
(
2011
),
Big Data: The Next Frontier for Innovation, Competition, and Productivity
,
McKinsey Global Institute Report
,
New York
.
Markou
,
M.
and
Singh
,
S.
(
2003a
), “
Novelty detection: a review – Part 1: statistical approaches
”,
Signal Processing
, Vol. 
83
No. 
12
, pp. 
2481
-
2497
, doi: .
Markou
,
M.
and
Singh
,
S.
(
2003b
), “
Novelty detection: a review – Part 2: neural network based approaches
”,
Signal Processing
, Vol. 
83
No. 
12
, pp. 
2499
-
2521
, doi: .
McAfee
,
A.
,
Brynjolfsson
,
E.
,
Davenport
,
T.H.
,
Patil
,
D.
and
Barton
,
D.
(
2012
), “
Big data. The management revolution
”,
Harvard Business Review
, Vol. 
90
No. 
10
, pp. 
60
-
66
.
Melo
,
A.I.
and
Mota
,
L.F.
(
2020
), “
Public sector reform and the state of performance management in Portugal: is there a gap between performance measurement and its use?
”,
International Journal of Public Sector Management
, Vol. 
33
Nos
6/7
, pp. 
613
-
627
, doi: .
Mergel
,
I.
,
Rethemeyer
,
R.K.
and
Isett
,
K.
(
2016
), “
Big data in public affairs
”,
Public Administration Review
, Vol. 
76
No. 
6
, pp. 
928
-
937
, doi: .
Mwita
,
J.I.
(
2000
), “
Performance management model. A system-based approach to public service quality
”,
International Journal of Public Sector Management
, Vol. 
13
No. 
1
, pp. 
19
-
37
, doi: .
OECD
(
2015
),
Data-Driven Innovation: Big Data for Growth and Well-Being
,
OECD Publishing
,
Paris
, doi: .
OECD
(
2019
),
Going Digital: Shaping Policies, Improving Lives
,
OECD Publishing
,
Paris
, doi: .
Pawsey
,
N.
,
Ananda
,
J.
and
Hoque
,
Z.
(
2018
), “
Rationality, accounting and benchmarking water businesses: an analysis of measurement challenges
”,
International Journal of Public Sector Management
, Vol. 
31
No. 
3
, pp. 
290
-
315
, doi: .
Picazo-Vela
,
S.
,
Gutiérrez-Martínez
,
I.
and
Luna-Reyes
,
L.F.
(
2012
), “
Understanding risks, benefits, and strategic alternatives of social media applications in the public sector
”,
Government Information Quarterly
, Vol. 
29
No. 
4
, pp. 
504
-
511
, doi: .
Pokrajac
,
D.
,
Lazarevic
,
A.
and
Latecki
,
L.J.
(
2007
), “
Incremental local outlier detection for data streams
”,
2007 IEEE symposium on computational intelligence and data mining
, pp. 
504
-
515
,
available at:
https://ieeexplore.ieee.org/xpl/conhome/4221263/proceeding
Pollitt
,
C.
and
Bouckaert
,
G.
(
2011
),
Public Management Reform: A Comparative Analysis
, (3th ed.) ,
Oxford University Press
,
Oxford
.
Price
,
R.
and
Shanks
,
G.
(
2005
), “
A semiotic information quality framework: development and comparative analysis
”,
Journal of Information Technology
, Vol. 
20
No. 
2
, pp. 
88
-
102
, doi: .
Rogge
,
N.
,
Agasisti
,
T.
and
De Witte
,
K.
(
2017
), “
Big data and the measurement of public organizations' performance and efficiency: the state-of-the-art
”,
Public Policy and Administration
, Vol. 
32
No. 
4
, pp. 
263
-
281
, doi: .
Romano
,
G.
and
Molinos-Senante
,
M.
(
2020
), “
Factors affecting eco-efficiency of municipal waste services in Tuscan municipalities: an empirical investigation of different management models
”,
Waste Management
, Vol. 
105
, pp. 
384
-
394
, doi: .
Roth
,
J.
,
Lim
,
B.
,
Jain
,
R.K.
and
Grueneich
,
D.
(
2020
), “
Examining the feasibility of using open data to benchmark building energy usage in cities: a data science and policy perspective
”,
Energy Policy
, Vol. 
139
, 111327, doi: .
Ruijer
,
E.
,
Porumbescu
,
G.
,
Porter
,
R.
and
Piotrowski
,
S.
(
2023
), “
Social equity in the data era: a systematic literature review of data-driven public service research
”,
Public Administration Review
, Vol. 
83
No. 
2
, pp. 
316
-
332
, doi: .
Sadiq
,
S.
and
Indulska
,
M.
(
2017
), “
Open data: quality over quantity
”,
International Journal of Information Management
, Vol. 
37
No. 
3
, pp. 
150
-
154
, doi: .
Sarra
,
A.
,
Mazzocchitti
,
M.
and
Rapposelli
,
A.
(
2017
), “
Evaluating joint environmental and cost performance in municipal waste management systems through data envelopment analysis: scale effects and policy implications
”,
Ecological Indicators
, Vol. 
73
, pp. 
756
-
771
, doi: .
Schmidthuber
,
L.
,
Stütz
,
S.
and
Hilgers
,
D.
(
2019
), “
Outcomes of open government: does an online platform improve citizens' perception of local government?
”,
International Journal of Public Sector Management
, Vol. 
32
No. 
5
, pp. 
489
-
507
, doi: .
Smiti
,
A.
(
2020
), “
A critical overview of outlier detection methods
”,
Computer Science Review
, Vol. 
38
, 100306, doi: .
Song
,
M.
,
Du
,
Q.
and
Zhu
,
Q.
(
2017
), “
A theoretical method of environmental performance evaluation in the context of big data
”,
Production Planning and Control
, Vol. 
28
Nos
11-12
, pp. 
976
-
984
, doi: .
Speklé
,
R.F.
and
Verbeeten
,
F.H.M.
(
2014
), “
The use of performance measurement systems in the public sector: effects on performance
”,
Management Accounting Research
, Vol. 
25
No. 
2
, pp. 
131
-
146
, doi: .
Steccolini
,
I.
,
Saliterer
,
I.
and
Guthrie
,
J.
(
2020
), “
The role(s) of accounting and performance measurement systems in contemporary public administration
”,
Public Administration
, Vol. 
98
No. 
1
, pp. 
3
-
13
, doi: .
Twizeyimana
,
J.D.
and
Andersson
,
A.
(
2019
), “
The public value of e-government – a literature review
”,
Government Information Quarterly
, Vol. 
36
No. 
2
, pp. 
167
-
178
, doi: .
Valle-Cruz
,
D.
(
2019
), “
Public value of e-government services through emerging technologies
”,
International Journal of Public Sector Management
, Vol. 
32
No. 
5
, pp. 
530
-
545
, doi: .
Van Ooijen
,
C.
,
Ubaldi
,
B.
and
Welby
,
B.
(
2019
), “
A data-driven public sector: enabling the strategic use of data for productive, inclusive and trustworthy governance
”,
OECD Working Papers on Public Governance
, No. 
33
, doi: .
Weerakkody
,
V.
,
Kapoor
,
K.
,
Balta
,
M.E.
,
Irani
,
Z.
and
Dwivedi
,
Y.K.
(
2017
), “
Factors influencing user acceptance of public sector big open data
”,
Production, Planning and Control
, Vol. 
28
Nos
11-12
, pp. 
891
-
905
, doi: .
Wilson
,
P.W.
(
1995
), “
Detecting influential observations in data envelopment analysis
”,
Journal of Productivity Analysis
, Vol. 
6
No. 
1
, pp. 
27
-
45
, doi: .
Worthington
,
A.C.
(
2000
), “
Cost efficiency in Australian local government: a comparative analysis of mathematical programming and econometrical approaches
”,
Financial Accountability and Management
, Vol. 
16
No. 
3
, pp. 
201
-
223
, doi: .
Worthington
,
A.C.
and
Dollery
,
B.E.
(
2001
), “
Measuring efficiency in local government: an analysis of New South Wales municipalities' domestic waste management function
”,
Policy Studies Journal
, Vol. 
29
No. 
2
, pp. 
232
-
249
, doi: .
Zhu
,
J.
(
2000
), “
Multi-factor performance measure model with an application to Fortune 500 companies
”,
European Journal of Operational Research
, Vol. 
23
No. 
1
, pp. 
105
-
124
, doi: .
Zhu
,
J.
(
2022
), “
DEA under big data: data enabled analytics and network data envelopment analysis
”,
Annals of Operations Research
, Vol. 
309
No. 
2
, pp. 
761
-
783
, doi: .
Zhu
,
Q.
,
Wu
,
J.
,
Li
,
X.
and
Xiong
,
B.
(
2017
), “
China's regional natural resource allocation and utilization: a DEA-based approach in a big data environment
”,
Journal of Cleaner Production
, Vol. 
142
No. 
part 2
, pp. 
809
-
818
, doi: .
Zuiderwijk
,
A.
and
Janssen
,
M.
(
2014
), “
Open data policies, their implementation and impact: a framework for comparison
”,
Government Information Quarterly
, Vol. 
31
No. 
1
, pp. 
17
-
29
, doi: .
Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at http://creativecommons.org/licences/by/4.0/legalcode

or Create an Account

Close subscription notice
Close access options