Purpose

This study develops an artificial intelligence (AI)-driven framework for using passenger-generated social media discourse as a complementary source of operational intelligence in railway service quality management. It examines whether Reddit discussions can reveal dimensions of passenger experience that are not adequately captured by official performance indicators.

Design/methodology/approach

Using the official Reddit API, the study collected 61,328 documents published between July 2024 and July 2026 across five railway communities. The framework combines lexicon-based classification into seven theory-informed service failure categories with non-negative matrix factorization topic modelling, VADER sentiment analysis, lexicon-based emotion detection and robust regression.

Findings

Delay-and-disruption discourse showed the highest proportion of negative documents, while safety discourse generated the greatest negativity intensity and visibility; comfort and cleanliness and ticketing and refunds were the most frequently discussed dimensions. Affective intensity and public engagement proved partially independent: delay and ticketing posts attracted significantly less engagement than average, safety, comfort posts significantly more. The index ranked comfort, cleanliness, staff, communication, ticketing and refunds as leading improvement priorities, with ticketing, delay dominating British, comfort dominating US discourse.

Research limitations/implications

Findings rest on self-selected Reddit communities and rule-based instruments validated internally rather than against human gold labels; they indicate digitally expressed passenger concerns, not representative estimates of the passenger population.

Practical implications

Operators and regulators can use the framework as a low-cost, continuously updated instrument complementing official performance indicators. Its extension into an early-warning system is set out as a design proposition requiring prospective validation against operational incident data, not as a capability demonstrated here.

Social implications

This study highlights social media as a democratic space for public accountability in transit management. By amplifying organic passenger discourse, the framework empowers commuters who's daily lived experiences, such as comfort, cleanliness and communication issues, are often overlooked by rigid official KPIs. Recognizing these grievances underscores how transit quality directly impacts public trust, equity and socio-economic well-being. Furthermore, addressing regional variations in passenger concerns helps operators design more inclusive services. Ultimately, integrating this digital feedback into policy ensures that vital public infrastructure better aligns with the diverse safety, accessibility and comfort needs of the communities it serves.

Originality/value

The study demonstrates how large-scale passenger discourse can be transformed into actionable railway service intelligence by integrating theory-informed classification, affective analysis, engagement modelling and a multidimensional operational priority measure.

Railway systems are evaluated, planned and regulated primarily through operational key performance indicators (KPIs), such as punctuality rates, delay minutes, cancellation counts and capacity utilization. While these indicators are indispensable for operational control and regulatory reporting, they capture the occurrence of service disruptions rather than their perception. Two delays of identical duration may produce very different passenger reactions depending on communication quality, crowding conditions, rebooking options and the availability of timely information. A cancelled early-morning commuter service with clear rebooking guidance may use less experiential damage than a 30-minute delay endured on an overcrowded platform without announcements. Consequently, a persistent gap exists between what operators measure and what passengers experience, a gap that service quality research has long emphasized but that operational monitoring systems have rarely closed.

Traditional instruments for measuring passenger experience, including periodic satisfaction surveys, complaint forms and mystery-shopper audits, suffer from well-documented limitations: low and declining response rates, sampling and self-selection bias, delayed availability, high administration costs and limited granularity with respect to specific journeys, routes and failure episodes (de Oña & de Oña, 2015). Cross-national survey programs additionally struggle with comparability across instruments and cultural response styles (Fellesson & Friman, 2008). By the time a quarterly satisfaction wave reports a decline in perceived reliability, the underlying operational episode, a timetable recast, an industrial dispute, a rolling-stock shortage may be months in the past, and the survey instrument rarely reveals which concrete aspects of the episode drove the deterioration.

In contrast, passengers increasingly narrate their travel experiences voluntarily and in near real time on online platforms. The transportation literature recognized early that this unsolicited discourse constitutes a policy- and operation-relevant data stream (Gal-Tzur et al., 2014; Rashidi, Abbasi, Maghrebi, Hasan, & Waller, 2017). Discussion communities such as Reddit host rich accounts of delays, cancellations, overcrowding, ticketing failures, staff interactions and comfort issues across multiple national railway systems. Compared with the microblogs that dominate existing transport studies, Reddit offers longer and more detailed posts, threaded discussions that preserve conversational context, topic-focused communities organized around specific rail systems and operators and voting signals that proxy community salience (Medvedev, Lambiotte, & Delvenne, 2019; Proferes, Jones, Gilbert, Fiesler, & Zimmer, 2021). These properties make Reddit particularly suited to experience research, yet it remains strikingly underused in transportation analytics.

Advances in natural language processing (NLP) make it feasible to convert large volumes of such unstructured discourse into structured, decision-relevant signals (Grimmer & Stewart, 2013). However, existing applications of social media analytics in transportation have predominantly focused on urban transit sentiment measurement (Casas & Delmelle, 2017; Collins, Hasan, & Ukkusuri, 2013), incident detection or generic satisfaction tracking, with railway-specific applications remaining scarce (Mogaji & Erkan, 2019). More fundamentally, most studies stop at descriptive sentiment reporting: they establish that passengers are unhappy about something, but they do not translate the discourse into instruments that operations managers can act upon – a prioritization of failure categories, an early-warning trigger, a segmentation by system and passenger type.

This study addresses that gap. Its central argument is that passenger-generated online discourse can serve as a low-cost, AI-assisted monitoring source that detects service failures invisible to official statistics, tracks shifts in passenger sentiment and identifies operational improvement priorities. Rather than treating passenger satisfaction solely as a consumer-behavior phenomenon, we frame it as an input to a decision-support system for railway operations management. Because the corpus spans multiple rail systems (the USA, the United Kingdom and Continental Europe, with smaller Asia-Pacific coverage) and multiple passenger types (daily commuters, occasional travelers and tourists, rail enthusiasts and industry professionals), the framework also enables cross-system and cross-segment comparison of failure salience, a comparative capability that no single national survey program provides.

The study is guided by four research questions:

RQ1.

What are the principal service failure themes that constitute railway passenger experience in Reddit discussions and does the inductively derived topic structure corroborate a theory-informed classification of failure clusters?

RQ2.

How do service dimensions such as delays and disruptions, crowding, ticketing and refunds, comfort and cleanliness, staff communication, safety and accessibility relate to passenger sentiment and emotion?

RQ3.

Which service failures generate the highest visibility, comment volume and engagement, and how does this vary across rail systems and passenger types?

RQ4.

Can an operational prioritization index be derived from passenger discourse to support railway service quality management?

The contribution of the paper is fourfold. Firstly, it provides one of the first large-scale, multi-community, multi-system NLP analyses of railway passenger discourse on Reddit, based on 61,328 documents spanning two years and three major rail regions. Secondly, it introduces a hybrid deductive–inductive design in which a theory-informed lexicon of service failure clusters, grounded in the transit service quality literature, is compared with unsupervised topic modeling, combining the auditability required for operational adoption with the discovery capacity of inductive methods. Thirdly, it links service failure themes to affective outcomes through dimension-targeted sentiment and emotion analysis, moving beyond single overall polarity scores and revealing that affective intensity and public salience are partially independent dimensions of failure severity. Fourthly, it introduces an Operational Priority Index (OPI) that translates unstructured discourse into an actionable prioritization tool, thereby bridging social media analytics and railway operations management.

The remainder of the paper is organized as follows. Section 2 reviews the relevant literature on railway service quality, service failure theory and NLP-based social media analytics in transportation. Section 3 describes the data and methodology. Section 4 presents the results. Section 5 discusses managerial and theoretical implications, and Section 6 concludes with limitations and directions for future research.

Service quality in public transport has traditionally been conceptualized through multi-dimensional frameworks derived from the SERVQUAL tradition. Parasuraman, Zeithaml, and Berry (1985) modelled perceived quality as the gap between expectations and performance across a set of generic service dimensions, later operationalized in the SERVQUAL instrument (Parasuraman, Zeithaml, & Berry, 1988). The underlying expectation-disconfirmation logic (Oliver, 1980) implies that identical objective performance can yield different satisfaction outcomes depending on what passengers expected a mechanism directly relevant to railway disruptions, where expectations are shaped by timetables, fare levels and prior experience.

Transit research has adapted these frameworks to the specific attribute structure of public transport. Eboli and Mazzulla (2007) identified reliability, frequency, information, cleanliness and personnel behavior as core satisfaction drivers for bus transit; dell’Olio, Ibeas, and Cecin (2011) showed that desired quality varies systematically across user segments, with waiting time, cleanliness and comfort weighted differently by commuters, students and occasional travelers a segmentation logic that this study extends to discourse data. Across public-transport contexts, reliability-related attributes such as punctuality, cancellations and connection integrity consistently emerge as major drivers of satisfaction (de Oña & de Oña, 2015; Eboli & Mazzulla, 2007; Fellesson & Friman, 2008), while comfort and information quality shape perceived service quality. Crowding occupies a special position: Tirachini, Hensher, and Rose (2013) document that crowding imposes measurable disutility on passengers, affects dwell times and operations and distorts demand estimates when ignored, making it both an experience variable and an operational one. Methodologically, de Oña and de Oña (2015) review the large survey-based literature and note its dependence on costly, infrequent data collection, motivating the search for continuous, passively generated indicators. Accessibility is retained here as a distinct operational dimension because failures in booked assistance and step-free access constitute observable service breakdowns in passenger discourse.

Service-failure research emphasizes the diagnostic importance of critical incidents and recovery episodes. Bitner, Booms, and Tetreault (1990) showed how favourable and unfavorable service encounters are organized around memorable incidents; Smith, Bolton, and Wagner (1999) linked failure type, magnitude and recovery to post-failure satisfaction; and Tax, Brown, and Chandrashekaran (1998) showed that complaint-handling evaluations shape satisfaction, trust and commitment. Two implications follow for railway monitoring. First, monitoring systems should focus specifically on failure discourse rather than average satisfaction, because failures can exert substantial influence on post-encounter evaluations. Secondly, the recovery dimension, compensation processes, communication during disruption and staff responsiveness deserve separate measurement because recovery quality can materially alter the evaluation of the underlying failure. The seven failure clusters operationalized in this study, and in particular the separation of staff-and-communication and ticketing-and-refund (compensation) discourse from the underlying delay events, are drawn directly from this literature.

Social media analytics entered transportation research through two routes: incident and disruption detection from real-time streams and satisfaction measurement from opinionated text. Collins et al. (2013) provided an early demonstration that transit rider sentiment can be measured at scale from social media proposing a rider satisfaction metric derived from message polarity. Casas and Delmelle (2017) analysed transit-related microblog content to characterize public perceptions of a transit system, and Gal-Tzur et al. (2014) mapped the broader potential of social media for transport policy goals, from service planning to customer relations. Kuflik et al. (2017) moved the field toward automation, proposing a framework for harvesting and analysing transport-related social media content while cataloguing its practical challenges: noise, relevance filtering and the gap between raw text and transport-meaningful categories. Rashidi et al. (2017) offered a systematic assessment of the capacity of social media data for modelling travel behaviour, concluding that such data are complementary to, not substitutes for, traditional sources, a conclusion this study echoes in the KPI-complementarity framing. In the railway domain specifically, Mogaji and Erkan (2019) analysed approximately 1.9 million tweets about UK train services and reported an overall positive customer-experience profile alongside meaningful variation in service-quality dimensions across train groups. Their study demonstrated the diagnostic value of social-media discourse but remained primarily descriptive rather than translating the signals into an operational prioritization tool.

Much of the cited transport social-media work relies on Twitter/X (Casas & Delmelle, 2017; Collins et al., 2013; Mogaji & Erkan, 2019). Reddit has remained peripheral despite structural advantages for experience research documented in the platform-studies literature: threaded, context-preserving discussions; communities organized by topic rather than by social graph; longer texts; pseudonymity; and voting mechanisms that surface community-salient content (Medvedev et al., 2019). Proferes et al. (2021), reviewing 727 Reddit studies, show that the platform supports a wide range of research designs but also highlight recurring ethical obligations around user expectations and anonymization. Baumgartner, Zannettou, Keegan, Squire, and Blackburn (2020) describe the data infrastructure that made large-scale Reddit research feasible. At the same time, the critical data-studies literature cautions against treating any platform corpus as a population sample: boyd and Crawford (2012) emphasize that big data are always partial, platform-shaped and demographically skewed, cautions that Section 6 operationalizes as explicit limitations.

The methodological foundation for this study is the text-as-data paradigm: automated content analysis can scale where human coding cannot, but every automated measurement is a model whose validity must be demonstrated rather than assumed (Grimmer & Stewart, 2013). Topic models are a central tool for inductive theme discovery. Latent Dirichlet allocation (Blei, Ng, & Jordan, 2003) represents documents as mixtures over latent word distributions; non-negative matrix factorization (Lee & Seung, 1999) provides a parts-based decomposition that can be applied to TF-IDF representations and yields directly inspectable term-weighted components. Transformer-based representations (Devlin, Chang, Lee, & Toutanova, 2019) enable contextual embedding approaches; BERTopic (Grootendorst, 2022) combines embedding-based clustering with class-based TF-IDF to represent topics.

Sentiment analysis has developed from general opinion-mining foundations (Liu, 2012; Pang & Lee, 2008) toward instruments tuned to specific text genres. VADER (Hutto & Gilbert, 2014) is a rule-based model constructed and validated specifically for social media language, handling intensifiers, negation, punctuation emphasis and slang. Here it is used as a transparent baseline where labeled training data are unavailable. Liu (2012) catalogues persistent hard cases, sarcasm, mixed valence within a document and comparative opinions that motivate two design choices adopted here: dimension-targeted aggregation, so that a document praising food while condemning delays contributes to both dimension profiles and the use of continuous scores rather than hard labels in inferential models. Emotion analysis extends polarity to discrete affective states; the crowdsourced word–emotion association lexicon of Mohammad and Turney (2013) established the feasibility of transparent, lexicon-based emotion measurement at scale. In the present framework, anger, anxiety and satisfaction are treated as distinct managerial signals: anger marks complaint intensity, anxiety flags risk-oriented concern and satisfaction identifies positively evaluated service elements.

Three gaps emerge from this review. First, railway-specific applications of NLP-based social-media analysis remain limited relative to the broader urban-transit literature and cross-system comparative designs are uncommon (Casas & Delmelle, 2017; Collins et al., 2013; Kuflik et al., 2017; Mogaji & Erkan, 2019; Rashidi et al., 2017). Secondly, much prior work emphasizes detection, sentiment measurement or descriptive service-quality analysis rather than translating discourse into operational prioritization tools (Collins et al., 2013; Mogaji & Erkan, 2019). Thirdly, engagement signals such as visibility and comment volume are not central feature to these transport studies, even though the service-failure literature suggests that evaluative severity and public response are conceptually distinct. This study addresses all three gaps through an integrated pipeline culminating in an Operational Priority Index.

Data were collected through the official Reddit API using the PRAW library in read-only mode. The corpus covers a two-year window from 10 July 2024 to 10 July 2026, a period encompassing post-pandemic ridership recovery, repeated industrial action in several European systems and continued fare and franchise reform debates in the United Kingdom. Five primary railway-experience communities were sampled (Table 1), selected to cover general railway discourse and the major English-language passenger systems: a large North American intercity community (r/Amtrak), the principal British rail community (r/uktrains), the main European leisure rail community (r/Interrail), a general public transport community (r/transit) and the general railway community (r/trains). Together, these communities span the three rail regions compared in Section 4.6 while remaining tractable for full comment-tree retrieval.

Table 1

Sampled subreddits, analytical purpose and collected volume

SubredditAnalytical purposePostsComments
r/AmtrakUS intercity passenger rail: delays, ticketing, comfort, staff64919,526
r/uktrainsUK rail system: cancellations, delays, capacity, fares, service quality60916,517
r/InterrailEuropean rail travel: connections, reservations, planning, ticketing5574,596
r/transitPublic transport and rail transit discussions21912,676
r/trainsGeneral railway experience, cross-system comparisons1395,840
Source(s): Authors’ own work

Two complementary retrieval strategies were employed. First, each subreddit was traversed through its new and top listing endpoints to a depth of up to 500 submissions per endpoint. The two endpoints have complementary biases: the new listing is recency-oriented, while the top listing is popularity-oriented, so their union mitigates the sampling distortion of either ranking alone. Secondly, 12 targeted search queries covering delay and disruption experiences, crowding, ticketing and refunds, comfort, staff communication, safety and general evaluative discourse (e.g. “worst train journey”, “train ticket refund”, “overcrowded train”) were executed within the sampled communities to recover older or lower-ranked experience threads that the listings miss. For every retained submission with at least three comments, the comment tree was retrieved to a depth of up to 300 comments, preserving the conversational context that distinguishes Reddit from microblog data. Records were deduplicated on post and comment identifiers across both strategies.

Relevance filtering followed the two-stage logic recommended for transport social media harvesting (Kuflik et al., 2017). Submissions were retained if their title or body matched at least one of the seven service failure keyword clusters described in Section 3.3; search-retrieved submissions were additionally required to contain an explicit railway reference term (e.g. train, railway, station, platform or a named operator), preventing contamination from non-rail service complaints. URLs and markup were stripped, and whitespace normalized. Bot and moderator accounts, deleted and removed content, exact duplicates and documents shorter than five words were excluded in a subsequent cleaning pass. The final analytical corpus comprises 61,328 documents: 2,173 posts and 59,155 comments, with a mean length of 48.2 words (median 28) – several times the length of the tweets analysed in most prior transport sentiment studies and long enough to carry multi-dimensional experiential detail.

The unit of analysis is the individual document (post or comment). Each document carries (1) platform metadata, (2) a theory-informed classification into seven service failure clusters implemented as transparent keyword lexicons and (3) automatically extracted contextual fields. Comments additionally inherit the metadata of their parent post, enabling thread-level visibility measures. Table 2 summarizes the variables as they appear in the dataset.

Table 2

Variables and operationalization

Variable (dataset field)Measurement
Service failure clusters (theme_clusters)Lexicon match to seven clusters: Delay & Disruption; Crowding & Capacity; Ticketing & Refunds; Comfort & Cleanliness; Staff & Communication; Safety & Security; Accessibility (multi-label)
Matched terms (…_kw_matched)Exact lexicon terms matched per cluster (audit trail for transparency)
Rail system context (rail_context)USA, UK, Continental Europe, Asia-Pacific; high-speed, commuter/urban, sleeper/night, intercity (multi-label)
Operator mentions (operators_mentioned)Named operators (e.g. Amtrak, ÖBB, SNCF, Eurostar, LNER) via alias dictionaries
Service dimensions (service_dimensions)Nine target dimensions for dimension-level sentiment: reliability/punctuality; crowding; ticketing/fares; comfort/cleanliness; food & catering; staff & communication; safety & security; accessibility; information systems
Sentiment valenceVADER compound score in [−1, 1]; negative-token proportion; categorical label (positive ≥0.05, negative ≤ −0.05, neutral otherwise)
Emotion cuesTransparent lexicons for anger, anxiety, and satisfaction (binary per document)
Delay duration mentionsRegex-extracted delay/waiting durations (e.g. “delayed by 45 min”)
Cost mentionsRegex-extracted fares and amounts ($, £, €)
Passenger type (user_type)Self-disclosed role: rail professional; daily commuter; rail enthusiast; tourist/occasional; skeptic/critic
Visibilitylog(1 + parent post score) + log(1 + parent post comment count)
Operational Priority IndexFrequency × negativity × visibility, each max-normalized (Section 3.4)
Source(s): Authors’ own work

The seven clusters were derived deductively from the transit service quality literature reviewed in Section 2.1: reliability and information-related failures (de Oña & de Oña, 2015; Eboli & Mazzulla, 2007), crowding (Tirachini et al., 2013), the tangibles and personnel/information dimensions (Parasuraman et al., 1988), safety (Fellesson & Friman, 2008) and ticketing/refund problems, whose compensation and recovery component is informed by service-recovery research (Smith et al., 1999; Tax et al., 1998). Accessibility is retained as an operationally relevant extension because assistance and step-free-access failures are directly observable in passenger discourse. Each cluster is operationalized as a curated list of words and phrases observed in railway discourse (e.g. “signal failure”, “rail replacement”, “delay repay”, “standing room only”, “booked assistance”) and every match is stored verbatim in the dataset, giving each automated classification a complete audit trail a transparency property that black-box classifiers cannot offer and that matters for managerial trust in AI-derived indicators. The complete cluster lexicons, the emotion lexicons, the operator alias dictionary and the regular expressions used for delay and cost extraction are reproduced in full in the online supplementary material, so that the classification layer can be inspected and replicated independently.

The keyword lexicons serve two purposes. Methodologically, they provide a deductive, fully transparent classification whose matched terms are stored per document, allowing complete auditability. Analytically, they act as weak labels against which the inductive topic solution is compared (RQ1), following the validation imperative of the text-as-data literature (Grimmer & Stewart, 2013). Overall, 23,471 documents (38.3% of the corpus) match at least one failure cluster and 6,453 (10.5%) match two or more, confirming both that failure discourse is a large minority of community talk and that multi-label treatment is warranted: passengers routinely narrate delays, crowding and staff communication within a single account of a disrupted journey.

The analysis proceeds in eight steps:

  1. Collection of Reddit posts and comment trees via the dual retrieval strategy (Section 3.1).

  2. Relevance filtering, deduplication and cleaning (bots, deleted content, very short texts).

  3. Deductive classification of each document into the seven service failure clusters via transparent keyword lexicons, with matched terms retained as an audit trail.

  4. Inductive extraction of themes with NMF (K = 10, selected by inspecting solutions over K = 5–20 and retaining the smallest K at which every deductive cluster had at least one interpretable counterpart without topic duplication; the selection criterion was therefore interpretability and non-redundancy rather than a quantitative coherence optimum, a judgement-based choice that is stated here explicitly because it is not reducible to a single fit statistic and whose consequences are bounded by the fact that the deductive layer, not the inductive one, carries the primary analysis) on TF-IDF representations (unigrams and bigrams, minimum document frequency 15, maximum document proportion 0.35) of the failure-relevant subcorpus (n = 23,471); LDA with identical K as robustness benchmark; correspondence between inductive topics and deductive clusters assessed via normalized mutual information (NMI) on single-cluster documents.

  5. Sentiment analysis with VADER (Hutto & Gilbert, 2014), yielding a compound valence score, a negative-token proportion and a categorical label per document; dimension-level profiles obtained by aggregating over documents tagged with each of the nine service dimensions.

  6. Emotion detection via transparent lexicons distinguishing anger, anxiety and satisfaction, following the lexicon-based tradition of Mohammad and Turney (2013).

  7. Regression modeling of which failure clusters predict negative sentiment and engagement, with heteroskedasticity-robust (HC1) standard errors, controlling for record type (post vs. comment), subreddit, mentioned delay duration, text length and year effects.

  8. Construction of the Operational Priority Matrix (frequency × negativity × visibility), including an explicit sensitivity analysis of its component weightings and derivation of a prioritization framework, with heterogeneity analysis across rail systems, service types and passenger types. A design for early-warning alerting is proposed on the basis of this architecture but is not empirically tested here.

Formally, the Operational Priority Index for cluster k is defined as OPIk = Fk × Nk × Vk, where F denotes cluster document frequency, N the mean negative-token proportion of documents in the cluster and V the mean visibility (log-scaled parent post scores and comment counts). Each component is normalized by its maximum across clusters, so that every component lies in (0, 1] and the index preserves multiplicative structure without degenerate zeros. The multiplicative form embodies a deliberate managerial logic: a failure category earns high priority only if it is simultaneously frequent, affectively damaging and publicly salient; a category that is extreme on one component but marginal on the others is flagged for targeted attention (Section 5.2) rather than headline priority. Because documents may belong to multiple clusters, frequencies are computed on multi-label counts; robustness to exclusive assignment, to an additive aggregation and to replacing N with the share of negative documents is reported in Section 4.5.

Following the injunction that automated text measurements must be validated rather than assumed (Grimmer & Stewart, 2013), three checks address the reliability of the pipeline. First, as an internal consistency check, not a test of external validity, since both instruments were constructed by the present authors, VADER labels were compared with a hand-curated cue lexicon (positive and negative phrase lists constructed at collection time, disjoint from the VADER dictionary) on the 5,052 documents carrying a single-polarity cue: raw agreement is 64.3% and Cohen's κ on non-neutral VADER labels (n = 4,793) is 0.36, which indicates limited agreement. This level is consistent with the known difficulty of sarcasm-rich forum text (Liu, 2012), but it is low enough that categorical polarity labels cannot be treated as reliable document-level measurements; all inferential models therefore use the continuous negative-token proportion and categorical shares are reported descriptively only. Secondly, the NMF topic solution was benchmarked against LDA with identical K (ARI = 0.17; NMI = 0.21), indicating partially overlapping but method-dependent partitions and against the deductive clusters (NMI = 0.23 on 17,018 single-cluster documents). Thirdly, all OPI rank orderings were subjected to the robustness battery described above; this battery establishes internal consistency across specification choices and should not be read as external validation of the index. No human gold-standard annotation was available. This is the principal validity limitation of the pipeline (Section 6) and its first extension priority: a stratified sample of documents double-coded for cluster membership and polarity, reported with Krippendorff's α, would place both the lexicon layer and the sentiment layer on an externally anchored footing.

Only publicly available content was collected in read-only mode through the official API, in accordance with Reddit's terms of service. The study follows established ethical guidance for social media research (Franzke et al., 2020; Townsend & Wallace, 2016) and the Reddit-specific recommendations of Proferes et al. (2021): usernames were removed before analysis and replaced with anonymous identifiers; no attempt was made to identify individuals; analysis is reported only in aggregate; and quoted excerpts are paraphrased where they could enable re-identification through search. Because the study analysed pre-existing, publicly available and pseudonymous content and involved no interaction or intervention with human participants, it did not fall within the remit of the institutional human-subjects review process.

Data and code availability. Reddit's terms of service preclude redistribution of post and comment text. The analysis code, the complete keyword and emotion lexicons and the list of document identifiers required to reconstruct the corpus through the official API will therefore be deposited in a public repository upon acceptance, allowing the pipeline to be re-executed end to end. No component of the analysis depends on proprietary software or on data unavailable to other researchers through the same API.

The corpus comprises 61,328 documents. Discussion is sustained across the entire two-year window (Figure 1), confirming the feasibility of continuous monitoring; the rising volume toward the end of the window partly reflects the recency-oriented new listing endpoint rather than organic growth in discourse and should not be read as a trend estimate. r/Amtrak contributes the largest share (20,175 documents), followed by r/uktrains (17,126), r/transit (12,895), r/trains (5,979) and r/Interrail (5,153). Automatic context tagging locates 5,221 documents in a US rail context, 2,474 in Continental Europe, 1,768 in the UK and 264 in Asia-Pacific; by service type, commuter/urban (3,495), high-speed (2,007), sleeper/night (1,456) and intercity/long-distance (1,254) discourse are all substantially represented. Context tagging is deliberately conservative: a document is assigned to a rail system only when it contains an explicit system, country or operator reference, so the great majority of documents in particular short replies within threads whose system is established by the parent post remain untagged and the tagged subsets (9,727 documents in total) are not a random sample of the corpus. The cross-system comparison in Section 4.6 rests on these subsets and is reported as a comparison of explicitly system-referencing discourse rather than of the full national sub-corpora. Subreddit membership offers an alternative and less demanding proxy that would recover the full corpus, but it attributes every document in a community to a single system, and because r/transit and r/trains are system-heterogeneous, together contributing 18,874 documents, that substitution would introduce misattribution on a scale exceeding the coverage it gains. The conservative rule is therefore retained, and the resulting loss of coverage is treated as a limitation rather than corrected by a weaker instrument. The service-type coverage nonetheless indicates that the corpus spans the full spectrum from daily commuting to long-distance leisure travel. Amtrak is by far the most frequently named operator (4,904 mentions), followed by ÖBB (537), SNCF (406), Eurostar (370), LNER (344), GWR (339) and Avanti West Coast (295); the prominence of ÖBB reflects the growth of night-train discourse in the study period. Self-disclosed passenger roles appear in 705 documents and are dominated by tourists and occasional travelers (490), followed by rail enthusiasts (123), skeptics (56), daily commuters (19) and rail professionals (17); the rarity of self-disclosure is expected in pseudonymous communities and cautions against strong inference from the segment analysis in Section 4.6.

Figure 1
A line graph showing the monthly volume of railway passenger discourse from July 2024 to July 2026.A line graph depicting the monthly volume of railway passenger discourse from July 2024 to July 2026. The horizontal axis represents the timeline from July 2024 to July 2026, while the vertical axis represents the number of documents per month, ranging from 0 to 10,000. The graph shows a generally low and stable volume of documents from July 2024 to early 2025, with a noticeable increase starting around mid-2025. There is a significant spike in the volume of documents around April 2026, reaching its peak, followed by a sharp decline towards July 2026.

Monthly volume of railway passenger discourse, July 2024 – July 2026. Source: Authors’ own work

Figure 1
A line graph showing the monthly volume of railway passenger discourse from July 2024 to July 2026.A line graph depicting the monthly volume of railway passenger discourse from July 2024 to July 2026. The horizontal axis represents the timeline from July 2024 to July 2026, while the vertical axis represents the number of documents per month, ranging from 0 to 10,000. The graph shows a generally low and stable volume of documents from July 2024 to early 2025, with a noticeable increase starting around mid-2025. There is a significant spike in the volume of documents around April 2026, reaching its peak, followed by a sharp decline towards July 2026.

Monthly volume of railway passenger discourse, July 2024 – July 2026. Source: Authors’ own work

Close Figure 1

Comfort & Cleanliness is the most prevalent failure cluster (9,147 documents; 14.9% of the corpus), followed by Ticketing & Refunds (8,476; 13.8%), Staff & Communication (5,345; 8.7%), Delay & Disruption (3,991; 6.5%), Safety & Security (2,137; 3.5%), Crowding & Capacity (1,694; 2.8%) and Accessibility (962; 1.6%) (Figure 2; Table 3). This frequency ordering itself is informative: the dimensions that dominate official KPI systems punctuality and cancellations are not the dimensions passengers write about most. The experiential surface of railway travel, as passengers narrate it, is dominated by the tangible and transactional layers of the journey.

Figure 2
A bar graph showing the prevalence of service failure clusters and the share of negative documents.The bar graph compares the prevalence of various service failure clusters and the share of negative documents. It features seven vertical bars, each representing a different cluster: Comfort & Cleanliness, Ticketing & Refunds, Staff & Communication, Delay & Disruption, Safety & Security, Crowding & Capacity, and Accessibility. The x-axis lists these clusters, while the y-axis indicates the prevalence percentage of the corpus. A secondary y-axis on the right shows the percentage of negative documents. The bars are colored in shades of blue. Comfort & Cleanliness has the highest prevalence at approximately 15 percentage, followed by Ticketing & Refunds at around 14 percentage. Staff & Communication is next at about 9 percentage, followed by Delay & Disruption at approximately 7 percentage. Safety & Security stands at around 4 percentage, Crowding & Capacity at about 3 percentage, and Accessibility at roughly 2 percentage. All values are approximated.

Prevalence of service failure clusters and share of negative documents. Source: Authors’ own work

Figure 2
A bar graph showing the prevalence of service failure clusters and the share of negative documents.The bar graph compares the prevalence of various service failure clusters and the share of negative documents. It features seven vertical bars, each representing a different cluster: Comfort & Cleanliness, Ticketing & Refunds, Staff & Communication, Delay & Disruption, Safety & Security, Crowding & Capacity, and Accessibility. The x-axis lists these clusters, while the y-axis indicates the prevalence percentage of the corpus. A secondary y-axis on the right shows the percentage of negative documents. The bars are colored in shades of blue. Comfort & Cleanliness has the highest prevalence at approximately 15 percentage, followed by Ticketing & Refunds at around 14 percentage. Staff & Communication is next at about 9 percentage, followed by Delay & Disruption at approximately 7 percentage. Safety & Security stands at around 4 percentage, Crowding & Capacity at about 3 percentage, and Accessibility at roughly 2 percentage. All values are approximated.

Prevalence of service failure clusters and share of negative documents. Source: Authors’ own work

Close Figure 2
Table 3

Service failure clusters: prevalence, corresponding NMF topics and visibility

Failure clusterDocs (n)Share (%)Corresponding NMF topics (share within cluster)Mean visibility
Comfort & Cleanliness9,14714.9T6 seating/class (17.6%); T4 onboard amenities; T2 reservations8.69
Ticketing & Refunds8,47613.8T9 passes & booking (29.9%); T1 purchase/validity (23.9%); T5 Delay Repay7.03
Staff & Communication5,3458.7T7 frontline staff & operations (17.4%)8.59
Delay & Disruption3,9916.5T5 Delay Repay/compensation (23.9%); T0 delay narratives (21.4%)6.71
Safety & Security2,1373.5T8 urban transit (15.5%); diffuse across narrative topics9.52
Crowding & Capacity1,6942.8T8 urban transit capacity (40.9%)9.20
Accessibility9621.6T8 urban transit (16.1%); diffuse across narrative topics9.07
Source(s): Authors’ own work

The inductive NMF solution on the failure-relevant subcorpus recovers clearly interpretable counterparts for five of the seven deductive clusters: delay and cancellation narratives (T0, 8.6% of subcorpus documents) and a distinct Delay Repay/compensation topic (T5, 4.7%) correspond to Delay & Disruption; ticket purchase and validity (T1, 8.6%) and passes, booking and reservations (T9, 15.8%) to Ticketing & Refunds; seat reservations (T2, 5.6%), onboard amenities, such as café, dining and sleeper cars (T4, 5.3%) and seating comfort and travel class (T6, 7.1%) to Comfort & Cleanliness; frontline staff and operations (T7, 5.6%) to Staff & Communication; and urban transit capacity (T8, 12.7%) to Crowding & Capacity. Safety & Security and Accessibility, by contrast, have no dedicated inductive counterpart: their documents distribute diffusely across the narrative topics (Table 3), indicating that these concerns surface as episodic elements embedded within journey narratives rather than as self-standing discursive fields – a substantive finding in its own right and the clearest illustration of why the deductive layer remains necessary. One residual conversational topic (T3, 26.0%) absorbs general evaluative talk. The emergence of Delay Repay as a discrete topic is theoretically noteworthy: the compensation process, a recovery mechanism in the sense of Smith et al. (1999), has become a failure category in its own right in British discourse. The cluster–topic correspondence is nonetheless weak in absolute terms (NMI = 0.23). The alignment reported here should therefore be read as qualitative convergence in topic content, not as statistical validation of the deductive scheme: the dominant-topic profile of each cluster is consistent with its theoretical content (Table 3), while the low mutual information and the large conversational residual jointly indicate that neither the deductive nor the inductive layer alone resolves the thematic structure of the discourse.

Corpus-wide sentiment is mildly positive (mean VADER compound = 0.125; 49.1% positive, 30.2% negative, 20.7% neutral documents), reflecting the mix of complaints, advice and enthusiast discussion typical of these communities, an important corrective to the assumption that transport social media is uniformly complaint-driven. Against this baseline, the failure dimensions separate sharply (Table 4; Figure 3). Reliability/punctuality (mean = 0.035; 44.9% negative) and safety & security (0.008; 46.3% negative) are the most negative dimensions, followed by staff & communication (0.062; 39.9% negative). Food & catering is the most positive dimension (0.274; 25.8% negative), with comfort/cleanliness and accessibility in between. At the cluster level, the contrast is starker still: documents in the Delay & Disruption cluster are 53.3% negative, with a mean compound of −0.060, the only cluster with net-negative valence alongside Safety & Security (−0.029; 49.5% negative). The pattern is fully consistent with the failure-asymmetry literature (Bitner et al., 1990): reliability failures are the affective core of negative passenger experience, while tangibles host substantial positive experience reporting alongside complaints.

Table 4

Sentiment and emotion profile by service dimension

DimensionMean sentiment% negativeAnger %Anxiety %Satisfaction %n
Reliability/punctuality0.03544.95.82.79.05,994
Safety & security0.00846.37.78.78.61,948
Staff & communication0.06239.96.33.17.75,105
Information systems0.13535.97.12.912.37,172
Ticketing/fares0.14632.85.02.37.19,869
Crowding/capacity0.15535.78.13.09.12,075
Comfort/cleanliness0.16832.66.12.517.26,993
Accessibility0.16833.711.23.910.41,541
Food & catering0.27425.88.72.815.31,760
Source(s): Authors’ own work
Figure 3
A bar graph comparing mean VADER compound sentiment across different service dimensions.A horizontal bar graph compares mean VADER compound sentiment across different service dimensions. The horizontal axis represents the mean VADER compound sentiment ranging from 0.00 to 0.25. The vertical axis lists the service dimensions: Food & Catering, Comfort / Cleanliness, Accessibility, Crowding / Capacity, Ticketing / Fares, Information Systems, Staff & Communication, Reliability / Punctuality, and Safety & Security. The bars are color-coded, with blue representing positive sentiment and red representing negative sentiment. Food & Catering has the highest positive sentiment at approximately 0.27, followed by Comfort / Cleanliness and Accessibility. Reliability / Punctuality and Safety & Security have the most negative sentiments, with values around 0.035 and 0.008, respectively. The graph highlights the contrast in sentiment across different service dimensions.

Mean sentiment by service dimension (red: most negative dimensions). Source: Authors’ own work

Figure 3
A bar graph comparing mean VADER compound sentiment across different service dimensions.A horizontal bar graph compares mean VADER compound sentiment across different service dimensions. The horizontal axis represents the mean VADER compound sentiment ranging from 0.00 to 0.25. The vertical axis lists the service dimensions: Food & Catering, Comfort / Cleanliness, Accessibility, Crowding / Capacity, Ticketing / Fares, Information Systems, Staff & Communication, Reliability / Punctuality, and Safety & Security. The bars are color-coded, with blue representing positive sentiment and red representing negative sentiment. Food & Catering has the highest positive sentiment at approximately 0.27, followed by Comfort / Cleanliness and Accessibility. Reliability / Punctuality and Safety & Security have the most negative sentiments, with values around 0.035 and 0.008, respectively. The graph highlights the contrast in sentiment across different service dimensions.

Mean sentiment by service dimension (red: most negative dimensions). Source: Authors’ own work

Close Figure 3

Emotion cues show distinctive affective signatures across dimensions. Anger peaks in accessibility discourse (11.2% of documents), where failed booked assistance and broken lifts constitute breaches of explicit service promises, precisely the failure type that the recovery literature identifies as most damaging to trust (Tax et al., 1998). Anxiety peaks in safety discourse (8.7%), a reputationally sensitive signal even at low complaint volumes. Satisfaction cues are most frequent in comfort (17.2%) and food & catering (15.3%) discourse, confirming that these dimensions are not merely complaint categories but arenas of positive differentiation, particularly in long-distance and night-train travel.

Table 5 reports OLS models with HC1-robust standard errors. Model 1 regresses the VADER negative-token proportion on cluster membership for all 61,328 documents, controlling for record type, subreddit, mentioned delay duration, text length and year. Membership in the Delay & Disruption cluster raises negativity by 3.6% points (b = 0.036, p < 0.001) and Safety & Security by 4.4 points (b = 0.044, p < 0.001) by far the largest effects and large relative to the corpus mean negative-token proportion followed by smaller positive effects of Staff & Communication (0.005, p < 0.001) and Comfort & Cleanliness (0.002, p < 0.01). Ticketing & Refunds membership is associated with lower negativity (−0.010, p < 0.001), consistent with the topic-level finding that a large share of ticketing discourse is transactional question-and-answer (validity, passes, reservations) rather than complaint. The overall explanatory power of Model 1 is modest (R2 = 0.028), as expected for a single-attribute model of document-level affect, but the cluster coefficients are precisely estimated and theoretically ordered.

Table 5

Regression results: predictors of negative sentiment and engagement (HC1 robust SEs)

PredictorModel 1: Negativity b (SE) [95% CI]Model 2: Engagement b (SE) [95% CI]p (M1)p (M2)
Delay & Disruption0.036 (0.001)
[0.034, 0.038]
−0.909 (0.133)
[−1.170, −0.648]
<0.001<0.001
Crowding & Capacity−0.001 (0.002)
[−0.005, 0.003]
0.311 (0.213)
[−0.107, 0.729]
0.4340.144
Ticketing & Refunds−0.010 (0.001)
[−0.012, −0.008]
−0.956 (0.126)
[−1.203, −0.709]
<0.001<0.001
Comfort & Cleanliness0.002 (0.001)
[0.000, 0.004]
0.635 (0.122)
[0.396, 0.874]
0.007<0.001
Staff & Communication0.005 (0.001)
[0.003, 0.007]
0.250 (0.151)
[−0.046, 0.546]
<0.0010.098
Safety & Security0.044 (0.002)
[0.040, 0.048]
0.818 (0.233)
[0.361, 1.275]
<0.001<0.001
Accessibility−0.002 (0.002)
[−0.006, 0.002]
0.075 (0.318)
[−0.548, 0.698]
0.2460.815
Delay duration (log min)−0.000 (0.001)
[−0.002, 0.002]
−0.162 (0.063)
[−0.285, −0.039]
0.8300.010
Controls (record type, subreddit, length, year)includedincluded
R2/N0.028/61,3280.325/2,173

Note(s): Intervals are Wald 95% confidence intervals computed as b ± 1.96 SE from the coefficients and standard errors reported above, both of which are rounded to three decimals; the limits are therefore accurate only to the precision of those inputs and narrow intervals may appear marginally inconsistent with the reported p-values. Because n = 61,328, conventional significance is attained by effects of negligible magnitude: Model 1 accounts for 2.8% of the variance in document-level negativity, so the substantive reading of this table rests on the relative ordering and magnitude of the coefficients rather than on their significance levels and no coefficient should be interpreted as establishing a practically important difference on the strength of its p-value alone.

Source(s): Authors’ own work

Model 2 regresses post-level engagement, log(1 + score) + log(1 + comments), on the same clusters (n = 2,173 posts; R2 = 0.33). Here the pattern inverts in a managerially important way: Safety & Security (b = 0.818, p < 0.001) and Comfort & Cleanliness (b = 0.635, p < 0.001) posts attract significantly more engagement, whereas Delay & Disruption (−0.909, p < 0.001) and Ticketing & Refunds (−0.956, p < 0.001) posts attract significantly less and longer self-reported delay durations further reduce engagement (−0.162, p < 0.01). Routine delay and ticketing complaints appear to elicit community fatigue; they are frequent, familiar and quickly resolved into advice threads, while safety incidents and comfort evaluations spark discussion, comparison and debate. High affective intensity therefore does not automatically translate into high visibility: the two most negative failure categories sit at opposite ends of the engagement distribution. This divergence is precisely what single-metric complaint counts collapse and precisely what the OPI is designed to reconcile.

Table 6 and Figure 4 report the OPI. Comfort & Cleanliness ranks first (OPI = 0.53), driven by the highest frequency combined with above-baseline negativity and visibility; Staff & Communication ranks second (0.36) on a balanced profile; and Ticketing & Refunds third (0.35), carried by frequency. Delay & Disruption (0.27) and Safety & Security (0.23) follow: both exhibit extreme negativity; safety additionally has the highest visibility of any cluster, but their lower document frequency moderates the composite score. Crowding & Capacity (0.10) and Accessibility (0.06) rank lowest on volume despite high visibility, with accessibility notable for its anger-dominant affective profile. The positioning of the clusters in the priority matrix (Figure 4) illustrates the diagnostic value of keeping the three components visible rather than reporting only the composite: comfort and ticketing occupy the high-frequency/moderate-negativity region where incremental service improvements accumulate large aggregate perception gains, while safety and delay occupy the low-frequency/high-negativity region where individual episodes carry disproportionate experiential and reputational weight.

Table 6

Operational Priority Index by service failure cluster (components max-normalized)

Failure clusterFrequency (F)Negativity (N)Visibility (V)OPIRank
Comfort & Cleanliness1.0000.5770.9130.5271
Staff & Communication0.5840.6760.9020.3572
Ticketing & Refunds0.9270.5130.7390.3513
Delay & Disruption0.4360.8790.7050.2714
Safety & Security0.2341.0001.0000.2345
Crowding & Capacity0.1850.5800.9660.1046
Accessibility0.1050.5900.9530.0597
Source(s): Authors’ own work
Figure 4
A scatter plot with bubbles representing different operational priorities.A scatter plot with bubbles representing different operational priorities. The x-axis represents the normalized frequency, ranging from 0.0 to 1.0, and the y-axis represents the normalized negativity, also ranging from 0.0 to 1.0. The size of each bubble indicates visibility. There are six labeled data points: Safety & Security, Delay & Disruption, Staff & Communication, Accessibility, Crowding & Capacity, Comfort & Cleanliness, and Ticketing & Refunds. Safety & Security has the highest visibility and is positioned at the top left with high negativity and low frequency. Delay & Disruption is positioned above the center with moderate visibility, high negativity, and moderate frequency. Staff & Communication is located near the center with moderate visibility, negativity, and frequency. Accessibility and Crowding & Capacity are clustered at the bottom left with low visibility, negativity, and frequency.

Operational Priority Matrix: frequency × negativity, bubble size = visibility. Source: Authors’ own work

Figure 4
A scatter plot with bubbles representing different operational priorities.A scatter plot with bubbles representing different operational priorities. The x-axis represents the normalized frequency, ranging from 0.0 to 1.0, and the y-axis represents the normalized negativity, also ranging from 0.0 to 1.0. The size of each bubble indicates visibility. There are six labeled data points: Safety & Security, Delay & Disruption, Staff & Communication, Accessibility, Crowding & Capacity, Comfort & Cleanliness, and Ticketing & Refunds. Safety & Security has the highest visibility and is positioned at the top left with high negativity and low frequency. Delay & Disruption is positioned above the center with moderate visibility, high negativity, and moderate frequency. Staff & Communication is located near the center with moderate visibility, negativity, and frequency. Accessibility and Crowding & Capacity are clustered at the bottom left with low visibility, negativity, and frequency.

Operational Priority Matrix: frequency × negativity, bubble size = visibility. Source: Authors’ own work

Close Figure 4

The ranking is highly robust: replacing mean negativity with the share of negative documents preserves the order almost perfectly (Spearman ρ = 0.96) and exclusive cluster assignment (ρ = 0.89) and additive aggregation (ρ = 0.75) yield only local reshuffling, chiefly between the delay and safety clusters. The only genuinely sensitive margin is thus the relative ordering of the two affect-intensive categories, which operators may legitimately resolve according to strategic priorities, for instance, elevating safety on reputational-risk grounds regardless of volume.

Because the index multiplies its three components with equal exponents, the implicit assumption that frequency, negativity and visibility contribute equally to operational priority warrants explicit scrutiny. Two considerations motivate retaining the equal-weight multiplicative form as the baseline specification. Firstly, the multiplicative structure encodes a conjunctive managerial logic: a dimension qualifies as an operational priority only if it scores on all three components simultaneously, so that a frequently discussed but affectively neutral theme, or an intensely negative but rare and invisible one, is correctly discounted – a property that additive aggregation does not preserve, since additive forms permit a high score on one component to compensate for near-zero scores on the others. Secondly, any differentiated weighting requires a normative judgement about the relative managerial value of volume, severity and salience, for which no external criterion is currently available; imposing one would embed an unstated policy preference in what is intended as a transparent diagnostic instrument.

Rather than fixing such weights arbitrarily, the sensitivity of the ranking to them was tested directly. Generalizing the index to OPIk = Fkα × Nkβ × Vkγ, five alternative weight vectors were evaluated against the equal-weight baseline (Table 7). Comfort & Cleanliness remains the first-ranked dimension under every specification that does not discount frequency. The leading group of Comfort & Cleanliness, Staff & Communication and Ticketing & Refunds is preserved intact when the exponent on frequency (ρ = 0.964) or on visibility (ρ = 0.964) is doubled, with only internal reordering of the second and third positions; when negativity is doubled (ρ = 0.893) the first two positions are unchanged and Ticketing & Refunds is displaced from third by Delay & Disruption. The ranking becomes materially sensitive only when frequency is actively discounted: halving its exponent (ρ = 0.750) raises Safety & Security to second place, and a severity-oriented specification that both halves frequency and doubles negativity (α = 0.5, β = 2, γ = 1) reorders the index substantially (ρ = 0.393), placing Safety & Security first and Delay & Disruption second. Crowding & Capacity and Accessibility occupy the sixth and seventh positions under all six specifications.

Table 7

Sensitivity of OPI rankings to alternative component weightings (cell entries are ranks; OPIk = Fkα × Nkβ × Vkγ)

Failure clusterEqual (1/1/1)Freq.↑ (2/1/1)Neg.↑ (1/2/1)Vis.↑ (1/1/2)Freq.↓ (0.5/1/1)Sev. (0.5/2/1)
Comfort & Cleanliness111114
Staff & Communication232233
Ticketing & Refunds325355
Delay & Disruption443542
Safety & Security554421
Crowding & Capacity666666
Accessibility777777
Spearman ρ vs. baseline1.0000.9640.8930.9640.7500.393

Note(s): Cell entries are ranks. Column headings denote the exponent vector (α/β/γ): Equal = equal weighting; Freq.↑ = frequency-weighted; Neg.↑ = negativity-weighted; Vis.↑ = visibility-weighted; Freq.↓ = frequency-discounted; Sev. = severity-oriented. Ranks are derived from the max-normalized components reported in Table 6. ρ is the Spearman rank correlation between the continuous index values under each specification and those under the equal-weight baseline.

Source(s): Authors’ own work

This boundary is interpretable rather than troubling, because it locates precisely the normative choice at stake. An operator prioritizing aggregate perceived service quality across the passenger base recovers the reported ranking; an operator prioritizing the severity of individual episodes on reputational or safety-regulatory grounds recovers the safety-led ordering. The index does not adjudicate between these positions and is not intended to: it makes the consequence of adopting either explicit, which is also why the three components are reported separately in Table 6 alongside the composite score. The equal-weight specification is therefore reported as the baseline on transparency grounds, with Table 7 supplied so that an operator holding a different weighting can read off the ranking implied by it directly.

The aggregate ranking conceals substantial cross-system variation (Figure 5). In UK-context discourse (n = 1,768; 39.5% negative documents, the highest of the three systems), Ticketing & Refunds leads the OPI (0.51) ahead of Delay & Disruption (0.32) and Staff & Communication (0.25), with the dedicated Delay Repay topic underscoring how compensation friction has become a failure category of its own. This Reddit profile contrasts with the overall positive Twitter sentiment reported by Mogaji and Erkan (2019), while extending their evidence of substantial service-quality variation across UK train groups; the difference also cautions against assuming that sentiment distributions transfer unchanged across platforms and sampling designs. In US-context discourse (n = 5,221; 32.1% negative), Comfort & Cleanliness dominates (0.76) – reflecting long-distance Amtrak travel in which seating, sleeper accommodation and onboard amenities are central to the experience, followed by Staff & Communication (0.42) and Ticketing (0.39). Continental European discourse (n = 2,474; 27.4% negative, the mildest profile) is led by Ticketing & Refunds (0.34), driven by reservation and pass complexity in Interrail travel, ahead of Comfort (0.26) and Delay (0.20).

Figure 5
A heat map comparing service failure clusters across different rail systems.A heat map compares service failure clusters across three rail systems: US Rail, UK Rail, and Continental Europe. The heat map is structured as a grid with rows representing different service failure clusters and columns representing the three rail systems. The rows are labeled as Delay & Disruption, Crowding & Capacity, Ticketing & Refunds, Comfort & Cleanliness, Staff & Communication, Safety & Security, and Accessibility. The columns are labeled as US Rail, UK Rail, and Continental Europe. The color scale on the right ranges from light blue to dark blue, indicating the OPI values from 0.1 to 0.7. Darker colors represent higher OPI values. Notable regions include Comfort & Cleanliness for US Rail with the highest value of 0.76, and Ticketing & Refunds for UK Rail with a value of 0.51. Overall, the heat map shows variation in service failure clusters across different rail systems, with some clusters having significantly higher OPI values in specific regions.

OPI by service failure cluster and rail system context. Source: Authors’ own work

Figure 5
A heat map comparing service failure clusters across different rail systems.A heat map compares service failure clusters across three rail systems: US Rail, UK Rail, and Continental Europe. The heat map is structured as a grid with rows representing different service failure clusters and columns representing the three rail systems. The rows are labeled as Delay & Disruption, Crowding & Capacity, Ticketing & Refunds, Comfort & Cleanliness, Staff & Communication, Safety & Security, and Accessibility. The columns are labeled as US Rail, UK Rail, and Continental Europe. The color scale on the right ranges from light blue to dark blue, indicating the OPI values from 0.1 to 0.7. Darker colors represent higher OPI values. Notable regions include Comfort & Cleanliness for US Rail with the highest value of 0.76, and Ticketing & Refunds for UK Rail with a value of 0.51. Overall, the heat map shows variation in service failure clusters across different rail systems, with some clusters having significantly higher OPI values in specific regions.

OPI by service failure cluster and rail system context. Source: Authors’ own work

Close Figure 5

Passenger-type profiles, though based on the small self-disclosed subset and therefore indicative rather than inferential, are consistent with these patterns and with the segment-specific quality preferences documented by dell’Olio et al. (2011): among tourists and occasional travelers (n = 490), 143 documents concern ticketing and 125 comfort; among the 19 self-identified daily commuters, 10 documents concern ticketing and 5 delays; 14 of the 17 documents contributed by self-identified rail professionals concern staff and communication, offering an insider perspective on the communication failures that passengers experience from the outside; and 36 of the 56 documents from self-identified skeptics are negative. Because the commuter and professional subsets comprise fewer than 20 documents each, these figures are reported as raw counts rather than percentages and support no inference beyond illustration; the segment comparison is retained only as an exploratory indication of where a purposively sampled follow-up study might focus.

A further axis of heterogeneity, and one bearing directly on the interpretation of the leading priority, is the type of service under discussion. Comfort and cleanliness is not homogeneous requirements across travel categories: the tolerance threshold a passenger applies to seating, thermal conditions, noise and sanitary facilities is conditioned by journey duration, by fare level and by whether the journey includes a night on board. A short urban commute and an overnight sleeper generate the same cluster label but describe substantively different service expectations, and an operational priority derived from the pooled corpus risks averaging across them. The service-type tags introduced in Section 3.3 commuter/urban (n = 3,495), high-speed (n = 2,007), sleeper/night (n = 1,456) and intercity/long-distance (n = 1,254) together with the inductive topic structure recovered in Section 4.2, permit a first characterization of this heterogeneity, though not a formal statistical disaggregation of it.

Three strands of evidence already established above bear on the question. First, the inductive layer does not recover comfort as a single undifferentiated theme but decomposes it into facets that track journey duration: seat reservations (T2, 5.6% of the failure-relevant subcorpus), onboard amenities including café, dining and sleeper cars (T4, 5.3%) and seating comfort and travel class (T6, 7.1%) emerge as separate topics, the first and third concerning the conditions of occupying a seat for a bounded interval and the second concerning provision for extended and overnight occupancy. Secondly, the cross-system comparison reported above is in part already a service-type comparison: the dominance of Comfort & Cleanliness in US discourse (OPI = 0.76) was attributed to the long-distance Amtrak network, on which seating, sleeper accommodation and onboard amenities are constitutive of the journey rather than incidental to it, whereas UK discourse, drawn from a predominantly commuter and intercity network, is led by Ticketing & Refunds (0.51) ahead of Delay & Disruption (0.32) and Staff & Communication (0.25), with comfort falling below all three. Thirdly, the prominence of ÖBB among named operators (537 mentions), reflecting the expansion of night-train services during the study window, indicates that sleeper travel is a substantial and growing locus of comfort discourse within the corpus.

Read together, these strands indicate that the aggregate first-rank position of Comfort & Cleanliness is not a single improvement target but a family of category-specific ones, differing both in what is demanded and in the threshold at which its absence becomes a complaint. On commuter and urban services the binding constraints are crowding, standing room and ventilation, endured over short exposure times and therefore tolerated at levels that would be unacceptable elsewhere; on high-speed services the premium fare raises the expectation baseline, and the salient facets are seat pitch, reservation integrity and onboard catering; on sleeper and night services exposure is continuous and the threshold correspondingly lowest, with berth quality, sanitary facilities and noise governing the assessment of the journey as a whole. The managerial consequence is that the OPI ranking should be disaggregated before it is used to allocate investment: rolling-stock interior specification for high-speed and sleeper services addresses a different failure mode from capacity and ventilation management on commuter services, and an operator running a mixed portfolio cannot treat the first-ranked dimension as a single programme of work.

This characterization is interpretive rather than estimated, and the present design does not support a quantitative disaggregation. The service-type tags are assigned by the same conservative explicit-reference rule as the system tags, so the four subsets are neither random samples of the corpus nor mutually exclusive; a single document may reference more than one service type, and the smallest of them (intercity/long-distance, n = 1,254) would in addition support only coarse comparison of cluster-level negativity. Estimating tolerance thresholds properly requires a design the present data cannot supply: journey-level attributes such as duration, fare class and departure time are not observable in the discourse, and separating a genuinely lower threshold from a higher incidence of the underlying failure would require them. A purposively sampled follow-up, in which service type is recorded rather than inferred and comfort facets are coded against journey duration, is therefore identified in Section 6 as the most immediate extension of this analysis.

Three findings carry theoretical weight. Firstly, five of the seven theory-informed failure clusters have clearly interpretable inductive counterparts, while Safety & Security and Accessibility remain diffusely distributed across the topic solution. This pattern offers partial naturalistic convergence with established service-quality dimensions (Eboli & Mazzulla, 2007; Parasuraman et al., 1988), but the weak cluster–topic correspondence (NMI = 0.23) and the large conversational residual show that the inductive solution does not statistically validate the deductive scheme. Instead, the two layers provide complementary representations: theory-guided measurement supplies stable operational categories, while topic modeling reveals how those categories are narrated and combined in open discourse.

Secondly, the results sharpen the failure-asymmetry thesis of the service encounter literature (Bitner et al., 1990; Smith et al., 1999). Delay and safety failures depress sentiment far more than any other dimension, consistent with reliability's dominance in satisfaction models, yet they generate less (delay) or only selectively more (safety) engagement. Affective intensity and public salience are therefore partially independent dimensions of failure severity – a distinction collapsed in single-metric complaint counts and invisible in survey means. The community-fatigue interpretation of the delay-engagement deficit suggests that repeated exposure normalizes even affectively intense failures, which has implications for how disruption “noise” should be baselined in monitoring systems.

Thirdly, the negative association between ticketing discourse and negativity, alongside its very high frequency, reveals a category of pre-failure friction: passengers invest substantial collective effort in understanding fares, validity, passes and reservations before anything has gone wrong. In expectation-disconfirmation terms (Oliver, 1980), this discourse is where expectations are formed and calibrated; its volume is a measure of the cognitive complexity cost that fare and reservation systems impose. This experienced complexity is invisible to both KPIs and complaint statistics, yet it shapes the expectation baseline against which every subsequent disruption is judged – and, per Tax et al. (1998), the fairness perceptions that govern post-failure trust.

The framework converts free social media data into three managerial instruments. The first is a continuous monitoring dashboard of passenger-perceived service failures, segmentable by rail system, operator and passenger type and updatable at negligible marginal cost, addressing the timeliness and granularity deficits of survey programs (de Oña & de Oña, 2015). The second is a proposed, and as yet untested design for early-warning alerting: because the pipeline scores negativity and visibility continuously, deviations of a cluster from its own baseline could in principle trigger alerts, with the safety cluster, combining maximal negativity and maximal visibility at low volume, the natural candidate for such a rule, since individual safety episodes propagate rapidly and carry reputational risk disproportionate to their frequency. It must be emphasized that no alerting rule was implemented or evaluated in the present study: neither a detection threshold, nor the lead time of a discourse signal relative to operational incidents, nor the false-alarm rate has been established, and no link to any operational outcome has been tested. This instrument is therefore advanced as a design and research agenda, whose validation route is specified in Section 6, rather than as a demonstrated result. The third is the OPI as a resource-allocation guide that ranks improvement areas by their combined frequency, affective impact and public salience.

The heterogeneity results translate directly into system-specific priorities. For British operators, simplification of ticketing and compensation processes promises the largest perception gains: Delay Repay friction emerged as a discrete discourse topic, meaning that the recovery mechanism itself, not only the underlying delays, is generating negative experience, a finding squarely in line with complaint-handling justice research (Tax et al., 1998). For Amtrak, onboard comfort and amenities dominate the discursive experience, implying that fleet renewal, seating and catering decisions carry perception weight comparable to schedule performance. For Continental operators serving international leisure travelers, reservation and pass complexity is the leading friction, suggesting integration and transparency investments in cross-border booking. Because the pipeline relies on open data, the official platform API and fully transparent lexicons whose matched terms are stored per document, every automated classification can be traced back to verbatim passenger language, a property that matters for managerial trust in and organizational adoption of, AI-derived indicators.

Two adoption caveats follow from the data-critical literature (boyd & Crawford, 2012). Reddit-derived indicators should be interpreted as measures of discourse salience among digitally active, English-speaking passengers, not as population-representative satisfaction estimates; and they are complements to, not substitutes for, KPI systems and statutory survey programs (Rashidi et al., 2017). Their comparative advantage lies in speed, granularity, verbatim traceability and the capture of failure categories, information breakdowns, app friction and assistance failures that official statistics under-record.

The hybrid deductive–inductive design proved essential: purely inductive topic modeling left roughly a quarter of failure-relevant discourse in a conversational residual, while the lexicon layer provided stable, auditable categories at the cost of missing unanticipated vocabulary. This trade-off is generic to text-as-data applications (Grimmer & Stewart, 2013) and the audit-trail design, storing every matched term per document – offers a practical template for operational deployments where classifications must be defensible. VADER performed adequately as a corpus-level instrument, its divergence from the independent cue lexicons (κ = 0.36) is concentrated in mixed-valence and sarcastic documents, the known hard cases of the genre (Hutto & Gilbert, 2014; Liu, 2012), supporting the design choice of continuous scores over hard labels in the regressions but also indicating clear headroom for transformer-based sentiment models (Devlin et al., 2019). The OPI proved highly robust to its main specification choices; the only sensitive margin is the relative ranking of the delay and safety clusters, which alternative weightings can interchange, and which operators may legitimately resolve according to strategic priorities. Finally, the comparison of NMF and LDA partitions (ARI = 0.17) is a reminder that topic solutions are method-dependent representations rather than discovered ground truth, reinforcing the case for anchoring them to theory-informed categories.

This study proposed and demonstrated an AI-driven framework for monitoring railway passenger experience through Reddit discourse, integrating a theory-informed failure lexicon, NMF topic modeling, rule-based sentiment and emotion analysis and an Operational Priority Index across USA, UK and Continental European rail contexts on a corpus of 61,328 documents. Delay and safety failures emerged as the affective core of negative passenger experience; comfort, ticketing and staff communication as the volume core; and the OPI reconciled these dimensions into actionable, system-specific priority rankings – ticketing and compensation friction in Britain, onboard comfort in the USA, reservation complexity in Continental leisure travel. The framework positions passenger-generated text not merely as consumer feedback but as an operational decision-support input that complements official performance indicators.

Several limitations qualify the findings. The corpus covers five primary English-language communities; systems with weak Reddit presence are under-represented, and the sample skews toward younger, digitally active, English-speaking travelers – the platform-shaped partiality that boyd and Crawford (2012) caution characterizes all big-data sources. Keyword-based retrieval and relevance filtering, while transparent, may miss failures expressed in unanticipated vocabulary, and listing-based sampling cannot guarantee exhaustive coverage of the window; the end-of-window volume increase visible in Figure 1 is a retrieval artifact rather than a discourse trend. Self-selected discourse cannot estimate absolute satisfaction levels, only relative failure salience; self-disclosed passenger types are available for only 1.1% of documents and the commuter and professional segments are very small. Sentiment and emotion measurement relies on rule-based instruments validated only against an internal cue lexicon (κ = 0.36) rather than human gold labels and the residual conversational topic indicates unresolved thematic content. Finally, the OPI weights, though robust to alternative specifications, involve normative choices and the robustness battery demonstrates internal consistency rather than validity: the index has not yet been benchmarked against an external criterion. Anchoring it to published complaint statistics or statutory passenger surveys for the same window, for instance, the category-level complaint data published by the British regulator would convert a plausible composite into a validated instrument and is the most direct route to establishing that discourse-derived priorities correspond to independently measured passenger concerns.

Future research could proceed along five lines. Firstly, re-estimate the pipeline with transformer-based components, BERTopic (Grootendorst, 2022) for topic modeling and fine-tuned contextual models (Devlin et al., 2019) for aspect-level sentiment and benchmark them against the transparent baseline established here, ideally anchored by human annotation of a stratified sample to obtain gold-standard validity estimates. Secondly, extend coverage to secondary system-specific communities, additional platforms and non-English sources such as consumer-complaint portals and app-store reviews, testing whether the priority structure replicates beyond the Anglophone Reddit population. Thirdly, link discourse signals to official incident logs and punctuality statistics of the named operators to test the early-warning capability causally, the design that would convert correlational salience into validated leading indicators. Fourthly, apply the pipeline longitudinally to evaluate the effect of operator interventions such as timetable recasts, fleet renewals, or compensation reforms, closing the loop between discourse monitoring and operations management. Fifthly, disaggregate the priority index by service type using a purposively sampled corpus in which journey duration, fare class and service category are recorded rather than inferred from explicit textual reference, so that the interpretive account of category-specific comfort thresholds offered in Section 4.6 short commuter exposure, premium high-speed expectation, continuous sleeper exposure can be replaced by estimated differences and so that a lower tolerance threshold can be distinguished empirically from a higher incidence of the underlying failure.

Elçin Noyan: Conceptualization, Methodology, Software, Data curation, Writing – original draft. Sezai Tunca: Validation, Formal analysis, Investigation, Supervision, Writing – review & editing.

The supplementary material for this article can be found online.

Baumgartner
,
J.
,
Zannettou
,
S.
,
Keegan
,
B.
,
Squire
,
M.
, &
Blackburn
,
J.
(
2020
).
The Pushshift Reddit dataset
. In
Proceedings of the International AAAI Conference on Web and Social Media
(Vol. 
14
(
1
), pp. 
830
839
). doi: .
Bitner
,
M. J.
,
Booms
,
B. H.
, &
Tetreault
,
M. S.
(
1990
).
The service encounter: Diagnosing favorable and unfavorable incidents
.
Journal of Marketing
,
54
(
1
),
71
84
. doi: .
Blei
,
D. M.
,
Ng
,
A. Y.
, &
Jordan
,
M. I.
(
2003
).
Latent Dirichlet allocation
.
Journal of Machine Learning Research
,
3
,
993
1022
.
Boyd
,
D.
, &
Crawford
,
K.
(
2012
).
Critical questions for big data: Provocations for a cultural, technological, and scholarly phenomenon
.
Information, Communication and Society
,
15
(
5
),
662
679
. doi: .
Casas
,
I.
, &
Delmelle
,
E. C.
(
2017
).
Tweeting about public transit—gleaning public perceptions from a social media microblog
.
Case Studies on Transport Policy
,
5
(
4
),
634
642
. doi: .
Collins
,
C.
,
Hasan
,
S.
, &
Ukkusuri
,
S. V.
(
2013
).
A novel transit rider satisfaction metric: Rider sentiments measured from online social media data
.
Journal of Public Transportation
,
16
(
2
),
21
45
. doi: .
dell’Olio
,
L.
,
Ibeas
,
A.
, &
Cecin
,
P.
(
2011
).
The quality of service desired by public transport users
.
Transport Policy
,
18
(
1
),
217
227
. doi: .
de Oña
,
J.
, &
de Oña
,
R.
(
2015
).
Quality of service in public transport based on customer satisfaction surveys: A review and assessment of methodological approaches
.
Transportation Science
,
49
(
3
),
605
622
. doi: .
Devlin
,
J.
,
Chang
,
M.-W.
,
Lee
,
K.
, &
Toutanova
,
K.
(
2019
).
BERT: Pre-training of deep bidirectional transformers for language understanding
. In
J. Burstein, C. Doran, & T. Solorio (Eds.)
.
Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)
(pp. 
4171
4186
).
Association for Computational Linguistics
. doi: .
Eboli
,
L.
, &
Mazzulla
,
G.
(
2007
).
Service quality attributes affecting customer satisfaction for bus transit
.
Journal of Public Transportation
,
10
(
3
),
21
34
. doi: .
Fellesson
,
M.
, &
Friman
,
M.
(
2008
).
Perceived satisfaction with public transport service in nine European cities
.
Journal of the Transportation Research Forum
,
47
(
3
),
93
103
. doi: .
Franzke
,
A. S.
,
Bechmann
,
A.
,
Zimmer
,
M.
,
Ess
,
C. M.
, &
Association of Internet Researchers
(
2020
).
Internet research: Ethical guidelines 3.0. Association of internet researchers
.
Available from:
 Link to the website
Gal-Tzur
,
A.
,
Grant-Muller
,
S. M.
,
Kuflik
,
T.
,
Minkov
,
E.
,
Nocera
,
S.
, &
Shoor
,
I.
(
2014
).
The potential of social media in delivering transport policy goals
.
Transport Policy
,
32
,
115
123
. doi: .
Grimmer
,
J.
, &
Stewart
,
B. M.
(
2013
).
Text as data: The promise and pitfalls of automatic content analysis methods for political texts
.
Political Analysis
,
21
(
3
),
267
297
. doi: .
Grootendorst
,
M.
(
2022
).
BERTopic: Neural topic modeling with a class-based TF-IDF procedure
. doi: .
Hutto
,
C. J.
, &
Gilbert
,
E.
(
2014
).
VADER: A parsimonious rule-based model for sentiment analysis of social media text
. In
Proceedings of the International AAAI Conference on Web and Social Media
(Vol. 
8
(
1
), pp. 
216
225
). doi: .
Kuflik
,
T.
,
Minkov
,
E.
,
Nocera
,
S.
,
Grant-Muller
,
S.
,
Gal-Tzur
,
A.
, &
Shoor
,
I.
(
2017
).
Automating a framework to extract and analyse transport related social media content: The potential and the challenges
.
Transportation Research Part C: Emerging Technologies
,
77
,
275
291
. doi: .
Lee
,
D. D.
, &
Seung
,
H. S.
(
1999
).
Learning the parts of objects by non-negative matrix factorization
.
Nature
,
401
(
6755
),
788
791
. doi: .
Liu
,
B.
(
2012
).
Sentiment analysis and opinion mining
.
Morgan & Claypool
. doi: .
Medvedev
,
A. N.
,
Lambiotte
,
R.
, &
Delvenne
,
J.-C.
(
2019
). The anatomy of Reddit: An overview of academic research. In
F.
 
Ghanbarnejad
,
R. S.
 
Roy
,
F.
 
Karimi
,
J.-C.
 
Delvenne
, &
B.
 
Mitra
(Eds.),
Dynamics on and of complex networks III: Machine learning and statistical physics approaches
(pp. 
183
204
).
Springer
. doi: .
Mogaji
,
E.
, &
Erkan
,
I.
(
2019
).
Insight into consumer experience on UK train transportation services
.
Travel Behaviour and Society
,
14
,
21
33
. doi: .
Mohammad
,
S. M.
, &
Turney
,
P. D.
(
2013
).
Crowdsourcing a word–emotion association lexicon
.
Computational Intelligence
,
29
(
3
),
436
465
. doi: .
Oliver
,
R. L.
(
1980
).
A cognitive model of the antecedents and consequences of satisfaction decisions
.
Journal of Marketing Research
,
17
(
4
),
460
469
. doi: .
Pang
,
B.
, &
Lee
,
L.
(
2008
).
Opinion mining and sentiment analysis
.
Foundations and Trends in Information Retrieval
,
2
(
1-2
),
1
135
. doi: .
Parasuraman
,
A.
,
Zeithaml
,
V. A.
, &
Berry
,
L. L.
(
1985
).
A conceptual model of service quality and its implications for future research
.
Journal of Marketing
,
49
(
4
),
41
50
. doi: .
Parasuraman
,
A.
,
Zeithaml
,
V. A.
, &
Berry
,
L. L.
(
1988
).
SERVQUAL: A multiple-item scale for measuring consumer perceptions of service quality
.
Journal of Retailing
,
64
(
1
),
12
40
.
Proferes
,
N.
,
Jones
,
N.
,
Gilbert
,
S.
,
Fiesler
,
C.
, &
Zimmer
,
M.
(
2021
).
Studying Reddit: A systematic overview of disciplines, approaches, methods, and ethics
.
Social Media + Society
,
7
(
2
),
1
14
. doi: .
Rashidi
,
T. H.
,
Abbasi
,
A.
,
Maghrebi
,
M.
,
Hasan
,
S.
, &
Waller
,
T. S.
(
2017
).
Exploring the capacity of social media data for modelling travel behaviour: Opportunities and challenges
.
Transportation Research Part C: Emerging Technologies
,
75
,
197
211
. doi: .
Smith
,
A. K.
,
Bolton
,
R. N.
, &
Wagner
,
J.
(
1999
).
A model of customer satisfaction with service encounters involving failure and recovery
.
Journal of Marketing Research
,
36
(
3
),
356
372
. doi: .
Tax
,
S. S.
,
Brown
,
S. W.
, &
Chandrashekaran
,
M.
(
1998
).
Customer evaluations of service complaint experiences: Implications for relationship marketing
.
Journal of Marketing
,
62
(
2
),
60
76
. doi: .
Tirachini
,
A.
,
Hensher
,
D. A.
, &
Rose
,
J. M.
(
2013
).
Crowding in public transport systems: Effects on users, operation and implications for the estimation of demand
.
Transportation Research Part A: Policy and Practice
,
53
,
36
52
. doi: .
Townsend
,
L.
, &
Wallace
,
C.
(
2016
).
Social media research: A guide to ethics
.
University of Aberdeen
.
Published in Railway Sciences. Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at Link to the terms of the CC BY 4.0 licence.

Supplementary data

or Create an Account

Close subscription notice
Close access options