This study aims to enhance the accuracy of key entity extraction from railway accident report texts and address challenges such as complex domain-specific semantics, data sparsity and strong inter-sentence semantic dependencies. A robust entity extraction method tailored for accident texts is proposed.
This method is implemented through a dual-branch multi-task mutual learning model named R-MLP, which jointly performs entity recognition and accident phase classification. The model leverages a shared BERT encoder to extract contextual features and incorporates a sentence span indexing module to align feature granularity. A cross-task mutual learning mechanism is also introduced to strengthen semantic representation.
R-MLP effectively mitigates the impact of semantic complexity and data sparsity in domain entities and enhances the model's ability to capture inter-sentence semantic dependencies. Experimental results show that R-MLP achieves a maximum F1-score of 0.736 in extracting six types of key railway accident entities, significantly outperforming baseline models such as RoBERTa and MacBERT.
This demonstrates the proposed method's superior generalization and accuracy in domain-specific entity extraction tasks, confirming its effectiveness and practical value.
1. Introduction
With the advancement of information technology in the field of railway safety regulation, the Railway Safety Supervision and Management Information System has been implemented across the entire network. It integrates a large volume of multimodal railway accident data, including structured data such as accident severity, line name, and line grade, as well as unstructured texts like accident overview, casualties, direct economic losses, and corrective measures. These multi-source heterogeneous safety data provide a comprehensive description of accident details, offering effective support for accident report text analysis and the quantitative analysis of accident losses (Shi et al., 2024). Existing railway accident texts are highly specialized in content organization, writing style, and terminology. To enhance analysis and information extraction capabilities for railway accident texts, it is necessary to develop a knowledge system in the railway safety supervision field. Entity extraction, as a key foundational task, faces challenges such as unclear entity boundaries, insufficient corpus acquisition, and a lack of high-quality samples. Additionally, influenced by the complexity of the railway system and the sporadic nature of accidents, accident-related entities exhibit significant sparsity in the corpus (Chen, 2023), demanding higher capabilities in entity feature analysis and semantic modeling.
In recent years, both domestically and internationally scholars have conducted extensive research on text entity extraction methods (Xiao & Chen, 2024). Traditional dictionary and rule-based pattern matching approaches, which rely on clear rule templates, struggle with comprehensive entity coverage. Statistical machine learning methods depend on high-quality, large-scale annotated corpora and expert-designed feature templates, limiting model generalization and transferability in railway-specific applications. Deep learning models, which do not require expert-designed entity features, have gradually become dominant in entity extraction, enhancing model performance. Significant progress has been made in general domains as well as in specific fields such as law (Lu & Li, 2024), medicine (Yang et al., 2016), military (Yin, Zhao, Zhao, Yao, & Huang, 2020), and finance (Xu, Zhu, Luo, & Dong, 2021). While entities like time, location, railway bureau, and equipment names in railway accident texts have been recognized (Du, Jin, Dai, Xue, & Wu, 2023), there remains a lack of analysis and extraction of complex causative factors, initial events, intermediate events, or final events (Shi et al., 2024). With the widespread application of large language models (e.g. GPT and ERNIE) in natural language processing, their robust semantic modeling capabilities offer new possibilities for entity extraction tasks. However, pre-trained large-scale models are typically trained on general-domain corpora and thus lack the capability to comprehend domain-specific terminology (Liang, Zhang, & Yan, 2024; Li et al., 2025), resulting in insufficient generalization and weak domain adaptability in railway-specific safety supervision text extraction. Given that existing general-purpose entity extraction methods fall short of meeting the requirements of safety text data analysis and mining, there is an urgent need for further in-depth exploration in areas such as extraction model architecture design, semantic modeling capability, and domain adaptability.
Compared with general-domain texts, railway accident texts exhibit a more pronounced stage-oriented narrative structure, typically recorded in accordance with the accident's evolution, including the occurrence process, specific causes, response measures, and rectification actions. Incorporating accident phase information into the entity extraction process can thus provide richer contextual and semantic relationship cues, thereby improving the accuracy of entity type recognition and extraction boundary determination. Based on these characteristics, this study proposes a railway accident text entity extraction model that integrates accident phase classification and mutual learning. The model jointly performs Entity Recognition (ER) and Accident Phase Classification (APC) using a shared Bidirectional Encoder Representations from Transformers (BERT) encoder, and introduces a Mutual Learning Mechanism (MLM) between the two task branches to enable bidirectional information exchange. This design fully exploits the potential associations of entities across different accident phases and enhances the accuracy and robustness of domain-specific entity extraction in scenarios with sparse samples.
2. Railway accident phase division and text entity definition
The railway sector has accumulated a large body of textual materials—such as determination documents, investigation reports, and handling reports—for various types of accidents, which systematically record the dynamic evolution of accidents over time (Shi, 2023). According to the literature (Shi et al., 2024), railway accidents can be divided, based on their development process, into different stages of an event chain: the initial event, the intermediate event, and the final event. The initial event refers to an incidental functional failure within the railway operation system; the intermediate event denotes a subsequent occurrence in the event chain that bridges the initial and final events; and the final event corresponds to the scenario that directly results in loss or damage, such as collisions, derailments, fires, or explosions.
Given that accident texts also contain information on the causes of accidents and post-incident handling measures, this study introduces two additional stages: the potential stage and the recovery stage. Accordingly, based on the accident overview texts, railway accidents are summarized into the following stages: potential causation stage, initial event stage, intermediate event stage, final event stage, and recovery/response stage. The initial and intermediate stages are described using patterns such as “human–operation–object,” “object–action–phenomenon,” and “location–existence–phenomenon.” The descriptions and examples of different accident stages are presented in Table 1.
Accident phases and examples in railway accident texts
| Accident phases | Description | Example |
|---|---|---|
| Potential Causation Stage | Accident has not yet occurred, but potential causative factors exist | Equipment aging, environmental changes, etc. |
| Initial Event Stage | Accident is triggered by an initial event | Human operational error, foreign object intrusion, etc. |
| Intermediate Event Stage | Relevant personnel or systems intervene to control accident progression | Emergency stop, train interception, etc. |
| Final Event Stage | Accident causes consequences | Collision, derailment, fire, explosion, etc. |
| Recovery/Response Stage | Post-incident site restoration, with impacts gradually eliminated | Rescue operations, manual handling |
| Accident phases | Description | Example |
|---|---|---|
| Potential Causation Stage | Accident has not yet occurred, but potential causative factors exist | Equipment aging, environmental changes, etc. |
| Initial Event Stage | Accident is triggered by an initial event | Human operational error, foreign object intrusion, etc. |
| Intermediate Event Stage | Relevant personnel or systems intervene to control accident progression | Emergency stop, train interception, etc. |
| Final Event Stage | Accident causes consequences | Collision, derailment, fire, explosion, etc. |
| Recovery/Response Stage | Post-incident site restoration, with impacts gradually eliminated | Rescue operations, manual handling |
Across different stages, accident texts provide comprehensive documentation of the overall incident, encompassing critical information such as the location of occurrence, causative factors, operational scope, responsible parties, response measures, and resulting impacts. Owing to the stochastic nature of accidents, entities associated with accident causation frequently exhibit pronounced sparsity. Based on the aforementioned informational dimensions, the key entity types within accident texts, along with representative examples, are defined as presented in Table 2.
Labels, descriptions, and examples of key entities in railway accident texts
| Entity type | Entity label | Example |
|---|---|---|
| Basic Information | Information | Time, location, personnel, organization, etc. |
| Causative Factors | Causality | Continuous heavy rainfall, strong winds, earthquakes, etc. |
| Initial Event | Initial Events | Catenary sagging into clearance limits, signalman misrouting turnout, etc. |
| Operation Type | Type | Train operation, shunting operation, etc. |
| Intermediate Event | Intermediate Events | Failure to stop, failure to intercept, etc. |
| Control Measures | Protective Measure | Driver-initiated stop, line closure, ground interception, etc. |
| Accident Consequences | Consequence | 12-min delay to the train, disruption of three passenger trains, etc. |
| Disposal Measures | Treatment Measures | Door securing, rescue and rerailing, power supply restoration, etc. |
| Entity type | Entity label | Example |
|---|---|---|
| Basic Information | Information | Time, location, personnel, organization, etc. |
| Causative Factors | Causality | Continuous heavy rainfall, strong winds, earthquakes, etc. |
| Initial Event | Initial Events | Catenary sagging into clearance limits, signalman misrouting turnout, etc. |
| Operation Type | Type | Train operation, shunting operation, etc. |
| Intermediate Event | Intermediate Events | Failure to stop, failure to intercept, etc. |
| Control Measures | Protective Measure | Driver-initiated stop, line closure, ground interception, etc. |
| Accident Consequences | Consequence | 12-min delay to the train, disruption of three passenger trains, etc. |
| Disposal Measures | Treatment Measures | Door securing, rescue and rerailing, power supply restoration, etc. |
In accident texts, there is a pronounced correlation between accident phases and entity types. In the potential causation stage, entities include the location of occurrence and the causative factors leading to the initial event. In the initial event stage, initial event entities are described in relation to attributes such as human operational errors, equipment failures, or environmental changes. In the intermediate event stage, one or more failed control measures, such as personnel response or equipment alarms, constitute intermediate event entities. In the final event stage, the focus shifts to the impacts triggered by accident consequences, describing the set of entities corresponding to different outcomes. In the recovery/response stage, the set of entities related to disposal measures is described.
Incorporating accident phases into entity extraction enables a clear delineation of the stages within the event chain to which entities belong, thereby reinforcing the intrinsic logical relationships of entities within the accident evolution process. It also provides constraint information for entity extraction. For example, crane equipment damage may be classified as a causative factor in the potential causation stage, but as an intermediate event in the intermediate stage. The associations between accident phases and key entities in railway accident texts are presented in Table 3.
Association between accident entities and accident phases in railway accident texts
| Accident stage | Common entities |
|---|---|
| Potential Causation Stage | Basic Information, Causative Factors |
| Initial Event Stage | Initial Event, Operation Type |
| Intermediate Event Stage | Basic Information, Intermediate Event, Control Measures |
| Final Event Stage | Accident Consequences |
| Recovery/Response Stage | Basic Information, Disposal Measures |
| Accident stage | Common entities |
|---|---|
| Potential Causation Stage | Basic Information, Causative Factors |
| Initial Event Stage | Initial Event, Operation Type |
| Intermediate Event Stage | Basic Information, Intermediate Event, Control Measures |
| Final Event Stage | Accident Consequences |
| Recovery/Response Stage | Basic Information, Disposal Measures |
To address the challenges of semantic complexity, entity sample sparsity, and strong inter-sentence dependencies in railway accident texts, this study proposes a multi-task mutual learning–based entity extraction model that jointly performs entity recognition and phase classification, leveraging phase information to enhance entity boundary detection and semantic discrimination. The framework adopts dual-branch architecture with a shared BERT encoder, employs sentence span indexing to achieve feature granularity alignment, and incorporates a mutual learning mechanism to promote semantic fusion between the two tasks. This design enables the extraction and structured output of key entities in the text, thereby improving the intelligent analysis and mining capabilities for railway accident texts. The overall framework of the proposed railway accident text entity extraction approach is shown in Figure 1.
The whole chart is inside a large outer rectangular boundary, which contains a smaller inner rectangular boundary. The “Text Input” is positioned below the inner boundary and has an upward arrow labeled “Entity Extraction Model” that connects it to the “Structured Output” box, which is positioned above the inner boundary. The “Structured Output” is in a long horizontal rectangular box at the top, reading “Basic Information, Causative Factors, Initial Event, Operation Type, Intermediate Event, Control Measures double ellipsis.” The “Text Input” is in a large rectangular box at the bottom, reading “The freight train driver reported that between K X X plus X X on the up line from pithead to X X Station, a landslide had occurred, with part of the debris sliding onto the ballast shoulder and infringing the clearance. Emergency braking measures were immediately implemented, and at X X colon X X the line was closed to clear the debris. Upon completion of the clearance at X X colon X X, the line was reopened with a speed limit of 25 kilometers per hour.” The inner rectangular boundary contains several connected rectangular boxes. A rectangular box at the bottom inside the inner boundary reads, “B E R T Backbone Feature Extraction.” Two upward arrows start from this box and go to two parallel rectangular boxes positioned above it: the one on the left reads, “Entity Recognition Branch,” and the one on the right reads, “Phase Classification Branch.” Upward arrows from both the “Entity Recognition Branch” and the “Phase Classification Branch” boxes connect to a rectangular box positioned above them that reads, “Feature Alignment.” Another rectangular box reading “Sentence Span Indexing” is placed to the left of the “Entity Recognition Branch” box, and an upward arrow from it leads to the “Feature Alignment” box. Upward arrows also go from the “Feature Alignment” box to a rectangular box reading “Mutual Learning” positioned above. A rectangular box above the “Mutual Learning” box reads, “Entity Extraction.” An upward arrow connects the “Mutual Learning” box to the “Entity Extraction” box.Framework for entity extraction from railway accident texts. Source(s): Authors’ own work
The whole chart is inside a large outer rectangular boundary, which contains a smaller inner rectangular boundary. The “Text Input” is positioned below the inner boundary and has an upward arrow labeled “Entity Extraction Model” that connects it to the “Structured Output” box, which is positioned above the inner boundary. The “Structured Output” is in a long horizontal rectangular box at the top, reading “Basic Information, Causative Factors, Initial Event, Operation Type, Intermediate Event, Control Measures double ellipsis.” The “Text Input” is in a large rectangular box at the bottom, reading “The freight train driver reported that between K X X plus X X on the up line from pithead to X X Station, a landslide had occurred, with part of the debris sliding onto the ballast shoulder and infringing the clearance. Emergency braking measures were immediately implemented, and at X X colon X X the line was closed to clear the debris. Upon completion of the clearance at X X colon X X, the line was reopened with a speed limit of 25 kilometers per hour.” The inner rectangular boundary contains several connected rectangular boxes. A rectangular box at the bottom inside the inner boundary reads, “B E R T Backbone Feature Extraction.” Two upward arrows start from this box and go to two parallel rectangular boxes positioned above it: the one on the left reads, “Entity Recognition Branch,” and the one on the right reads, “Phase Classification Branch.” Upward arrows from both the “Entity Recognition Branch” and the “Phase Classification Branch” boxes connect to a rectangular box positioned above them that reads, “Feature Alignment.” Another rectangular box reading “Sentence Span Indexing” is placed to the left of the “Entity Recognition Branch” box, and an upward arrow from it leads to the “Feature Alignment” box. Upward arrows also go from the “Feature Alignment” box to a rectangular box reading “Mutual Learning” positioned above. A rectangular box above the “Mutual Learning” box reads, “Entity Extraction.” An upward arrow connects the “Mutual Learning” box to the “Entity Extraction” box.Framework for entity extraction from railway accident texts. Source(s): Authors’ own work
3. Railway accident text entity extraction model based on accident phase classification and mutual learning
Railway accident texts have domain-specific characteristics, with entities typically demonstrating complex structures, fine-grained semantics, and sparse distributions. To improve the accuracy of key entity recognition in railway safety accident texts, this study proposes a text entity extraction model called R-MLP (Railway Entity Extraction with Mutual Learning and Accident Phase Classification), which models entity recognition and phase classification as independent tasks, while enhancing extraction performance through inter-task information exchange. A pre-trained language model, BERT, is employed as a shared encoder to capture the initial semantic features of the text, which are then fed into the entity recognition branch and the accident phase classification branch, respectively. In the entity recognition branch, a Conditional Random Field (CRF) (Shi, 2023; Pooja & Jagadeesh, 2024) is used to model label dependencies, thereby improving the consistency and accuracy of boundary recognition. In the phase classification branch, semantic information is aggregated to achieve effective accident phase identification. Considering that the same entity may carry different semantics across phases, a feature alignment and mutual learning mechanism (Wang, Li, & Ge, 2023) is introduced between the two tasks to facilitate information sharing and interaction. This mechanism provides contextual constraints for entity recognition, enhances the model's semantic parsing capability, and improves robustness and generalization under conditions of sample sparsity. Ultimately, the model achieves accurate extraction of key accident entities (Entity Extraction, EE). The framework of the proposed R-MLP model is illustrated in Figure 2.
The schematic illustrates a model architecture divided into four main modules by dashed lines: the “Sentence Span Indexing Module” at the top, the “Backbone Feature Extraction Module” at the bottom left in a yellow background, the “Feature Alignment Module” in the middle with a blue background, and the “Feature Reconstruction Module” on the right. In the “Sentence Span Indexing Module,” a yellow rectangular box labeled “Span Split” has an arrow pointing to a vertical series of four violet rectangular boxes labeled “Span subscript 1,” “Span subscript 2,” “Span subscript 3,” and “Span subscript m,” separated by a dotted line. In the “Backbone Feature Extraction Module,” a vertical series of elements, upper T subscript n on the left, and an upward arrow connect to a yellow rectangular box labeled “Span Split.” In the “Sentence Span Indexing Module,” a rightward arrow connects to a column of green rectangular boxes labeled “C subscript 1,” “C subscript 2,” “C subscript 3,” and “C subscript 4,” with dotted lines indicating continuation, and “C subscript n,” representing input tokens. These input boxes connect horizontally to a pink rectangular box labeled “Tokenizer” in the left-center. From the “Tokenizer” box, multiple horizontal arrows extend rightward to a column of green rectangular boxes in the center labeled “T o k subscript C L S,” “T o k subscript 1,” “T o k subscript 2,” “T o k subscript 3,” “T o k subscript 4,” with dotted lines indicating continuation, and “T o k subscript n” and “T o k subscript s e p” at the bottom, representing tokenized outputs. These token boxes feed rightward into a light blue rectangular box labeled “BERT Model.” From the “BERT Model,” two lines rightwards: the upper arrow leads to a blue parallelogram, and the lower arrow leads to a blue parallelogram. The upper arrow points to a rectangular box labeled “Linear.” The rightward arrow points to the violet parallelogram, which then points to the upper input of a rectangular box labeled “Feature Alignment” in the center. The lower arrow from “BERT Model” points to another rectangular box labeled “Linear.” The rightward arrow points to the yellow parallelogram, which then points to the lower input of “Feature Alignment.” An arrow from “Span subscript 1” in the “Sentence Span Indexing Module” also points to the upper part of the “Feature Alignment.” Two rightward arrows, one from the top and another from the bottom of “Feature Alignment,” connect to two streams of multiple, stacked feature blocks in purple and orange. The “Feature Reconstruction Module” is divided into three shaded columns; the first one on the left, from the “Feature Alignment” box, has two main outputs: an upper output that leads to both a rectangular box in light blue labeled “Cross Attention” and a rectangular box labeled “Add,” and a lower output that leads only to the “Add” box. The “Cross Attention” box has an additional input line from the upper output of the “Feature Alignment” box, indicated by an arrow. The “Add” box receives an input from “Cross Attention” and an input from the lower output of “Feature Alignment.” Both “Cross Attention” and “Add” have outputs pointing to a large rectangular box labeled in the second shaded column in the middle, “Feature Concat.” The “Feature Concat” box has two outputs: an upper output that points to a violet parallelogram and a lower output that points to a yellow parallelogram. The upper violet parallelogram points to a rectangular box labeled “C R F” in the third shaded column on the right. An arrow labeled “E R” leads out from the “C R F” box and points to the right; the lower yellow parallelogram points to an output line labeled “A P C.”R-MLP model framework. Source(s): Authors’ own work
The schematic illustrates a model architecture divided into four main modules by dashed lines: the “Sentence Span Indexing Module” at the top, the “Backbone Feature Extraction Module” at the bottom left in a yellow background, the “Feature Alignment Module” in the middle with a blue background, and the “Feature Reconstruction Module” on the right. In the “Sentence Span Indexing Module,” a yellow rectangular box labeled “Span Split” has an arrow pointing to a vertical series of four violet rectangular boxes labeled “Span subscript 1,” “Span subscript 2,” “Span subscript 3,” and “Span subscript m,” separated by a dotted line. In the “Backbone Feature Extraction Module,” a vertical series of elements, upper T subscript n on the left, and an upward arrow connect to a yellow rectangular box labeled “Span Split.” In the “Sentence Span Indexing Module,” a rightward arrow connects to a column of green rectangular boxes labeled “C subscript 1,” “C subscript 2,” “C subscript 3,” and “C subscript 4,” with dotted lines indicating continuation, and “C subscript n,” representing input tokens. These input boxes connect horizontally to a pink rectangular box labeled “Tokenizer” in the left-center. From the “Tokenizer” box, multiple horizontal arrows extend rightward to a column of green rectangular boxes in the center labeled “T o k subscript C L S,” “T o k subscript 1,” “T o k subscript 2,” “T o k subscript 3,” “T o k subscript 4,” with dotted lines indicating continuation, and “T o k subscript n” and “T o k subscript s e p” at the bottom, representing tokenized outputs. These token boxes feed rightward into a light blue rectangular box labeled “BERT Model.” From the “BERT Model,” two lines rightwards: the upper arrow leads to a blue parallelogram, and the lower arrow leads to a blue parallelogram. The upper arrow points to a rectangular box labeled “Linear.” The rightward arrow points to the violet parallelogram, which then points to the upper input of a rectangular box labeled “Feature Alignment” in the center. The lower arrow from “BERT Model” points to another rectangular box labeled “Linear.” The rightward arrow points to the yellow parallelogram, which then points to the lower input of “Feature Alignment.” An arrow from “Span subscript 1” in the “Sentence Span Indexing Module” also points to the upper part of the “Feature Alignment.” Two rightward arrows, one from the top and another from the bottom of “Feature Alignment,” connect to two streams of multiple, stacked feature blocks in purple and orange. The “Feature Reconstruction Module” is divided into three shaded columns; the first one on the left, from the “Feature Alignment” box, has two main outputs: an upper output that leads to both a rectangular box in light blue labeled “Cross Attention” and a rectangular box labeled “Add,” and a lower output that leads only to the “Add” box. The “Cross Attention” box has an additional input line from the upper output of the “Feature Alignment” box, indicated by an arrow. The “Add” box receives an input from “Cross Attention” and an input from the lower output of “Feature Alignment.” Both “Cross Attention” and “Add” have outputs pointing to a large rectangular box labeled in the second shaded column in the middle, “Feature Concat.” The “Feature Concat” box has two outputs: an upper output that points to a violet parallelogram and a lower output that points to a yellow parallelogram. The upper violet parallelogram points to a rectangular box labeled “C R F” in the third shaded column on the right. An arrow labeled “E R” leads out from the “C R F” box and points to the right; the lower yellow parallelogram points to an output line labeled “A P C.”R-MLP model framework. Source(s): Authors’ own work
3.1 Sentence span indexing module
In dual-branch multi-task learning models, the entity recognition (ER) task is typically annotated at the character (token) level, whereas the accident phase classification (APC) task is performed at the sentence level. This difference in granularity prevents feature alignment between the two branches during the mutual learning process. To address this issue, this study designs a sentence span indexing module, as shown in Equation (1). The primary function of this module is to segment the input text according to natural language clause boundaries indicated by punctuation marks (e.g., commas, periods), and to extract the start and end indices of the original tokens for each clause. These index mappings serve as the key basis for achieving feature alignment, and the process can be expressed as follows:
In the equation, is a text sequence of length n, and represents the set of boundary indices for sentence segmentation. Each specifies the start token index and the end token index of the -th sentence. The sentence span indexing results are illustrated in Table 4, where each pair of numbers in parentheses indicates the start and end character positions of a clause in the original text. These boundary indices provide essential alignment information between token-level and sentence-level features in the subsequent dual-branch mutual learning process.
Results of clause index extraction
| Text | Sentence span Indexing() |
|---|---|
| Due to a landslide halting at K122 + 863, the locomotive's left-side pilot (cowcatcher) struck the debris during the stop. At 03:31, the driver requested rescue assistance | (0,15) |
| (17,32) | |
| (34,44) |
| Text | Sentence span Indexing( |
|---|---|
| Due to a landslide halting at K122 + 863, the locomotive's left-side pilot (cowcatcher) struck the debris during the stop. At 03:31, the driver requested rescue assistance | (0,15) |
| (17,32) | |
| (34,44) |
3.2 Backbone feature extraction module
Although the objectives of the entity recognition task and the phase classification task differ, both operate on the same text sequence and thus exhibit a high degree of overlap in underlying linguistic features. Sharing backbone features not only enables the information-interaction model to improve its understanding of one task while optimizing the other, but also enhances overall modeling efficiency and effectiveness. As shown in Equation (2), the backbone feature extraction module primarily encodes the input text to obtain the initial semantic features . First, the Tokenizer module segments and encodes the text , automatically appending special tokens such as [CLS] and [SEP] to meet the structural input requirements of the BERT model. Then, the BERT model employs a multi-layer Transformer encoder to perform representation learning on the input sequence, effectively modeling bidirectional dependencies and semantic relationships within the context (He, Zhao, & Tang, 2025). This process generates high-dimensional semantic feature vectors for each character, enriched with contextual information, and can be expressed as:
As shown in Figure 3, the initial semantic features are simultaneously fed into both the ER branch and the APC branch. By employing a shared backbone feature extraction module, the framework not only improves parameter utilization efficiency but also provides a unified and high-quality representation space for subsequent tasks, thereby facilitating effective multi-task collaborative learning.
The flowchart displays a multi-stage processing architecture for natural language processing tasks. On the far left, a label “T subscript n” points with a rightwards arrow to a column of green rectangular boxes labeled “C subscript 1,” “C subscript 2,” “C subscript 3,” and “C subscript 4,” with dotted lines indicating continuation, and “C subscript n,” representing input tokens. These input boxes are connected horizontally to a pink rectangular box labeled “Tokenizer” in the left center. From the “Tokenizer” box, multiple horizontal arrows extend rightward to a column of green rectangular boxes in the center labeled “T o k subscript C L S,” “T o k subscript 1,” “T o k subscript 2,” “T o k subscript 3,” “T o k subscript 4,” with dotted lines indicating continuation, and “T o k subscript n” and “T o k subscript s e p” at the bottom, representing tokenized outputs. These token boxes feed rightward into a light blue rectangular box labeled “BERT Model.” From the “BERT Model,” two lines rightwards: the upper arrow leads to a blue parallelogram labeled “F subscript blackbone” and continues to “E R”; the lower arrow leads to a blue parallelogram labeled “F subscript blackbone” and continues to “A P C.”Backbone feature extraction module. Source(s): Authors’ own work
The flowchart displays a multi-stage processing architecture for natural language processing tasks. On the far left, a label “T subscript n” points with a rightwards arrow to a column of green rectangular boxes labeled “C subscript 1,” “C subscript 2,” “C subscript 3,” and “C subscript 4,” with dotted lines indicating continuation, and “C subscript n,” representing input tokens. These input boxes are connected horizontally to a pink rectangular box labeled “Tokenizer” in the left center. From the “Tokenizer” box, multiple horizontal arrows extend rightward to a column of green rectangular boxes in the center labeled “T o k subscript C L S,” “T o k subscript 1,” “T o k subscript 2,” “T o k subscript 3,” “T o k subscript 4,” with dotted lines indicating continuation, and “T o k subscript n” and “T o k subscript s e p” at the bottom, representing tokenized outputs. These token boxes feed rightward into a light blue rectangular box labeled “BERT Model.” From the “BERT Model,” two lines rightwards: the upper arrow leads to a blue parallelogram labeled “F subscript blackbone” and continues to “E R”; the lower arrow leads to a blue parallelogram labeled “F subscript blackbone” and continues to “A P C.”Backbone feature extraction module. Source(s): Authors’ own work
3.3 Feature alignment module
The feature alignment module, as a core component of the dual-branch mutual learning framework, is designed to ensure semantic consistency between the token-level modeling of the ER task and the sentence-level modeling of the APC task. The module first segments the input sequence into sentence-level units based on sentence boundary indices (). For the ER task, this segmentation provides explicit entity boundary constraints, preventing incorrect recognition of entities that span multiple sentences. For the APC task, the module generates stable sentence representations by aggregating (via mean pooling) the token features within each sentence, treating each clause as an independent classification unit to reduce semantic ambiguity and phase overlap. More importantly, through a structured index mapping mechanism, the module strengthens the correspondence between token-level and sentence-level features despite the difference in task granularity, thereby achieving explicit cross-task representation space alignment. The specific process is as follows:
The current dual branches share the same feature dimensionality, the features of the branches can be defined as:
where denotes the total number of tokens and the feature dimension. Based on the sentence boundary indices , the token feature set for the same sentence can be extracted as:
where denotes the token-level feature block after segmentation, and represents the length of the -th sentence. Ultimately, the token-level feature block set for the ER branch and the token-level feature block set for the APC branch are obtained, which can be expressed as:
For the phase classification branch, each is further aggregated using mean pooling to obtain the sentence-level vector :
where denotes the sentence-level feature vector, and the final sentence-level feature vector set is
As shown in Figure 4, the feature alignment module uses to denote the feature segmentation symbol. Through a structured segmentation approach, the ER branch can satisfy the strict boundary requirements of entity recognition, while the APC branch can extract sentence-level feature representations to meet the overall semantic understanding needs of the phase classification task. At the same time, the boundary index–based method preserves the alignment relationship between the token-level features in the ER branch and the sentence-level features in the APC branch.
The flowchart displays a multi-branch architecture with two parallel processing streams. On the top left, a vertical stack of four rectangular boxes labeled “Span subscript 1,” “Span subscript 2,” “Span subscript 3,” with dotted lines indicating continuation, and “Span subscript m,” connected to a vertical line. The vertical line connecting the two circled nodes is labeled “SPAN subscript m.” From this vertical line, two circular enclosure nodes branch the flow into two paths. The upper path receives input from “H superscript e r” on the left, connects horizontally through the upper circled node, and outputs rightward through stacked rectangular boxes labeled “S superscript e r,” with dotted lines indicating continuation. The lower path receives input from “H superscript a p c” on the left, connects horizontally through the lower circled node, and outputs rightward to stacked rectangular boxes labeled “S superscript a p c,” which then feed into a yellow rectangular box labeled “Mean Pooling.” From the “Mean Pooling” box, an arrow extends rightward to the final output “G superscript a p c,” represented as stacked rectangular boxes with dotted lines.Feature alignment module. Source(s): Authors’ own work
The flowchart displays a multi-branch architecture with two parallel processing streams. On the top left, a vertical stack of four rectangular boxes labeled “Span subscript 1,” “Span subscript 2,” “Span subscript 3,” with dotted lines indicating continuation, and “Span subscript m,” connected to a vertical line. The vertical line connecting the two circled nodes is labeled “SPAN subscript m.” From this vertical line, two circular enclosure nodes branch the flow into two paths. The upper path receives input from “H superscript e r” on the left, connects horizontally through the upper circled node, and outputs rightward through stacked rectangular boxes labeled “S superscript e r,” with dotted lines indicating continuation. The lower path receives input from “H superscript a p c” on the left, connects horizontally through the lower circled node, and outputs rightward to stacked rectangular boxes labeled “S superscript a p c,” which then feed into a yellow rectangular box labeled “Mean Pooling.” From the “Mean Pooling” box, an arrow extends rightward to the final output “G superscript a p c,” represented as stacked rectangular boxes with dotted lines.Feature alignment module. Source(s): Authors’ own work
3.4 Asymmetric mutual learning module
The entity recognition task is a fine-grained sequence labeling task that classifies each token, whereas the phase classification task is a coarse-grained sentence classification task that assigns labels to entire sentences. Owing to the differences in task granularity and information requirements, this study adopts an asymmetric mutual learning module design.
In the ER branch, to enhance the model's ability to learn contextual information, this study designs a cross-attention–based feature fusion module (citation 14), where the feature from the APC branch serves as the Query, and the feature from the entity recognition branch serves as both the Key and the Value. Through the attention mechanism, the feature is adaptively weighted and adjusted, enabling each token in to learn the overall semantic context of its corresponding sentence, thereby improving entity recognition accuracy in complex contexts. Since and , the feature is first replicated along the token dimension of , which can be expressed as:
where the resulting query matrix has the same dimensionality as , and together with serving as the key and value , is fed into the cross-attention module. The cross-attention process can be expressed as:
To mitigate gradient vanishing and accelerate model convergence, the attention-weighted features are added to the original features via a residual connection, forming the fused features . This process can be expressed as:
As shown in Figure 5, the cross-attention–based feature fusion module uses to denote the repeat operation, in which the sentence-level feature is replicated along the token dimension of . This fusion strategy preserves the fine-grained token-level semantic representation in the ER branch while incorporating sentence-level information, thereby enhancing the entity recognition branch's ability to capture the semantics of the containing sentence. In addition, the residual connection design facilitates optimized information flow and alleviates the vanishing gradient problem in deep networks.
The flowchart contains two input streams that process and merge into a final output. On the left, a stack of four rectangular boxes labeled “S superscript e r,” with dotted lines indicating continuation, connected by a horizontal rightwards arrow to “S subscript i superscript e r.” From “S subscript i superscript e r,” a dashed vertical line extends downward labeled “l subscript i.” On the lower left, a stack of three rectangular boxes labeled “G superscript a p c,” with dotted lines indicating continuation, connected by a horizontal rightwards arrow to “G subscript i superscript a p c,” which then connects to a circled multiplication symbol. From this multiplication operation, a horizontal rightward arrow extends rightward to “q subscript i,” which then connects upward into the attention box. In the center-right, an enclosed rectangular box labeled “Attention” is written vertically on a light blue shaded region with “v subscript i” at the top and “K subscript i.” Three arrows feed into the attention mechanism: “k subscript i” and “v subscript i” both originate from “S subscript i superscript e r,” and “q subscript i” originates from “q subscript i.” The attention outputs rightward through a circled addition symbol, and a horizontal arrow from “S subscript i superscript e r” also connects directly to this addition symbol. From the addition symbol, a horizontal arrow extends rightward to produce the final output “s subscript i superscript fused.”Feature fusion module based on cross-attention. Source(s): Authors’ own work
The flowchart contains two input streams that process and merge into a final output. On the left, a stack of four rectangular boxes labeled “S superscript e r,” with dotted lines indicating continuation, connected by a horizontal rightwards arrow to “S subscript i superscript e r.” From “S subscript i superscript e r,” a dashed vertical line extends downward labeled “l subscript i.” On the lower left, a stack of three rectangular boxes labeled “G superscript a p c,” with dotted lines indicating continuation, connected by a horizontal rightwards arrow to “G subscript i superscript a p c,” which then connects to a circled multiplication symbol. From this multiplication operation, a horizontal rightward arrow extends rightward to “q subscript i,” which then connects upward into the attention box. In the center-right, an enclosed rectangular box labeled “Attention” is written vertically on a light blue shaded region with “v subscript i” at the top and “K subscript i.” Three arrows feed into the attention mechanism: “k subscript i” and “v subscript i” both originate from “S subscript i superscript e r,” and “q subscript i” originates from “q subscript i.” The attention outputs rightward through a circled addition symbol, and a horizontal arrow from “S subscript i superscript e r” also connects directly to this addition symbol. From the addition symbol, a horizontal arrow extends rightward to produce the final output “s subscript i superscript fused.”Feature fusion module based on cross-attention. Source(s): Authors’ own work
In the APC branch, to incorporate fine-grained local semantic information from the ER branch, this study employs a lightweight compression-fusion module. The token-level features from the ER branch are aggregated via mean pooling to obtain the compressed feature which can be expressed as:
where has the same dimensionality as , and is directly added to to achieve semantic fusion:
As shown in Figure 6, compared with the complex cross-attention mechanism adopted in the ER branch, the APC branch employs a more lightweight strategy for feature fusion. By simply applying feature compression followed by an addition operation, it can incorporate supplementary semantic information from the ER branch. This asymmetric fusion design ensures effective information interaction while significantly reducing computational overhead, thereby avoiding structural redundancy in the model and achieving a balance between efficiency and performance.
The flowchart displays two input streams merging through a processing pipeline. On the top left, a stack of four rectangular boxes labeled “S superscript e r” with dotted lines indicating continuation, connected by a horizontal arrow to “S subscript i superscript e r” on the right. Below this, “S subscript i superscript e r” feeds downward into a yellow rectangular box labeled “Mean Pooling,” which then outputs downward to “S subscript i superscript i.” On the lower left, four stacked rectangular boxes labeled “G superscript a p c” with dotted lines indicating continuation, connected by a horizontal arrow to “G subscript i superscript a p c,” which continues rightward to a circled plus symbol. From the circled plus point, a horizontal arrow extends rightward to “G subscript i superscript fused,” the combined output of the two processing branches.Compressed fusion model. Source(s): Authors’ own work
The flowchart displays two input streams merging through a processing pipeline. On the top left, a stack of four rectangular boxes labeled “S superscript e r” with dotted lines indicating continuation, connected by a horizontal arrow to “S subscript i superscript e r” on the right. Below this, “S subscript i superscript e r” feeds downward into a yellow rectangular box labeled “Mean Pooling,” which then outputs downward to “S subscript i superscript i.” On the lower left, four stacked rectangular boxes labeled “G superscript a p c” with dotted lines indicating continuation, connected by a horizontal arrow to “G subscript i superscript a p c,” which continues rightward to a circled plus symbol. From the circled plus point, a horizontal arrow extends rightward to “G subscript i superscript fused,” the combined output of the two processing branches.Compressed fusion model. Source(s): Authors’ own work
This study proposes an asymmetric mutual learning strategy that fully accounts for the fundamental differences between the entity recognition task and the phase classification task in terms of granularity level, semantic focus, and modeling objectives. In the ER branch, a cross-attention–based learning mechanism is introduced, whereby phase-related contextual information is provided to help the model more accurately locate and distinguish key entities. In the APC branch, a compression-fusion learning mechanism is employed to effectively incorporate local entity information, thereby enhancing the precision of sentence-level semantic representations. This asymmetric learning strategy not only enables effective interaction between features of different granularities but also balances computational efficiency with structural simplicity.
3.5 Feature reconstruction module
During the mutual learning process, the model performs segmentation of the original continuous features and local feature fusion, resulting in the features being divided into multiple independent segments. The feature reconstruction module accurately concatenates these segmented features in their original order to restore a complete and coherent feature representation, thereby providing the subsequent entity recognition and phase classification tasks with complete input representations. In the ER branch, the segmented features are reassembled to restore the complete token feature sequence , and the process is expressed as:
Similarly, in the APC branch, the segmented features are reassembled to restore the complete sentence feature sequence .
3.6 Dual-branch prediction and joint loss function design
In the entity recognition task, the model performs fine-grained label prediction at the token level, i.e., it identifies and classifies each position in the reconstructed feature sequence . A fully connected layer maps the feature dimension of the fused representation to the number of entity categories yielding the label distribution for each token. Given the strong structural constraints of the entity recognition task, this study incorporates a Conditional Random Field (CRF) as the label prediction layer in the decoding stage to enhance the model's ability to discriminate entity boundaries and categories (Zhang et al., 2024), producing the optimal label sequence for the entire sentence. This process can be expressed as:
In addition, during training, the loss function of the CRF models the entire label sequence by maximizing the log-likelihood of the ground truth label sequence. This process can be expressed as:
where denotes the ground truth entity labels, and represents the prediction loss value for the entity recognition task.
In the phase classification task, label prediction is performed on the sentence-level feature vectors . Similarly, a fully connected layer maps the fused feature dimension to the number of phase categories , yielding the class label distribution for each sentence. Since the phase classification task does not involve label structures with contextual dependencies, the cross-entropy loss function (Shen, Wang, & Zhao, 2025) is adopted for loss computation. This process can be expressed as:
where denotes the ground truth phase labels, and represents the prediction loss value for the phase classification task.
To fully leverage the complementary advantages of the dual-branch structure in multi-task learning, this study adopts a joint optimization strategy to collaboratively model and train the entity recognition and phase classification tasks simultaneously. By integrating the task losses from both branches, the model facilitates the sharing of learned knowledge and semantic enhancement, thereby effectively improving overall performance. This process can be expressed as:
where is an adjustable weighting coefficient used to balance the relative contributions of the two tasks during training. Given that the phase classification task serves as an auxiliary task in this study, is set to 0.7.
4. Experimental validation and results analysis
4.1 Dataset description
This study collects a total of 3,824 accident summary text records from the past four years (2020–2023) from the Railway Safety Supervision and Management Information System. The dataset primarily contains information such as date and time, accident code, railway line, responsible bureau, and detailed summary descriptions. The raw dataset was cleaned by removing blank entries, eliminating invalid symbols and stop words, performing paragraph segmentation, basic error correction, and encoding conversion. Certain sensitive data were replaced with symbols. This process resulted in the formation of a preprocessed dataset for pre-training. Examples of accident summary descriptions are presented in Table 5.
Example of pre-training dataset
| Number | Detailed overview |
|---|---|
| 1 | On XX, XX, 20XX, at XX:XX, Train XX of Line 6 operated by XX Bureau was traveling on the up line between Yanjin North and Linjiangxi when it collided with fallen rocks at kilometer marker K211 + 985 and came to a stop. Upon inspection, it was found that the locomotive's obstacle removal device and the right-side sand spreader were damaged. After maintenance work was completed, the train resumed operation at XX:XX, and the section was reopened at XX:XX. Cause: Due to various factors such as rainfall and earthquakes, a landslide occurred on the local road embankment, causing rocks to pile up on the road surface. Some rocks crossed the road's crash barrier and rolled down the left-hand embankment of the railway onto the tracks. Upon discovering the anomaly, the on-site inspection personnel alerted the driver. However, the train's braking distance was insufficient, and a collision occurred during the braking process |
| 2 | On XX, XX, 20XX, at XX:XX, the XX freight train on the XX line operated by the XX Bureau was traveling between Mengdu and Sanyuanba when it collided with fallen rocks and came to a stop at K157 + 303, affecting four freight trains. Cause: Due to recent continuous rainfall, the soil on the natural slope at the top of the first-class embankment on the left side of the XX Line at K157 + 400 became saturated with water. Localized areas were eroded by rainwater, causing the soil around and beneath the boulders in the overlying layer to soften, leading to a localized landslide. Due to the close proximity, the train collided with the fallen rocks during braking |
| Number | Detailed overview |
|---|---|
| 1 | On XX, XX, 20XX, at XX:XX, Train XX of Line 6 operated by XX Bureau was traveling on the up line between Yanjin North and Linjiangxi when it collided with fallen rocks at kilometer marker K211 + 985 and came to a stop. Upon inspection, it was found that the locomotive's obstacle removal device and the right-side sand spreader were damaged. After maintenance work was completed, the train resumed operation at XX:XX, and the section was reopened at XX:XX. Cause: Due to various factors such as rainfall and earthquakes, a landslide occurred on the local road embankment, causing rocks to pile up on the road surface. Some rocks crossed the road's crash barrier and rolled down the left-hand embankment of the railway onto the tracks. Upon discovering the anomaly, the on-site inspection personnel alerted the driver. However, the train's braking distance was insufficient, and a collision occurred during the braking process |
| 2 | On XX, XX, 20XX, at XX:XX, the XX freight train on the XX line operated by the XX Bureau was traveling between Mengdu and Sanyuanba when it collided with fallen rocks and came to a stop at K157 + 303, affecting four freight trains. Cause: Due to recent continuous rainfall, the soil on the natural slope at the top of the first-class embankment on the left side of the XX Line at K157 + 400 became saturated with water. Localized areas were eroded by rainwater, causing the soil around and beneath the boulders in the overlying layer to soften, leading to a localized landslide. Due to the close proximity, the train collided with the fallen rocks during braking |
Since the field of railway safety supervision covers multiple professional subfields such as engineering, electrical engineering, power supply, rolling stock, and locomotives, it has strong industry characteristics. Currently, there are no publicly available entity annotation datasets covering this scenario. Therefore, this paper constructs a specialized dataset based on railway accident overview texts and annotates the texts using the BIO tagging method. An example of the stage annotation is shown in Table 6.
Examples of entity annotation and phase annotation
| Original text | Due to rockfall, the train collided with the stationary vehicle | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Entity Annotation | O | O | B-Initial Event | I- Initial Event | I- Initial Event | I- Initial Event | O | O | O | O | O | B-Intermediate Event | 0- Intermediate Event |
| Stage Annotation | Initial Event Stage | Intermediate Event Stage | |||||||||||
| Original text | Due to rockfall, the train collided with the stationary vehicle | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Entity Annotation | O | O | B-Initial Event | I- Initial Event | I- Initial Event | I- Initial Event | O | O | O | O | O | B-Intermediate Event | 0- Intermediate Event |
| Stage Annotation | Initial Event Stage | Intermediate Event Stage | |||||||||||
4.2 Evaluation metrics description
To evaluate the performance of the proposed model, three metrics: precision , recall and score, are employed. These can be expressed as:
where (True Positive)denotes the number of samples correctly predicted as positive, (False Positive)denotes the number of negative samples incorrectly predicted as positive, and (False Negative)denotes the number of positive samples incorrectly predicted as negative.
4.3 Ablation study
To comprehensively evaluate the specific contributions of each key module to the model's entity extraction performance, this study builds a multi-task fusion model by incorporating the accident phase classification task, mutual learning, and feature alignment mechanisms on top of entity recognition. To further analyze the role of each component in improving overall performance, a series of ablation experiments are designed. By progressively adding phase classification, the mutual learning mechanism, and the feature alignment module, the impact of each structural enhancement on entity extraction effectiveness is systematically examined, thereby verifying the effectiveness of the model design and the synergistic gains of its components.
The structural configurations of the ablation experiments are shown in Table 7. Model 1 is the baseline model (BERT + ER), which only includes the entity recognition task branch and does not incorporate any auxiliary information. Model 2 introduces the accident phase classification (APC) task on top of the baseline model, forming a dual-branch structure. The ER and APC tasks share the BERT encoder to achieve the sharing of basic semantic features. Furthermore, Model 3 incorporates a mutual learning mechanism (MLM) into the dual-task architecture to promote high-level semantic information interaction and collaborative optimization between the two tasks. Finally, Model 4 (R-MLP) introduces a feature alignment module (FAM) on the aforementioned structure, achieving unified alignment between different tasks at the semantic granularity level through sentence-level segmentation of the dual branches and construction of segment mappings, thereby enhancing feature fusion and context modeling capabilities.
Structural differences of models in ablation study
| Model | Description | Stage classification task | Mutual learning mechanism | Feature alignment |
|---|---|---|---|---|
| Model 1 | BERT + CRF(Baseline) | ✕ | ✕ | ✕ |
| Model 2 | Baseline + APC | ✓ | ✕ | ✕ |
| Model 3 | Baseline + APC + MLM | ✓ | ✓ | ✕ |
| R-MLP | Baseline + APC + MLM + FAM | ✓ | ✓ | ✓ |
| Model | Description | Stage classification task | Mutual learning mechanism | Feature alignment |
|---|---|---|---|---|
| Model 1 | BERT + CRF(Baseline) | ✕ | ✕ | ✕ |
| Model 2 | Baseline + APC | ✓ | ✕ | ✕ |
| Model 3 | Baseline + APC + MLM | ✓ | ✓ | ✕ |
| R-MLP | Baseline + APC + MLM + FAM | ✓ | ✓ | ✓ |
The results of the model ablation experiments are shown in Table 8, illustrating the performance changes after gradually introducing structural modules based on the Baseline model. The initial model BERT + CRF achieved an F1 score of 0.564, relying solely on basic sequence modeling capabilities, resulting in limited entity extraction performance. After introducing the accident stage classification task to construct a dual-branch structure, the F1 score of Model 2 improved to 0.645, indicating that auxiliary tasks can enhance entity recognition, but there is still a lack of deep semantic synergy between tasks. Model 3 further incorporates a mutual learning mechanism to achieve high-dimensional semantic interaction, but due to unaligned features, the F1 score drops back to 0.586, reflecting that semantic fusion imbalance may introduce noise. Finally, R-MLP introduced a feature alignment module on this basis to achieve semantic granularity unification and task collaboration, significantly improving the F1 score to 0.736, validating the key role of the alignment mechanism in multi-task mutual learning. Additionally, with structural improvements, the model's accuracy and recall rates both improved, indicating a comprehensive enhancement in overall recognition performance. By combining a dual-branch structure with mutual learning mechanisms and feature alignment design, the model effectively leverages complementary information across multiple tasks, significantly enhancing entity extraction model performance and demonstrating the advantages of multi-task mutual learning in complex information extraction tasks.
4.4 Comparative experiment
To verify the effectiveness of the proposed model, R-MLP is compared with several BERT-based optimized models developed in recent years. The comparative experimental results of different models are shown in Table 9. The RoBERTa + CRF model achieves an F1-score of only 0.151, significantly lower than expected, because RoBERTa is well suited for intra-sentence modeling (Mei, Wu, Huang, & Ling, 2023) but struggles to handle inter-sentence semantic dependencies present in railway accident texts. The BERT + CRF and BERT + BiLSTM + CRF models show comparable performance, with F1-scores of 0.564 and 0.596, respectively; although BiLSTM improves contextual modeling capability, its effectiveness in capturing long-range dependencies is limited. The MacBERT + CRF model achieves an F1-score of 0.605, benefiting from its optimized masking strategy, which better approximates real language usage. The proposed R-MLP model achieves the highest F1-score of 0.736, substantially outperforming the other models. This improvement is attributed to its multi-task learning strategy, which jointly performs entity recognition and accident phase classification, and to the incorporation of a cross-task mutual learning mechanism, enabling effective utilization of correlated information from both tasks to enhance feature representation and improve the model's understanding of low-resource semantic units.
Experimental results of different models
| Model | P | R | F1 |
|---|---|---|---|
| RoBERTa + CRF | 0.212 | 0.117 | 0.151 |
| BERT + CRF | 0.613 | 0.522 | 0.564 |
| BERT + LSTM + CRF | 0.610 | 0.583 | 0.596 |
| MACBERT + CRF | 0.695 | 0.536 | 0.605 |
| R-MLP | 0.772 | 0.703 | 0.736 |
| Model | P | R | F1 |
|---|---|---|---|
| RoBERTa + CRF | 0.212 | 0.117 | 0.151 |
| BERT + CRF | 0.613 | 0.522 | 0.564 |
| BERT + LSTM + CRF | 0.610 | 0.583 | 0.596 |
| MACBERT + CRF | 0.695 | 0.536 | 0.605 |
| R-MLP | 0.772 | 0.703 | 0.736 |
To further assess the recognition capability of various models across different entity categories, this experiment conducts a detailed comparative analysis on the identification performance of six types of railway accident–related entities. The F1-scores of each model for different entity categories are compared in Figure 7. The RoBERTa + CRF model performs worst overall, lacking adequate contextual modeling capability. BERT + CRF shows notable improvements for labels such as “Causative Factors” and “Disposal Measures,” benefiting from CRF's ability to model label dependencies. BERT + BiLSTM + CRF achieves better performance on “Intermediate Events,” reflecting its stronger context awareness. MacBERT + CRF further improves on “Causative Factors” and “Accident Consequences,” owing to its optimized pre-training strategy. The proposed R-MLP outperforms all other models across all labels, with particularly significant gains in “Initial Events” and “Control Measures,” demonstrating its strong adaptability to complex contexts and diverse expressions.
The clustered vertical bar graph has two axes. The vertical axis is labeled “F 1 score” and ranges from 0.0 to 1.0 in increments of 0.2. The horizontal axis contains six category groups: Causative Factors, Initial Event, Disposal Measures, Control Measures, Intermediate Event, and Accident Consequences. In the graph, there are six clusters of five bars in distinct colors. The bars are labeled in the legend as “R o B E R T a plus C R F” in blue, “B E R T plus C R F” in orange, “B E R T plus B i L S T M plus C R F” in green, “M a c B E R T plus C R F” in red, and “R-M L P” in purple. For each research question category, the five bars are placed side by side, showing the F1-score values for each model. The approximate F 1 score values for each configuration from left to right are as follows: Causative Factors: R o B E R T a plus C R F: 0.122, B E R T plus C R F: 0.645, B E R T plus B i L S T M plus C R F: 0.683, M a c B E R T plus C R F: 0.76, R-M L P: 0.787. Initial Event: R o B E R T a plus C R F: 0.094, B E R T plus C R F: 0.46, B E R T plus B i L S T M plus C R F: 0.481, M a c B E R T plus C R F: 0.46, R-M L P: 0.672. Disposal Measures: R o B E R T a plus C R F: 0.129, B E R T plus C R F: 0.652, B E R T plus B i L S T M plus C R F: 0.659, M a c B E R T plus C R F: 0.659, R-M L P: 0.746. Control Measures: R o B E R T a plus C R F: 0.077, B E R T plus C R F: 0.265, B E R T plus B i L S T M plus C R F: 0.153, M a c B E R T plus C R F: 0.247, R-M L P: 0.533. Intermediate Event: R o B E R T a plus C R F: 0.105, B E R T plus C R F: 0.397, B E R T plus B i L S T M plus C R F: 0.512, M a c B E R T plus C R F: 0.467, R-M L P: 0.631. Accident Consequences: R o B E R T a plus C R F: 0.188, B E R T plus C R F: 0.425, B E R T plus B i L S T M plus C R F: 0.338, M a c B E R T plus C R F: 0.516, R-M L P: 0.585. Note: All numerical data values are approximated.Comparison of F1 scores across different entity categories for each model. Source(s): Authors’ own work
The clustered vertical bar graph has two axes. The vertical axis is labeled “F 1 score” and ranges from 0.0 to 1.0 in increments of 0.2. The horizontal axis contains six category groups: Causative Factors, Initial Event, Disposal Measures, Control Measures, Intermediate Event, and Accident Consequences. In the graph, there are six clusters of five bars in distinct colors. The bars are labeled in the legend as “R o B E R T a plus C R F” in blue, “B E R T plus C R F” in orange, “B E R T plus B i L S T M plus C R F” in green, “M a c B E R T plus C R F” in red, and “R-M L P” in purple. For each research question category, the five bars are placed side by side, showing the F1-score values for each model. The approximate F 1 score values for each configuration from left to right are as follows: Causative Factors: R o B E R T a plus C R F: 0.122, B E R T plus C R F: 0.645, B E R T plus B i L S T M plus C R F: 0.683, M a c B E R T plus C R F: 0.76, R-M L P: 0.787. Initial Event: R o B E R T a plus C R F: 0.094, B E R T plus C R F: 0.46, B E R T plus B i L S T M plus C R F: 0.481, M a c B E R T plus C R F: 0.46, R-M L P: 0.672. Disposal Measures: R o B E R T a plus C R F: 0.129, B E R T plus C R F: 0.652, B E R T plus B i L S T M plus C R F: 0.659, M a c B E R T plus C R F: 0.659, R-M L P: 0.746. Control Measures: R o B E R T a plus C R F: 0.077, B E R T plus C R F: 0.265, B E R T plus B i L S T M plus C R F: 0.153, M a c B E R T plus C R F: 0.247, R-M L P: 0.533. Intermediate Event: R o B E R T a plus C R F: 0.105, B E R T plus C R F: 0.397, B E R T plus B i L S T M plus C R F: 0.512, M a c B E R T plus C R F: 0.467, R-M L P: 0.631. Accident Consequences: R o B E R T a plus C R F: 0.188, B E R T plus C R F: 0.425, B E R T plus B i L S T M plus C R F: 0.338, M a c B E R T plus C R F: 0.516, R-M L P: 0.585. Note: All numerical data values are approximated.Comparison of F1 scores across different entity categories for each model. Source(s): Authors’ own work
4.5 Case analysis
To investigate the entity extraction results and their differences, this study conducts a case analysis, with the extraction performance of different models presented in Table 10. The RoBERTa + CRF model shows a high misclassification rate and struggles to handle inter-sentence dependencies and domain-specific terminology. BERT + CRF and its BiLSTM variant achieve relatively complete extraction for fields such as “Control Measures” and “Disposal Measures,” but perform poorly on “Intermediate Events.” MacBERT + CRF performs well in “Causative Factors” and “Intermediate Events,” but its extraction of “Collapse” in the “Initial Event” category is overly generalized, lacking semantic refinement. In contrast, R-MLP achieves the best performance across all six entity types, accurately extracting complex intermediate events such as “rocks rolling onto the track” and producing outputs that closely align with human annotations in multiple key fields. Although its performance on “Intermediate Events” still has considerable room for improvement, R-MLP significantly outperforms other models. This advantage stems from its dual-branch architecture and mutual learning mechanism, which leverage the contextual information provided by the accident phase classification task to enhance the semantic perception capability of entity recognition, effectively addressing challenges such as sample scarcity, semantic complexity, and strong inter-sentence dependencies.
Entity extraction performance of different models
| Model | Causative factors | Initial event | Intermediate event | Control measures | Accident consequences | Disposal measures |
|---|---|---|---|---|---|---|
| Human Annotation | Rainfall, Earthquake | Slope Collapse | Rocks Rolling onto the Track | Warning to the Driver | Striking Fallen Rocks | Maintenance Handling |
| RoBERTa + CRF | Rainfall | Earthquake | / | Warning | Striking Fallen Rocks | Detecting Anomaly |
| BERT + CRF | Slope Collapse | Collapse | Crossing the Crash Barrier | Warning to the Driver | Stop | Maintenance Handling |
| BERT + LSTM + CRF | Earthquake | Highway Slope Collapse | Rocks Covering the Highway | Warning to the Driver | Striking Fallen Rocks and Stopping | Maintenance Handling |
| MACBERT + CRF | Rainfall, Earthquake | Collapse | Rolling onto the Track | Warning | Stop | Maintenance Handling |
| R-MLP | Rainfall, Earthquake | Slope Collapse | Rolling onto the Track | Warning to the Driver | Striking Fallen Rocks | Maintenance Handling |
| Model | Causative factors | Initial event | Intermediate event | Control measures | Accident consequences | Disposal measures |
|---|---|---|---|---|---|---|
| Human Annotation | Rainfall, Earthquake | Slope Collapse | Rocks Rolling onto the Track | Warning to the Driver | Striking Fallen Rocks | Maintenance Handling |
| RoBERTa + CRF | Rainfall | Earthquake | / | Warning | Striking Fallen Rocks | Detecting Anomaly |
| BERT + CRF | Slope Collapse | Collapse | Crossing the Crash Barrier | Warning to the Driver | Stop | Maintenance Handling |
| BERT + LSTM + CRF | Earthquake | Highway Slope Collapse | Rocks Covering the Highway | Warning to the Driver | Striking Fallen Rocks and Stopping | Maintenance Handling |
| MACBERT + CRF | Rainfall, Earthquake | Collapse | Rolling onto the Track | Warning | Stop | Maintenance Handling |
| R-MLP | Rainfall, Earthquake | Slope Collapse | Rolling onto the Track | Warning to the Driver | Striking Fallen Rocks | Maintenance Handling |
Note(s): Original text: Train XX came to a stop at the K211 + 985 mark on the up line due to a collision with fallen rocks. Upon inspection, it was found that the locomotive’s obstacle removal device and the right-side sand spreader were damaged. After maintenance work was completed, the train departed at 8:20 AM, and the section was reopened at 9:21 AM. Cause: Due to various factors such as rainfall and earthquakes, a landslide occurred on the local road embankment, causing rocks to pile up on the road surface. Some rocks crossed the road's crash barrier and rolled down the left-hand embankment of the railway onto the tracks. Upon discovering the anomaly, the on-site inspection personnel alerted the driver. However, the train's braking distance was insufficient, and a collision occurred during the braking process
5. Conclusion
This study focuses on the entity extraction problem in railway accident texts, which present challenges such as strong domain specificity, complex semantic structures, and sparse entity distribution. To address these issues, a dual-branch mutual learning model, R-MLP, based on multi-task learning is proposed. The model employs a shared BERT encoder and jointly performs entity recognition and accident phase classification. A sentence span indexing module is introduced to resolve feature alignment issues arising from differences in task granularity, while a cross-task mutual learning mechanism is designed to achieve deep semantic fusion. Incorporating accident phase classification as an auxiliary task in entity recognition enriches the semantic context and helps the model capture stage-specific characteristics in the accident evolution process, thereby improving the accuracy and stability of entity recognition and extraction. Experimental results demonstrate that the proposed model significantly outperforms mainstream baselines in extracting multiple key entity categories, achieving a maximum F1-score of 0.736, and exhibiting stronger generalization and robustness. Ablation studies further validate the effectiveness of multi-task collaborative learning and cross-task mutual learning in enhancing entity extraction performance, providing a feasible and efficient solution for information extraction from domain texts with complex structures.

