With the rapid and explosive development of the new energy vehicle industry, power battery recycling has evolved into a vital component in fulfilling resource circulation and the “dual carbon” goals. The current industry faces challenges, including poor coordination between market entities and unbalanced profit distribution. This research aims to improve recycling efficiency through optimised reward mechanisms.
This study presents a tripartite evolutionary game model. Manufacturers, third-party recyclers and cascade utilisation enterprises are the key participants. It takes government subsidies and reward-penalty mechanisms within the supply chain into account. It examines the evolutionary rules of strategy selection for each participant, along with the stability of the system. Then it carries out simulation validation.
Static reward-penalty mechanism results in cyclical fluctuation within the system and impedes stable cooperation. Linear dynamic reward and penalty mechanisms provide some motivation for manufacturers and cascade utilisation enterprises but fail to impose effective incentives or constraints on third-party recyclers. The nonlinear dynamic reward-penalty mechanism, by dynamically linking reward and penalty intensity to strategy selection, more effectively guides third-party recyclers to improve their processing standards while significantly enhancing manufacturers' recycling willingness, thereby achieving relatively ideal system equilibrium.
The research incorporates a nonlinear dynamic reward and penalty mechanism within supply chain incentives. The method modulates the severity of penalties based on participants' strategic decisions, promoting stable collaborative partnerships among them.
1. Introduction
With the explosive worldwide growth of the new energy vehicle industry, the amount of retired power batteries is increasing. Their large-scale recycling and efficient utilisation have become critical to resource circulation and achieving the “dual carbon” goals. By 2040, China is expected to generate 1.5–3.3 million tons of discarded EV batteries (Jiang et al., 2021). Their efficient recycling and utilisation are critical for resource circulation and the “dual carbon” goals. Improper handling causes heavy metal contamination and waste. A closed-loop supply chain with cascade utilisation enhances resource efficiency. However, the industry lacks standardisation, suffers from poor coordination and imbalanced profit distribution. The complex interplay of interests across the industrial chain necessitates scientific policy design and supply chain coordination mechanisms to promote healthy industry development. Against this backdrop, examining the strategic interactions and reward mechanisms among multiple stakeholders within the power battery recycling system holds significant importance for optimising resource allocation and balancing economic and environmental benefits.
Facing these complex challenges, the integrated management model of the power battery closed-loop supply chain provides new insights into how to resolve industrial bottlenecks. Some scholars have focused on the decision-making and coordination mechanisms among supply chain members. Ma et al. (2026) found that higher consumer sensitivity to cascaded product performance encourages better design standards and greater recycling participation. Yu and Wang (2025) focused on the synergistic impact of cascaded utilisation and the Extended Producer Responsibility (EPR) system on decision-making, demonstrating that enhancing data transparency and process verifiability improves operational efficiency across the entire chain. Yan et al. (2024) examined cascade utilisation and EPR systems via three pricing models, concluding that lenient EPR regulation is ineffective in governing waste battery disposal during high-revenue market stages. Further exploring the balance between competition and collaboration, Hu et al. (2025) adopted a non-cooperative-cooperative hybrid game to examine pricing and profit distribution in cascade utilisation, attempting to establish a balance between competition and collaboration. Xu et al. (2023a) analysed profits and revenue, showing that applying low-carbon innovation in cascaded-use manufacturing increases profit growth. Furthermore, factors such as fairness concerns and corporate social responsibility (CSR) have been explored. Zhang and Liang (2023) proposed a contractual mechanism that coordinates CSR investment while mitigating fairness concerns, which were shown to hinder supply chain efficiency. Tian et al. (2022) explored recycling coordination and selection strategies while taking demand uncertainties and CSR into consideration, analysing optimal strategies under different recycling modes.
Government subsidies play an important part in power battery recycling. Related research examines how these policies influence supply chain entities' behaviour and overall operations from a variety of perspectives. Wu et al. (2025) investigated recycling models that maximise overall corporate or supply chain benefits under government subsidies and extended producer responsibility schemes. They compared manufacturer-led, retailer-led, third-party, and joint recycling approaches. Wen et al. (2025) investigated the feasibility of online recycling networks and government funding options, and explored how governments might design subsidy policies to incentivise online recycling and promote cascading utilisation. Zhang et al. (2023a) compared manufacturer subsidies, recycler subsidies, and no subsidies, analysing how the recipient type affects green technology investment, recycling outcomes, and environmental benefits. Examining behavioural preferences, Liu and Zhu (2024) found that battery capacity determines optimal subsidy policy choices and that preferences may hinder recycling efficiency. Taking a long-term view, Yu and Hou (2023) constructed a differential game model and found that the impact of cost subsidy policies on supply chain profits strengthens over time.
As policy research deepens, scholars increasingly focus on the differentiated effects of various policy tools, comparing reward-penalty mechanisms with traditional subsidies. Yang et al. (2022) contrasted no intervention, subsidies, and reward-penalty mechanisms, finding reward-penalty mechanisms more effective than subsidies at incentivising power battery recycling and increasing recovery rates. Tang et al. (2019) examined how reward-penalty and subsidy schemes affected the promotion of power battery recycling. Their findings show these mechanisms have more noticeable effects on raising recycling rates and social welfare. Zhang et al. (2022) constructed three models under government reward-penalty frameworks, demonstrating that automakers implementing reward-penalty policies can maximise both recycling rates and supply chain profits. Wei and Qi (2025) compared reward-penalty and subsidy mechanisms, finding that reward-penalty mechanisms are more effective in motivating recycling behaviour and improving recycling rates. Zhang et al. (2023b) built and compared several closed-loop supply chain models. Their research explores how government subsidies, deposit-refund systems, and reward-penalty policies affect recycling rates.
In summary, the power battery recycling in China has progressed, but the regulatory structure remains incomplete. Existing research mainly focuses on analysing the impact of governmental policies on the power battery recycling, and the reward-plenty mechanism designs are centred on static subsidies. Designing reward mechanisms within the supply chain under government subsidy frameworks becomes a key area for future research.
With this context, the main contents of this study are structured as follows: First, a tripartite evolutionary game model is created. It includes manufacturers, third-party recyclers, and cascade utilisation enterprises. Government subsidies for manufacturers are included in the model. By designing internal supply chain incentive mechanisms, the dynamic evolutionary patterns of each party's strategy selection are analysed. It investigates the equilibrium point stability conditions using Lyapunov stability theory. Simulate the evolutionary path of each game participant. It reveals cyclical fluctuations under the static reward-penalty mechanism. Second, optimisation techniques for dynamic reward and penalty mechanisms are devised. They address the limitations of static mechanisms. These methods investigate how different reward-penalty combinations influence system stability. Finally, the nonlinear reward-penalty mechanism is used to improve the dynamic reward-penalty control method. This approach supports the sustainable development of resource circulation systems under the dual carbon goals.
2. Scenario analysis and model construction
The recycling process for power batteries encompasses two key stages: cascade utilisation and dismantling recycling. This establishes a closed-loop supply chain system centred on three core nodes: manufacturers, third-party recyclers, and cascade utilisation enterprises. All these entities are dedicated to the recovery and reuse of power batteries. Under this framework, manufacturers employ differentiated pricing strategies to determine whether recycled battery materials should be utilised for remanufacturing, subsequently deploying newly produced power batteries.
According to the Industry Specification for Comprehensive Utilisation of Waste Power Batteries (Ministry of Industry and Information Technology et al., 2024) (hereinafter referred to as the Specification), recyclers achieving technical benchmarks such as lithium recovery rates of no less than 90% during smelting, electrode powder recovery rates of no less than 98% after crushing and separation, and impurity aluminium content below 1.5% shall be deemed to have implemented high-level processing. Third-party recyclers gather retired power batteries from the electric vehicle market. They consider whether to implement technological innovations to enhance the quality standards of processed remanufactured batteries. Third-party recyclers then sell remanufactured batteries of varying processing standards to cascade utilisation enterprises. They are responsible for collecting all waste batteries that have completed their cascade use cycle and delivering them to manufacturers for subsequent processing.
According to the Specification, cascade utilisation operators shall be deemed to engage in proactive cascade utilisation if the annual volume of cascade-utilised waste power batteries reaches no less than 60% (by weight) (Ministry of Industry and Information Technology et al., 2024) of the actual volume of waste power batteries recovered. Manufacturers, after receiving the waste batteries, hand them over to specialised dismantling facilities for processing.
To encourage coordination and long-term growth of this closed-loop supply chain, the government pays subsidies to manufacturers. Manufacturers design and implement reward and penalty mechanisms within the supply chain. These techniques are intended to stimulate active participation from third-party recyclers and cascade utilisation enterprises. Figure 1 illustrates the logical relationships among all three participants in the evolutionary game.
A flowchart illustrating the logic relationships of a power battery closed-loop supply chain. The government provides subsidies to manufacturers and implements the EPR system. The manufacturer is involved in the green disposal of discarded batteries and the payment of environmental treatment fees. The manufacturer also engages in recycled material remanufacturing and new material manufacturing, supplying the electric vehicle market. The cascade utilization enterprise processes recycling waste batteries and engages in active or passive cascade utilization. The cascade utilization enterprise sends recycled waste batteries back to the manufacturer and processes recycling batteries after cascade utilization to the third-party recycler. The third-party recycler recycles power batteries and sends them back to the manufacturer. The manufacturer reduces or does not affect disassembly costs and receives rewards or penalties.Logic relationships of power battery closed-loop supply chain
A flowchart illustrating the logic relationships of a power battery closed-loop supply chain. The government provides subsidies to manufacturers and implements the EPR system. The manufacturer is involved in the green disposal of discarded batteries and the payment of environmental treatment fees. The manufacturer also engages in recycled material remanufacturing and new material manufacturing, supplying the electric vehicle market. The cascade utilization enterprise processes recycling waste batteries and engages in active or passive cascade utilization. The cascade utilization enterprise sends recycled waste batteries back to the manufacturer and processes recycling batteries after cascade utilization to the third-party recycler. The third-party recycler recycles power batteries and sends them back to the manufacturer. The manufacturer reduces or does not affect disassembly costs and receives rewards or penalties.Logic relationships of power battery closed-loop supply chain
2.1 Model assumptions
To build the game model, each participant's options for strategy and the stability of the system must be considered. This will aid in determining the optimal set of circumstances when the three participants can use socially stable methods. The following assumptions are made based on practical contexts and with reference to the relevant literature (Liu and Ma, 2021; Guan et al., 2023):
Three participants are involved: manufacturers, third-party recyclers, and cascade utilisation enterprises. All three are bounded rational individuals. Their strategy choices shift and become stable as time passes. They adjust their own strategies by observing each other's strategy choices and payoff outcomes, focusing on the dynamic evolutionary process of their group behaviour.
Manufacturers, third-party recyclers, and cascade utilisation enterprises each have two strategy options. The manufacturer's strategy set for power battery production materials is {recycled material remanufacturing, new material manufacturing}. The probability of selecting remanufacturing of recycled materials is and the probability of selecting manufacturing of new materials is ; Third-party recyclers have the following strategy set for retired battery recycling: {high-level processing, low-level processing}. The probability of selecting high-level processing is , while the probability of selecting low-level processing is . Cascade utilisation enterprises have the following strategy set for actively developing cascade utilisation products: {active cascade utilisation, passive cascade utilisation}. The probability of selecting active cascade utilisation is , while the probability of selecting passive cascade utilisation is , .
The manufacturer's sales revenue from remanufacturing recycled materials is , while the sales revenue from manufacturing new materials is . If the manufacturer chooses to remanufacture recycled materials, it needs to gather all wasted power batteries from the cascade utilisation enterprise and third-party recycler, and chemically disassemble the recycled batteries for a cost . If the manufacturer decides to manufacture additional materials, the procurement cost is .
When manufacturers opt for new material production, the government mandates battery recycling under the EPR policy, requiring handover to specialised dismantling enterprises (Xu et al., 2023b). Manufacturers bear the associated recycling and processing costs, calculated as . Should third-party recyclers perform high-level processing on recovered batteries, this partially offsets the green processing fees paid by manufacturers, with an offset ratio of .
Third-party recyclers experience profit variations influenced by cascade utilisation strategies when processing waste power batteries. If third-party recyclers implement high-level processing and cascade utilisation enterprises take an active role in cascade utilisation, recyclers earn profits . If cascade utilisation enterprises passively utilise batteries, recyclers lose additional profits due to reduced sales volume . High-level processing by recyclers causes additional costs , including technological improvements and equipment maintenance. When third-party recyclers perform low-level processing, if cascade utilisation enterprises actively engage in cascade utilisation, their profit is . If cascade utilisation enterprises passively engage in cascade utilisation, third-party recyclers lose profit . If manufacturers remanufacture recycled materials, high-level processing by recyclers partially reduces dismantling costs by factor .
The cascade utilisation market remains in its early development stage, with relatively low demand for spent battery reuse. The volume of spent batteries recovered from the electric vehicle market can satisfy cascade utilisation enterprises for active cascade utilisation. If cascade utilisation operators engage in active cascade utilisation, the profit from selling high-level processed and remanufactured batteries is , while selling low-level processed batteries results in a profit loss of . If they engage in passive cascade utilisation, the profit from selling high-level processed and remanufactured batteries is , while selling low-level processed batteries results in a profit loss of . Given poor consistency and lifespan estimation difficulties, active cascade utilisation requires additional R&D costs (Yuan and Zhang, 2025).
If manufacturers choose to use recycled materials for remanufacturing, the government provides manufacturers with reward subsidies . Manufacturers use these rewards to guide enterprises toward technological upgrades and shared environmental responsibility, avoiding low-cost, low-quality competition and promoting the greening of power batteries across the whole life cycle. Consequently, manufacturers incentivise recyclers to achieve high-level processing and share subsidies , and active cascade utilisation enterprises to actively engage in cascade utilisation and share subsidies . According to relevant regulations (Ministry of Industry and Information Technology et al., 2026), manufacturing enterprises must establish recycling service outlets within their sales regions and assume corresponding responsibility for the flow of power batteries. In line with the EPR requirements, manufacturers should translate external regulatory pressures into internal supply chain management requirements, effectively constraining non-compliant partners. Therefore, if recyclers fail to meet processing standards, manufacturers impose penalties ; if cascade utilisation enterprises demonstrate passive utilisation, they face sanctions .
2.2 Model construction
A game matrix is generated based on the model assumptions stated above. It includes manufacturers, third-party recyclers, and cascade utilisation enterprises. The matrix is detailed in Table 1.
Tripartite game matrix for cascaded utilisation of retired power batteries
| Third-party recycler | Cascade utilisation enterprise | |||
|---|---|---|---|---|
| Active cascade Utilisation | Passive cascade utilisation | |||
| Battery Manufacturer | Recycled Material Remanufacturing | High-Level Processing | , , | , , |
| Low-Level Processing | , , | , , | ||
| New Material Manufacturing | High-level Processing | , , | , , | |
| Low-level processing | , , | , , | ||
| Third-party recycler | Cascade utilisation enterprise | |||
|---|---|---|---|---|
| Active cascade | Passive cascade utilisation | |||
| Battery Manufacturer | Recycled Material Remanufacturing | High-Level Processing | ||
| Low-Level Processing | ||||
| New Material Manufacturing | High-level Processing | |||
| Low-level processing | ||||
Construct replication dynamic equation functions based on the game matrix involving manufacturers, third-party recyclers, and cascade utilisation enterprises.
2.2.1 Manufacturer dynamic equation function construction
The expected benefits for manufacturers using recycled materials for remanufacturing , the expected returns for using new materials in manufacturing and the average expected returns are, respectively:
The dynamic replication equation for the manufacturer's chosen strategy is produced by the system of equations:
2.2.2 Third-party recyclers' dynamic equation function construction
The expected benefits for third-party recyclers using high-level processing , the expected returns for low-level processing by third-party recyclers and the average expected returns are, respectively:
For third-party recyclers, the dynamics replication equation is:
2.2.3 Cascade utilisation enterprise dynamic equation function construction
The expected benefits for active and passive cascade utilisation by cascade utilisation enterprises are and , and the average expected benefits are, respectively:
The dynamic replication equation for cascade utilisation enterprise strategy selection is given below:
2.3 Equilibrium point stability analysis for the tripartite evolutionary game model
From , , and , the equilibrium points of the system are obtained as: , , , , , , , and . The Jacobian matrix is expressed by the equation J below:
According to Lyapunov's indirect method, the stability of an equilibrium point is determined by the sign of the eigenvalues of its Jacobian matrix: if all eigenvalues are negative, the equilibrium is asymptotically stable; conversely, if at least one eigenvalue is positive, it is unstable. For the purpose of facilitating Stability analysis, set , , , . The eigenvalues of the Jacobian matrix corresponding to each equilibrium point are listed in Table 2.
Equilibrium point stability evaluation
| Equilibrium point | Eigenvalue | Sign | Stability condition |
|---|---|---|---|
| Uncertain point; Stable for | |||
| Uncertain point; Stable for | |||
| Saddle point | |||
| Saddle point | |||
| Uncertain point; Stable for | |||
| Uncertain point; Stable for | |||
| Saddle point | |||
| Uncertain point; Stable for |
| Equilibrium point | Eigenvalue | Sign | Stability condition |
|---|---|---|---|
| Uncertain point; Stable for | |||
| Uncertain point; Stable for | |||
| Saddle point | |||
| Saddle point | |||
| Uncertain point; Stable for | |||
| Uncertain point; Stable for | |||
| Saddle point | |||
| Uncertain point; Stable for |
Note(s): + for positive sign; − for negative sign; unknown signs marked
2.4 Simulation analysis
To verify the validity of the evolutionary stability analysis, the model is parameterised with numerical values based on realistic conditions and relevant literature (Wu and Zhang, 2025; Zhang et al., 2024a, b). Table 3 Initial Parameter Settings presents the initial parameter settings.
Initial parameter settings
| Parameter | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Initial Value | 173 | 221 | 81 | 24 | 59 | 44 | 64 | 11 | 7 | 20 | 30 |
| Parameters | |||||||||||
| Initial Value | 14 | 20 | 3 | 2 | 5 | 0.8 | 0.9 | 20 | 10 | 5 | 10 |
| Parameter | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Initial Value | 173 | 221 | 81 | 24 | 59 | 44 | 64 | 11 | 7 | 20 | 30 |
| Parameters | |||||||||||
| Initial Value | 14 | 20 | 3 | 2 | 5 | 0.8 | 0.9 | 20 | 10 | 5 | 10 |
Using MATLAB R2024a software, data simulations were conducted on the evolutionary track of each game entity. Figure 2 displays the tripartite evolutionary game's initial results.
A three-dimensional line graph with multiple colored lines forming a spiral pattern. The graph features three axes labeled x, y, and z, each ranging from 0 to 1. The lines converge towards the center of the graph, creating a dense, colorful spiral. The graph appears to illustrate a complex, dynamic system with multiple trajectories.Initial tripartite evolutionary game path
A three-dimensional line graph with multiple colored lines forming a spiral pattern. The graph features three axes labeled x, y, and z, each ranging from 0 to 1. The lines converge towards the center of the graph, creating a dense, colorful spiral. The graph appears to illustrate a complex, dynamic system with multiple trajectories.Initial tripartite evolutionary game path
The system's evolutionary game path forms a closed curve with periodic motion around a stable centre point, lacking any stable equilibrium points. When initial values for x, y, z are set to (0.5, 0.5, 0.5) and (0.2, 0.5, 0.8) respectively, the system's evolutionary paths are shown in Figures 3 and 4. The strategic choices of third-party recyclers have gradually trended towards low-level treatment over time, and the strategic choices of manufacturers and cascade utilisation enterprises exhibit a pattern of continuous fluctuation. This demonstrates that under the current model settings, the interplay of interests among various entities is rather complex. Further optimisation of the reward and penalty mechanisms or adjustment of other relevant parameters is required to improve the system's evolutionary trajectory.
A line graph displays the evolutionary path with initial values of 0.5, 0.5, and 0.5. The horizontal axis represents time (t) ranging from 0 to 10, and the vertical axis represents the variable p ranging from 0 to 1. There are three distinct lines: a blue line with star markers representing x equals 0.2, a green line with square markers representing y equals 0.5, and a pink line with diamond markers representing z equals 0.8. The blue and pink lines show oscillatory behavior with peaks and troughs, while the green line starts high and quickly drops to near zero, remaining flat for the rest of the time period. All values are approximated.Evolutionary path with initial values (0.5, 0.5, 0.5)
A line graph displays the evolutionary path with initial values of 0.5, 0.5, and 0.5. The horizontal axis represents time (t) ranging from 0 to 10, and the vertical axis represents the variable p ranging from 0 to 1. There are three distinct lines: a blue line with star markers representing x equals 0.2, a green line with square markers representing y equals 0.5, and a pink line with diamond markers representing z equals 0.8. The blue and pink lines show oscillatory behavior with peaks and troughs, while the green line starts high and quickly drops to near zero, remaining flat for the rest of the time period. All values are approximated.Evolutionary path with initial values (0.5, 0.5, 0.5)
A line graph displays the evolutionary path with initial values of 0.2, 0.5, and 0.8. The horizontal axis represents time (t) ranging from 0 to 10, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines: a blue line with star markers for x equals 0.2, a green line with square markers for y equals 0.5, and a magenta line with circle markers for z equals 0.8. The blue line shows a periodic pattern with peaks around 0.9 and troughs around 0.1. The magenta line also shows a periodic pattern but with slightly lower peaks around 0.85 and higher troughs around 0.3. The green line remains close to zero throughout the time range.Evolutionary path with initial values (0.2, 0.5, 0.8)
A line graph displays the evolutionary path with initial values of 0.2, 0.5, and 0.8. The horizontal axis represents time (t) ranging from 0 to 10, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines: a blue line with star markers for x equals 0.2, a green line with square markers for y equals 0.5, and a magenta line with circle markers for z equals 0.8. The blue line shows a periodic pattern with peaks around 0.9 and troughs around 0.1. The magenta line also shows a periodic pattern but with slightly lower peaks around 0.85 and higher troughs around 0.3. The green line remains close to zero throughout the time range.Evolutionary path with initial values (0.2, 0.5, 0.8)
We conducted a sensitivity analysis on key parameters. By setting the penalty parameters for manufacturers towards third-party recyclers to respectively, and those for cascade utilisation to respectively, the effects of these parameter changes on the system's evolution are illustrated in Figures 5 and 6. The impact of changes in manufacturers' penalty on third-party recyclers. When the penalty level is low, the volatility of the system increases. When the penalty level imposed on third-party recyclers is high, cascade utilisation enterprises reach a stable state, but manufacturers and third-party recyclers remain in a volatile state. When the penalty level imposed on cascade utilisation enterprises is high, third-party recyclers remain in a stable state characterised by a tendency towards low-level recycling, whilst the volatility of manufacturers and cascade utilisation enterprises intensifies.
The image contains three line graphs side by side, each representing different values of n (10, 20, and 40) and their respective x, y, and z coordinates over time t. The x-axis represents time t ranging from 0 to 10, and the y-axis represents the probability p ranging from 0 to 1. Each graph shows three lines: blue for x, green for y, and magenta for z. The lines exhibit periodic behavior with varying amplitudes and frequencies. For n equals 10, the blue line (x) shows the highest amplitude, followed by the magenta line (z), while the green line (y) remains close to zero. For n equals 20, the blue and magenta lines (x and z) show similar amplitudes with more frequent oscillations, while the green line (y) remains close to zero. For n equals 40, the blue line (x) shows the highest amplitude, followed by the magenta line (z), while the green line (y) shows more pronounced oscillations compared to the previous graphs. All values are approximated.The impact of changes in manufacturers' penalty on third-party recyclers
The image contains three line graphs side by side, each representing different values of n (10, 20, and 40) and their respective x, y, and z coordinates over time t. The x-axis represents time t ranging from 0 to 10, and the y-axis represents the probability p ranging from 0 to 1. Each graph shows three lines: blue for x, green for y, and magenta for z. The lines exhibit periodic behavior with varying amplitudes and frequencies. For n equals 10, the blue line (x) shows the highest amplitude, followed by the magenta line (z), while the green line (y) remains close to zero. For n equals 20, the blue and magenta lines (x and z) show similar amplitudes with more frequent oscillations, while the green line (y) remains close to zero. For n equals 40, the blue line (x) shows the highest amplitude, followed by the magenta line (z), while the green line (y) shows more pronounced oscillations compared to the previous graphs. All values are approximated.The impact of changes in manufacturers' penalty on third-party recyclers
Three line graphs depict the impact of changes in manufacturers' penalty on cascade utilisation enterprises. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different frequencies (f) and variables (x, y, z). Panel A shows the impact for f=5, Panel B for f=10, and Panel C for f=20. Each panel contains three lines representing different variables: x in blue, y in green, and z in magenta. The lines show oscillatory behavior with varying amplitudes and frequencies. In Panel A, the blue line (f=5, x) has the highest amplitude, followed by the magenta line (f=5, z), while the green line (f=5, y) remains constant at zero. In Panel B, the blue line (f=10, x) and the magenta line (f=10, z) show similar amplitudes, while the green line (f=10, y) remains constant at zero. The graphs illustrate how the penalty changes affect the utilization enterprises over time.The impact of changes in manufacturers' penalty on cascade utilisation enterprises
Three line graphs depict the impact of changes in manufacturers' penalty on cascade utilisation enterprises. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different frequencies (f) and variables (x, y, z). Panel A shows the impact for f=5, Panel B for f=10, and Panel C for f=20. Each panel contains three lines representing different variables: x in blue, y in green, and z in magenta. The lines show oscillatory behavior with varying amplitudes and frequencies. In Panel A, the blue line (f=5, x) has the highest amplitude, followed by the magenta line (f=5, z), while the green line (f=5, y) remains constant at zero. In Panel B, the blue line (f=10, x) and the magenta line (f=10, z) show similar amplitudes, while the green line (f=10, y) remains constant at zero. The graphs illustrate how the penalty changes affect the utilization enterprises over time.The impact of changes in manufacturers' penalty on cascade utilisation enterprises
Varying and gives the impact shown in Figures 7 and 8. When the reward parameters for third-party recyclers are altered, there is no significant change in the system's volatility. When the incentive level for cascade utilisation enterprises is low, the system's volatility increases. When the incentive level is high, the volatility of both manufacturers and cascade utilisation enterprises is significantly reduced.
The image contains three line graphs side by side, each depicting the impact of changes in manufacturers' reward on third-party recyclers. The x-axis represents time (t) ranging from 0 to 10, and the y-axis represents the probability (p) ranging from 0 to 1. Each graph shows three different scenarios labeled as Q equals 2x, Q equals 2y, and Q equals 2z for the first graph; Q equals 5x, Q equals 5y, and Q equals 5z for the second graph; and Q equals 10x, Q equals 10y, and Q equals 10z for the third graph. The lines are color-coded: blue for x, green for y, and magenta for z. The blue and magenta lines exhibit oscillatory behavior with varying amplitudes, while the green lines start at a high probability and quickly drop to zero, remaining constant thereafter. The blue and magenta lines show periodic peaks and troughs, indicating fluctuating probabilities over time. All values are approximated.The impact of changes in manufacturers' reward on third-party recyclers
The image contains three line graphs side by side, each depicting the impact of changes in manufacturers' reward on third-party recyclers. The x-axis represents time (t) ranging from 0 to 10, and the y-axis represents the probability (p) ranging from 0 to 1. Each graph shows three different scenarios labeled as Q equals 2x, Q equals 2y, and Q equals 2z for the first graph; Q equals 5x, Q equals 5y, and Q equals 5z for the second graph; and Q equals 10x, Q equals 10y, and Q equals 10z for the third graph. The lines are color-coded: blue for x, green for y, and magenta for z. The blue and magenta lines exhibit oscillatory behavior with varying amplitudes, while the green lines start at a high probability and quickly drop to zero, remaining constant thereafter. The blue and magenta lines show periodic peaks and troughs, indicating fluctuating probabilities over time. All values are approximated.The impact of changes in manufacturers' reward on third-party recyclers
Three line graphs depict the impact of changes in manufacturers' reward on cascade utilisation enterprises. Each graph shows the probability (p) on the vertical axis and time (t) on the horizontal axis. The graphs are labeled with different parameters: P equals 5, 10, and 20, with corresponding x, y, and z values. Panel A: The first line graph shows the probability (p) over time (t) for P equals 5. The blue line represents P equals 5, x, the green line represents P equals 5, y, and the magenta line represents P equals 5, z. The blue and magenta lines show a rapid increase and then a decrease, while the green line remains relatively constant. Panel B: The second line graph shows the probability (p) over time (t) for P equals 10. The blue line represents P equals 10, x, the green line represents P equals 10, y, and the magenta line represents P equals 10, z. The blue and magenta lines exhibit oscillatory behavior, while the green line remains relatively constant. Panel C:The impact of changes in manufacturers' reward on cascade utilisation enterprises
Three line graphs depict the impact of changes in manufacturers' reward on cascade utilisation enterprises. Each graph shows the probability (p) on the vertical axis and time (t) on the horizontal axis. The graphs are labeled with different parameters: P equals 5, 10, and 20, with corresponding x, y, and z values. Panel A: The first line graph shows the probability (p) over time (t) for P equals 5. The blue line represents P equals 5, x, the green line represents P equals 5, y, and the magenta line represents P equals 5, z. The blue and magenta lines show a rapid increase and then a decrease, while the green line remains relatively constant. Panel B: The second line graph shows the probability (p) over time (t) for P equals 10. The blue line represents P equals 10, x, the green line represents P equals 10, y, and the magenta line represents P equals 10, z. The blue and magenta lines exhibit oscillatory behavior, while the green line remains relatively constant. Panel C:The impact of changes in manufacturers' reward on cascade utilisation enterprises
Varying and gives the impact shown in Figures 9 and 10. When the additional investment costs for high-level recycling by third-party recyclers vary, there is no significant change in the system's volatility. When the additional investment costs associated with high-level recycling by third-party recyclers are high, the system's volatility increases. When these costs are low, the volatility of both manufacturers and cascade utilisation enterprises decreases.
Three line graphs depict the impact of cost changes among third-party recyclers. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different cost change scenarios: c21=10, c21=20, and c21=30. Each scenario is represented by three lines: x, y, and z. The lines are color‑coded: blue for x, green for y, and magenta for z. In the three sets of subgraphs, the blue line (x) and the magenta line (z) exhibit sinusoidal oscillation patterns with peaks and troughs, while the green line (y) remains close to zero. The amplitude and phase show slight differences among the three groups of graphs.The impact of cost changes among third-party recyclers
Three line graphs depict the impact of cost changes among third-party recyclers. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different cost change scenarios: c21=10, c21=20, and c21=30. Each scenario is represented by three lines: x, y, and z. The lines are color‑coded: blue for x, green for y, and magenta for z. In the three sets of subgraphs, the blue line (x) and the magenta line (z) exhibit sinusoidal oscillation patterns with peaks and troughs, while the green line (y) remains close to zero. The amplitude and phase show slight differences among the three groups of graphs.The impact of cost changes among third-party recyclers
Three line graphs depict the impact of cost changes among cascade utilisation enterprises. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different cost change scenarios: c31=2, c31=5, and c31=10. Each scenario is represented by three lines: x, y, and z. The lines are color-coded: blue for x, green for y, and magenta for z. Panel A shows the impact for c31=2. The blue line (x) and the magenta line (z) exhibit sinusoidal patterns with peaks and troughs, while the green line (y) remains close to zero. Panel B shows the impact for c31=5. The blue line (x) and the magenta line (z) again exhibit sinusoidal patterns, but with different amplitudes and phases compared to Panel A. The green line (y) remains close to zero. Panel C shows the impact for c31=10. The graphs illustrate how the variable p changes over time under different cost change scenarios.The impact of cost changes among cascade utilisation enterprises
Three line graphs depict the impact of cost changes among cascade utilisation enterprises. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different cost change scenarios: c31=2, c31=5, and c31=10. Each scenario is represented by three lines: x, y, and z. The lines are color-coded: blue for x, green for y, and magenta for z. Panel A shows the impact for c31=2. The blue line (x) and the magenta line (z) exhibit sinusoidal patterns with peaks and troughs, while the green line (y) remains close to zero. Panel B shows the impact for c31=5. The blue line (x) and the magenta line (z) again exhibit sinusoidal patterns, but with different amplitudes and phases compared to Panel A. The green line (y) remains close to zero. Panel C shows the impact for c31=10. The graphs illustrate how the variable p changes over time under different cost change scenarios.The impact of cost changes among cascade utilisation enterprises
3. Model analysis under linear dynamic reward and penalty system
The static system's evolutionary strategy exhibits fluctuating patterns. In order to develop an evolutionarily stable strategy, the fundamental model is enhanced. Dynamic reward and penalty mechanism is implemented in internal incentives. Manufacturers may regulate the optimum penalties and rewards for cascade utilisation enterprises and third-party recyclers according to their strategic decisions due to the dynamic system. The reward and penalty amounts are optimised from fixed constants , , , to dynamic linear functions. Based on different combinations of reward and penalty mechanisms, the three mechanisms are as follows: dynamic linear penalty, dynamic linear reward, and dynamic linear reward-penalty.
3.1 Model construction
3.1.1 Construction of the linear dynamic penalty mechanism
Assume manufacturers' penalty strategies toward third-party recyclers and cascade utilisation enterprises dynamically correlate with their respective strategy choices. Specifically, the manufacturer's penalty towards third-party recyclers is , and towards cascade utilisation enterprises is . and are the linear dynamic penalty coefficients, , , and reflect the chance of third-party recyclers using low-level recycling methods, and represents the probability of cascade utilisation enterprises using passive cascade utilisation strategies. The reward for cascade utilisation enterprises and third-party recyclers remains a fixed constant and .
In the expected payoffs of the game's participants, the dynamic penalty term appears as a lump sum. Third-party recyclers and cascade utilisation enterprises have and deducted respectively for their non-compliant behaviour. Meanwhile, manufacturers see a corresponding increase in their payoffs as a result of collecting the fines. The pure strategy payoff expression , , , , and , under the static mechanism remains unchanged. the dynamic reward and penalty terms are simply superimposed on the overall expected payoffs of each participant. The expected returns for each participant are obtained:
The dynamic equation is derived based on the expected return described above, as shown below:
3.1.2 Construction of the linear dynamic reward mechanism
Assume the manufacturer's reward strategies for the three participants are dynamically linked to their respective strategy choices. Specifically, the manufacturer's reward for third-party recyclers is , and for cascade utilisation enterprises is , where and are linear dynamic reward coefficients, and . Penalties for both third-party recyclers and cascade utilisation enterprises remain fixed constants and . In terms of expected returns, third-party recyclers and cascade utilisation enterprises receive additional returns of and respectively as a result of their proactive behaviour. The expected return expressions are:
The replicator dynamics equations for the three participants are, respectively:
3.1.3 Construction of the linear dynamic reward-penalty mechanism
The penalty and reward mechanisms are combined. The expected return expressions are as follows:
The replicator dynamic equations are shown below:
3.2 Stability analysis
3.2.1 Stability analysis of the linear dynamic penalty mechanism
According to differential equation stability theory, a strategy combination achieves evolutionary equilibrium when the replication dynamics equation equals zero and its derivative is less than zero. For manufacturer, the derivative of the replication dynamics equation with respect to strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the manufacturer, the system is in an evolutionarily stable state, and the strategy selection remains unchanged over time. When , is the evolutionarily stable strategy for the manufacturer, choosing to recycle materials for remanufacturing. When , is the evolutionarily stable strategy for the manufacturer choosing to manufacture with new materials.
For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: . Let , then , and . When , . At this stage, regardless of the approach chosen by the third-party recycler, it is in an evolutionary stable state, with strategy selection remaining constant throughout time. When , is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when , this is the third-party recycler's evolutionary stable method, who prefers low-level recycling.
The derivative of the cascade utilisation enterprise's replicator dynamic equation with respect to strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system is evolutionarily stable, and strategy selection remains unchanged over time. When , is the evolutionarily stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, when , this is the evolutionarily stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation.
Following Lyapunov's indirect method, the stability analysis is conducted on the equilibrium points of the system for the linear dynamic penalty mechanism. From , , and , the equilibrium points of the system are obtained as: , , , , , , , and . The equilibrium points and eigenvalues of the linear dynamic penalty mechanism are shown in Table 4.
3.2.2 Stability analysis of the linear dynamic reward mechanism
For the manufacturer, the dynamic equations are identical to those in Section 3.2.1. Consequently, the conclusions of the stability analysis for the manufacturer are exactly the same. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the third-party recycler, it is in an evolutionary stable state, with strategy selection remaining stable throughout time. When , is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when , this is the third-party recycler's evolutionary stable method, who prefers low-level recycling.
The derivative of the cascade utilisation enterprise's replicator dynamic equation with respect to strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionary stable state, and the strategy selection remains unchanged over time. When , is the evolutionary stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, is the evolutionary stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation. The equilibrium points and eigenvalues of the linear dynamic reward mechanism are shown in Table 5.
3.2.3 Stability analysis of the linear dynamic reward-penalty mechanism
Manufacturer's stability analysis is identical to Section 3.2.1. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: . Let , then , and . When , . The third-party recycler remains in an evolutionary stable state regardless of the strategy it chooses, and this state remains unchanged over time. When , is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, is the third-party recycler's evolutionary stable strategy choosing low-level recycling.
For cascade utilisation enterprise, the derivative of the replicator dynamic equation with respect to the strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionarily stable state, and the strategy selection remains unchanged over time. When , is the evolutionarily stable strategy for the cascade utilisation enterprise, choosing active cascade utilisation. Conversely, is the evolutionarily stable strategy for the cascade utilisation enterprise, indicating the selection of passive cascade utilisation. The equilibrium points and eigenvalues of the linear dynamic penalty-reward mechanism are shown in Table 6.
3.3 Model simulation analysis under linear dynamic reward and penalty system
The initial parameters of the linear dynamic reward and penalty system are the same as those of the static reward-penalty system. The dynamic penalty coefficient is set to , , and the dynamic reward coefficient is set to , . The simulation results of the system's evolution over time under the three reward and penalty mechanisms are shown in Figures 11–13, respectively.
A line graph displays the evolutionary path under a linear dynamic penalty mechanism. The x-axis represents time (t) ranging from 0 to 3, and the y-axis represents the variable p ranging from 0 to 1. The graph includes three data lines: one in blue for x equals 0.5, one in green for y equals 0.5, and one in magenta for z equals 0.5. The blue line starts at a p value of approximately 0.4 and decreases to around 0.3, stabilizing around 0.3 at t equals 2.69839. The magenta line starts at a p value of approximately 0.4 and increases to around 0.7, stabilizing around 0.7 at t equals 2.55938. The green line starts at a p value of approximately 0.4 and decreases to around 0.0, stabilizing around 0.0 at t equals 1.5. All values are approximated.Evolutionary path under linear dynamic penalty mechanism
A line graph displays the evolutionary path under a linear dynamic penalty mechanism. The x-axis represents time (t) ranging from 0 to 3, and the y-axis represents the variable p ranging from 0 to 1. The graph includes three data lines: one in blue for x equals 0.5, one in green for y equals 0.5, and one in magenta for z equals 0.5. The blue line starts at a p value of approximately 0.4 and decreases to around 0.3, stabilizing around 0.3 at t equals 2.69839. The magenta line starts at a p value of approximately 0.4 and increases to around 0.7, stabilizing around 0.7 at t equals 2.55938. The green line starts at a p value of approximately 0.4 and decreases to around 0.0, stabilizing around 0.0 at t equals 1.5. All values are approximated.Evolutionary path under linear dynamic penalty mechanism
A line graph depicts the evolutionary path under a linear dynamic reward mechanism. The horizontal axis represents time (t) ranging from 0 to 2, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines labeled x equals 0.5, y equals 0.5, and z equals 0.5. The blue line labeled x equals 0.5 starts at the origin and rapidly increases, reaching a value close to 1 around t equals 1.5. The green line labeled y equals 0.5 starts at the origin and quickly decreases to near 0, remaining flat thereafter. The magenta line labeled z equals 0.5 starts at the origin, decreases slightly, and then gradually increases to a value around 0.4, remaining relatively flat after t equals 1. A data point is marked at approximately t equals 1.57577 and p equals 0.399139 on the magenta line.Evolutionary path under linear dynamic reward mechanism
A line graph depicts the evolutionary path under a linear dynamic reward mechanism. The horizontal axis represents time (t) ranging from 0 to 2, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines labeled x equals 0.5, y equals 0.5, and z equals 0.5. The blue line labeled x equals 0.5 starts at the origin and rapidly increases, reaching a value close to 1 around t equals 1.5. The green line labeled y equals 0.5 starts at the origin and quickly decreases to near 0, remaining flat thereafter. The magenta line labeled z equals 0.5 starts at the origin, decreases slightly, and then gradually increases to a value around 0.4, remaining relatively flat after t equals 1. A data point is marked at approximately t equals 1.57577 and p equals 0.399139 on the magenta line.Evolutionary path under linear dynamic reward mechanism
A line graph displays the evolutionary path under a linear dynamic reward-penalty mechanism. The x-axis represents time (t) ranging from 0 to 4, and the y-axis represents the variable p ranging from 0 to 1. The graph includes three data lines: one in blue for x equals 0.5, one in green for y equals 0.5, and one in magenta for z equals 0.5. The blue line starts at the origin, rises steeply, and then levels off around 0.926679 at t equals 3.5875. The magenta line starts at 0.4, rises gradually, and levels off around 0.665992 at t equals 3.5875. The green line starts at 0.4, drops steeply, and then levels off around 0.0742487 at t equals 3.77748. All values are approximated.Evolutionary path under linear dynamic reward-penalty mechanism
A line graph displays the evolutionary path under a linear dynamic reward-penalty mechanism. The x-axis represents time (t) ranging from 0 to 4, and the y-axis represents the variable p ranging from 0 to 1. The graph includes three data lines: one in blue for x equals 0.5, one in green for y equals 0.5, and one in magenta for z equals 0.5. The blue line starts at the origin, rises steeply, and then levels off around 0.926679 at t equals 3.5875. The magenta line starts at 0.4, rises gradually, and levels off around 0.665992 at t equals 3.5875. The green line starts at 0.4, drops steeply, and then levels off around 0.0742487 at t equals 3.77748. All values are approximated.Evolutionary path under linear dynamic reward-penalty mechanism
Under the linear dynamic penalty mechanism, the probability of manufacturers adopting a strategy involving the use of recycled materials stabilises at 0.3, whilst the probability of third-party recyclers opting for low-level processing rapidly converges to 0. This indicates that the penalty mechanism has limited effectiveness in incentivising manufacturers to recycle and fails to effectively curb recyclers' low-level processing behaviour. The probability of cascade utilisation enterprises engaging in active cascade utilisation stabilises at 0.7, reflecting that the dynamic penalty has a relatively significant deterrent effect on the negative behaviour of such enterprises. Overall, this mechanism can only partially improve the willingness to engage in cascade utilisation and lacks effective constraints on recyclers.
Under the linear dynamic reward mechanism, the probability of manufacturers using recycled materials stabilises at 1, indicating that this dynamic reward effectively encourages manufacturers to adopt recycled materials. The probability of third-party recyclers opting for high-level processing still converges to 0, suggesting that rewards alone cannot reverse the tendency of recyclers to choose low-level processing due to cost pressures. The probability of cascade utilisation enterprises engaging in active cascade utilisation stabilises at 0.4, indicating that the incentive effect of the rewards on these enterprises is limited and fails to fully stimulate their enthusiasm. Although this mechanism has enhanced manufacturers' sense of responsibility, it has failed to coordinate other links in the industrial chain.
Under the linear dynamic reward-penalty mechanism, the probability of manufacturers using recycled materials stabilised at 0.93, approaching full adoption of recycled materials. The probability of third-party recyclers employing high-standard processing rose slightly to 0.07, but remained at an extremely low level, indicating that whilst this combined reward-penalty mechanism brought about some improvement for recyclers, the effect was not significant. The probability of cascade utilisation enterprises actively engaging in cascade utilisation stabilised at 0.67, representing an improvement compared to a single reward mechanism.
Overall, the linear dynamic reward and penalty mechanisms can encourage manufacturers to fulfil their recycling responsibilities and exert a certain degree of incentive and constraint on cascade utilisation enterprises. However, they still lack sufficient driving force to promote high-level processing behaviour among third-party recyclers.
4. Optimisation of dynamic reward and penalty control schemes
The linear dynamic reward-penalty mechanism, as previously indicated, can suppress fluctuations. However, under the three dynamic reward and penalty configurations described, third-party recyclers consistently prefer low-level processing of recovered power batteries. These are not optimal control solutions. Therefore, this study proposes optimising the dynamic reward and penalty control scheme by introducing nonlinear dynamic reward and penalty functions (Chang et al., 2017). The three mechanisms are as follows: nonlinear dynamic penalty, nonlinear dynamic reward, and nonlinear dynamic reward-penalty.
4.1 Model construction
4.1.1 Construction of the nonlinear dynamic penalty mechanism
The manufacturer's nonlinear dynamic penalty function for third-party recyclers is defined as , while the function for cascade utilisation enterprises is , where , , , are the dynamic penalty coefficients, , , and , . This nonlinear function signifies that the severity of penalties imposed on manufacturers is proportional to the additional investment costs paid for high-level processing of retired power batteries and active cascade utilisation. Specifically, the higher the probability that manufacturers choose to recycle materials for remanufacturing, the greater the penalty. Conversely, the lower the additional investment costs for third-party recyclers and cascade utilisation enterprises, the greater the penalty. The total penalty is reflected in the expected return as and . The pure strategy payoff expression , , , , and , under the static mechanism remains unchanged. The reward for cascade utilisation enterprises and third-party recyclers remains a fixed constant , . The expected return expressions are as follows:
The replicator dynamics equations for the three participants are, respectively:
4.1.2 Construction of the nonlinear dynamic reward mechanism
The manufacturer's nonlinear dynamic reward function for third-party recyclers is set as , and for cascade utilisation enterprises as , where , , , are the dynamic reward coefficients, , , , and . This nonlinear function signifies that manufacturers' reward levels correlate with the additional investment costs incurred by high-level processing of the retired power batteries and active cascade utilisation. Specifically, manufacturers receive higher rewards when they choose to recycle materials for remanufacturing, while third-party recyclers and cascade utilisation enterprises receive higher rewards when their additional investment costs increase. Third-party recyclers and cascade utilisation enterprises receive additional returns of and respectively as a result of their proactive behaviour. The penalty for third-party recyclers and cascade utilisation enterprise remains a fixed constant , . The expected return expressions are:
The replicator dynamics equations for the three participants are, respectively:
4.1.3 Construction of the nonlinear dynamic reward-penalty mechanism
The penalty and reward mechanisms are combined. The expected return expressions are as follows:
The replicator dynamics equations for the three participants are, respectively:
4.2 Stability analysis
4.2.1 Stability analysis of the nonlinear dynamic penalty mechanism
For manufacturer, stability analysis is consistent with Section 3.2.1. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the third-party recycler, it remains in an evolutionary stable state, and the strategy selection remains unchanged all the time. When , , , , and . is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when , this is the third-party recycler's evolutionary stable strategy choosing low-level recycling.
The derivative of the cascade utilisation enterprise's replicator dynamic equation with respect to strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionarily stable state, and the strategy selection remains unchanged over time. When , is the evolutionarily stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, when , this is the evolutionarily stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation. Following Lyapunov's indirect method, the equilibrium points and eigenvalues of the nonlinear dynamic penalty mechanism are shown in Table 7.
4.2.2 Stability analysis of the nonlinear dynamic reward mechanism
The manufacturer's stability analysis is consistent with Section 3.2.1. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the third-party recycler, it remains in an evolutionary stable state, and the strategy selection remains unchanged over time. When , , , , and . is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when , this is the third-party recycler's evolutionary stable strategy, choosing low-level recycling.
For cascade utilisation enterprise, the derivative of the replicator dynamic equation with respect to the strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionarily stable state, and the strategy selection remains unchanged over time. When , , , , . At this point, is the evolutionarily stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, when , this is the evolutionarily stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation. The equilibrium points and eigenvalues are shown in Table 8.
4.2.3 Stability analysis of the nonlinear dynamic reward-penalty mechanism
The stability analysis for the manufacturer is consistent with Section 3.2.1. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the third-party recycler, it remains in an evolutionary stable state, and the strategy selection remains unchanged over time. When , , , , and . is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when , this is the third-party recycler's evolutionary stable strategy, choosing low-level recycling.
The derivative of the cascade utilisation enterprise's replicator dynamic equation with respect to strategy probability is: . Let , then , and . When , . At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionarily stable state, and the strategy selection remains stable all the time. When , is the evolutionarily stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, is the evolutionarily stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation. The equilibrium points and eigenvalues of the nonlinear dynamic reward-penalty mechanism are shown in Table 9.
4.3 Model simulation analysis of the optimised reward and penalty system
4.3.1 Model simulation analysis of nonlinear dynamic penalty mechanism
Initial parameters are consistent with those used in the static mechanism. The dynamic penalty coefficients are set as follows: , , , and . Figure 14 shows the simulation results of the nonlinear dynamic penalty mechanism model.
A line graph displays the evolutionary path under a nonlinear dynamic penalty mechanism. The horizontal axis represents time (t) ranging from 0 to 1, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines labeled x equals 0.5, y equals 0.5, and z equals 0.5. The blue line labeled x equals 0.5 shows a decreasing trend, starting at around 0.5 and approaching 0 as time increases. The green line labeled y equals 0.5 shows a slight decrease, starting at around 0.5 and stabilizing just below 0.5. The magenta line labeled z equals 0.5 shows an increasing trend, starting at around 0.5 and stabilizing just above 0.6. Two data points are highlighted at t equals 0.971857, with p values of approximately 0.66482 and 0.546261 for the green and blue lines, respectively.Evolutionary path under nonlinear dynamic penalty mechanism
A line graph displays the evolutionary path under a nonlinear dynamic penalty mechanism. The horizontal axis represents time (t) ranging from 0 to 1, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines labeled x equals 0.5, y equals 0.5, and z equals 0.5. The blue line labeled x equals 0.5 shows a decreasing trend, starting at around 0.5 and approaching 0 as time increases. The green line labeled y equals 0.5 shows a slight decrease, starting at around 0.5 and stabilizing just below 0.5. The magenta line labeled z equals 0.5 shows an increasing trend, starting at around 0.5 and stabilizing just above 0.6. Two data points are highlighted at t equals 0.971857, with p values of approximately 0.66482 and 0.546261 for the green and blue lines, respectively.Evolutionary path under nonlinear dynamic penalty mechanism
According to simulation results, under the non-linear dynamic penalty mechanism, the probability of manufacturers using recycled materials stabilised at 0, indicating that this penalty alone is insufficient to motivate manufacturers to proactively adopt recycled materials. The probability of high-level processing by third-party recyclers stabilises at 0.54, a significant improvement over the linear mechanisms, indicating that the non-linear penalty design effectively curbs low-level processing behaviour and encourages more than half of recyclers to opt for technological upgrades. The probability of proactive cascade utilisation by cascade utilisation enterprises stabilises at 0.66, comparable to the level under linear mechanisms. This mechanism significantly improves the constraint effect on recyclers, but provides insufficient incentives for manufacturers.
4.3.2 Model simulation analysis of nonlinear dynamic reward mechanism
Initial parameters were set consistently with those under the static mechanism, with dynamic reward coefficients defined as , , , and . Figure 15 shows the simulation results of the nonlinear dynamic reward mechanism model.
A line graph with three data lines representing different variables over time. The x-axis is labeled 't' and ranges from 0 to 1. The y-axis is labeled 'p' and ranges from 0 to 1. The blue line with diamond markers represents 'x equals 0.5', the green line with square markers represents 'y equals 0.5', and the magenta line with cross markers represents 'z equals 0.5'. The blue line starts at approximately 0.45 on the y-axis and gradually decreases to near 0 as time progresses. The green line starts at approximately 0.45 on the y-axis and rapidly decreases to near 0 within the first 0.2 units of time. The magenta line starts at approximately 0.45 on the y-axis and quickly increases to near 1 within the first 0.1 units of time, remaining constant thereafter. All values are approximated.Evolutionary path under nonlinear dynamic reward mechanism
A line graph with three data lines representing different variables over time. The x-axis is labeled 't' and ranges from 0 to 1. The y-axis is labeled 'p' and ranges from 0 to 1. The blue line with diamond markers represents 'x equals 0.5', the green line with square markers represents 'y equals 0.5', and the magenta line with cross markers represents 'z equals 0.5'. The blue line starts at approximately 0.45 on the y-axis and gradually decreases to near 0 as time progresses. The green line starts at approximately 0.45 on the y-axis and rapidly decreases to near 0 within the first 0.2 units of time. The magenta line starts at approximately 0.45 on the y-axis and quickly increases to near 1 within the first 0.1 units of time, remaining constant thereafter. All values are approximated.Evolutionary path under nonlinear dynamic reward mechanism
In this mechanism, the probability of manufacturers using recycled materials and the probability of third-party recyclers achieving high-level processing both stabilise at 0, indicating that the non-linear rewards have virtually no incentive effect on manufacturers or recyclers. Conversely, the probability of cascade utilisation enterprises actively engaging in cascade utilisation stabilises at 1, suggesting that the non-linear rewards generate a strong positive incentive for these enterprises, causing them to be fully inclined towards active cascade utilisation. This mechanism is capable of driving the cascade utilisation market to an optimal state on its own, but it fails to improve manufacturers' willingness to recycle or recyclers' processing standards.
4.3.3 Model simulation analysis of nonlinear dynamic reward-penalty mechanism
Initial parameters were set consistently with those under the static mechanism. The dynamic penalty coefficients were defined as , , , and , while dynamic reward coefficients were set as , , , and . Figure 16 shows the simulation results of the nonlinear dynamic reward-penalty mechanism model.
A line graph depicts the evolutionary path under a nonlinear dynamic reward-penalty mechanism. The horizontal axis represents time (t) ranging from 0 to 3, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines, each representing different initial conditions: x equals 0.5, y equals 0.5, and z equals 0.5. The blue line with diamond markers represents x equals 0.5, showing an upward trend starting from approximately 0.5 and stabilizing around 0.86. The green line with square markers represents y equals 0.5, showing a slight upward trend starting from approximately 0.5 and stabilizing around 0.65. The magenta line with cross markers represents z equals 0.5, showing a downward trend starting from approximately 0.5 and stabilizing around 0.33. The graph includes labels for specific data points, such as X 2.87805 and Y 0.864839 for the blue line, X 1.68639 and Y 0.65443 for the green line, and X 2.74513 and Y 0.325632 for the magenta line.Evolutionary path under nonlinear dynamic reward-penalty mechanism
A line graph depicts the evolutionary path under a nonlinear dynamic reward-penalty mechanism. The horizontal axis represents time (t) ranging from 0 to 3, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines, each representing different initial conditions: x equals 0.5, y equals 0.5, and z equals 0.5. The blue line with diamond markers represents x equals 0.5, showing an upward trend starting from approximately 0.5 and stabilizing around 0.86. The green line with square markers represents y equals 0.5, showing a slight upward trend starting from approximately 0.5 and stabilizing around 0.65. The magenta line with cross markers represents z equals 0.5, showing a downward trend starting from approximately 0.5 and stabilizing around 0.33. The graph includes labels for specific data points, such as X 2.87805 and Y 0.864839 for the blue line, X 1.68639 and Y 0.65443 for the green line, and X 2.74513 and Y 0.325632 for the magenta line.Evolutionary path under nonlinear dynamic reward-penalty mechanism
Under the non-linear dynamic reward-penalty mechanism, the probability of manufacturers producing recycled materials stabilises at 0.86, a significant improvement on the previous two single mechanisms. This demonstrates that combining rewards and penalties can effectively incentivise manufacturers to fulfil their recycling responsibilities, as they face both revenue from penalties imposed on non-compliant parties and expenditure on rewards for compliant parties; their net revenue is closely linked to the behaviour of recyclers and cascade utilisation enterprises. The probability of high-level processing by third-party recyclers stabilises at 0.65, the highest level across all mechanisms, indicating that the combination of non-linear rewards and penalties simultaneously discourages low-level processing and incentivises high-level processing, thereby creating a powerful driving force for recyclers. The probability of proactive cascade utilisation by cascade utilisation enterprises stabilises at 0.33, which is relatively low. Nevertheless, this mechanism performs well in balancing the interests of the three parties, enhancing the technical capabilities of recyclers and the participation of manufacturers, thereby achieving a relatively ideal state of systemic equilibrium.
As the central entity, manufacturers should design internal reward and penalty systems to translate government subsidies into dynamic incentives and constraints for upstream and downstream players. For recyclers, a combination of rewards and penalties is more effective than a single approach in driving technological upgrades; for cascade utilisation enterprises, a medium-level equilibrium may arise when the intensity of rewards and penalties is mismatched. Regulators need to fine-tune the reward and penalty coefficients to avoid unduly dampening enthusiasm. In practice, enterprises can dynamically adjust reward and penalty parameters according to the cost structures of different stages: offering higher rewards for costly technological upgrades and imposing progressively stricter penalties on behaviour prone to “free-riding”, thereby achieving overall optimisation at the supply chain level. For policymakers, they should encourage the adoption of dynamic internal incentive and penalty schemes by providing technical guidance and promoting information sharing across the supply chain, and validate the practical effectiveness of different reward and penalty functions through pilot projects, thereby providing empirical evidence to support industry-wide adoption.
5. Conclusion
This research centres on the reward and penalty methods in the power battery recycling. A tripartite evolutionary game model involving manufacturers, third-party recyclers, and cascaded utilisation enterprises is developed to examine the revenues, costs, and strategic stability of each party under different strategy choices. The model's validity was evaluated using numerical simulation analysis. This research investigates how various reward and penalty mechanisms influence the evolutionary trajectory of the system. The main conclusions include:
The system exhibits cyclical fluctuations and struggles to achieve stable cooperation under the static reward-penalty mechanism. The linear dynamic mechanism provides some motivation for manufacturers and cascade utilisation enterprises to recycle materials for remanufacturing and to actively engage in cascade utilisation, but it fails to provide effective incentives or constraints for third-party recyclers. Nonlinear dynamic reward-penalty mechanisms, by dynamically linking reward and penalty intensity to strategy selection, is more effective in guiding third-party recyclers to improve their processing standards, whilst significantly enhancing manufacturers' willingness to recycle, thereby achieving a relatively ideal systemic equilibrium.
Based on these findings, manufacturers should proactively assume primary responsibility for recycling. By thoroughly considering stakeholders' interests and behavioural patterns, they should design internal reward and penalty mechanisms. Concurrently, manufacturers can establish long-term, stable partnerships with third-party recyclers and cascade utilisation enterprises, and jointly invest resources in technological research and development to enhance the quality of recycled materials and the market competitiveness of cascade utilisation products. Furthermore, governments and relevant institutions should provide the technology and platforms to facilitate information sharing within the supply chain.
This paper has limitations: First, the research focuses only on manufacturers, recyclers, and utilisation enterprises, excluding other stakeholders like governments and consumers. Second, the model assumes bounded rationality with static parameters, ignoring dynamic market fluctuations.

