Purpose

With the rapid and explosive development of the new energy vehicle industry, power battery recycling has evolved into a vital component in fulfilling resource circulation and the “dual carbon” goals. The current industry faces challenges, including poor coordination between market entities and unbalanced profit distribution. This research aims to improve recycling efficiency through optimised reward mechanisms.

Design/methodology/approach

This study presents a tripartite evolutionary game model. Manufacturers, third-party recyclers and cascade utilisation enterprises are the key participants. It takes government subsidies and reward-penalty mechanisms within the supply chain into account. It examines the evolutionary rules of strategy selection for each participant, along with the stability of the system. Then it carries out simulation validation.

Findings

Static reward-penalty mechanism results in cyclical fluctuation within the system and impedes stable cooperation. Linear dynamic reward and penalty mechanisms provide some motivation for manufacturers and cascade utilisation enterprises but fail to impose effective incentives or constraints on third-party recyclers. The nonlinear dynamic reward-penalty mechanism, by dynamically linking reward and penalty intensity to strategy selection, more effectively guides third-party recyclers to improve their processing standards while significantly enhancing manufacturers' recycling willingness, thereby achieving relatively ideal system equilibrium.

Originality/value

The research incorporates a nonlinear dynamic reward and penalty mechanism within supply chain incentives. The method modulates the severity of penalties based on participants' strategic decisions, promoting stable collaborative partnerships among them.

With the explosive worldwide growth of the new energy vehicle industry, the amount of retired power batteries is increasing. Their large-scale recycling and efficient utilisation have become critical to resource circulation and achieving the “dual carbon” goals. By 2040, China is expected to generate 1.5–3.3 million tons of discarded EV batteries (Jiang et al., 2021). Their efficient recycling and utilisation are critical for resource circulation and the “dual carbon” goals. Improper handling causes heavy metal contamination and waste. A closed-loop supply chain with cascade utilisation enhances resource efficiency. However, the industry lacks standardisation, suffers from poor coordination and imbalanced profit distribution. The complex interplay of interests across the industrial chain necessitates scientific policy design and supply chain coordination mechanisms to promote healthy industry development. Against this backdrop, examining the strategic interactions and reward mechanisms among multiple stakeholders within the power battery recycling system holds significant importance for optimising resource allocation and balancing economic and environmental benefits.

Facing these complex challenges, the integrated management model of the power battery closed-loop supply chain provides new insights into how to resolve industrial bottlenecks. Some scholars have focused on the decision-making and coordination mechanisms among supply chain members. Ma et al. (2026) found that higher consumer sensitivity to cascaded product performance encourages better design standards and greater recycling participation. Yu and Wang (2025) focused on the synergistic impact of cascaded utilisation and the Extended Producer Responsibility (EPR) system on decision-making, demonstrating that enhancing data transparency and process verifiability improves operational efficiency across the entire chain. Yan et al. (2024) examined cascade utilisation and EPR systems via three pricing models, concluding that lenient EPR regulation is ineffective in governing waste battery disposal during high-revenue market stages. Further exploring the balance between competition and collaboration, Hu et al. (2025) adopted a non-cooperative-cooperative hybrid game to examine pricing and profit distribution in cascade utilisation, attempting to establish a balance between competition and collaboration. Xu et al. (2023a) analysed profits and revenue, showing that applying low-carbon innovation in cascaded-use manufacturing increases profit growth. Furthermore, factors such as fairness concerns and corporate social responsibility (CSR) have been explored. Zhang and Liang (2023) proposed a contractual mechanism that coordinates CSR investment while mitigating fairness concerns, which were shown to hinder supply chain efficiency. Tian et al. (2022) explored recycling coordination and selection strategies while taking demand uncertainties and CSR into consideration, analysing optimal strategies under different recycling modes.

Government subsidies play an important part in power battery recycling. Related research examines how these policies influence supply chain entities' behaviour and overall operations from a variety of perspectives. Wu et al. (2025) investigated recycling models that maximise overall corporate or supply chain benefits under government subsidies and extended producer responsibility schemes. They compared manufacturer-led, retailer-led, third-party, and joint recycling approaches. Wen et al. (2025) investigated the feasibility of online recycling networks and government funding options, and explored how governments might design subsidy policies to incentivise online recycling and promote cascading utilisation. Zhang et al. (2023a) compared manufacturer subsidies, recycler subsidies, and no subsidies, analysing how the recipient type affects green technology investment, recycling outcomes, and environmental benefits. Examining behavioural preferences, Liu and Zhu (2024) found that battery capacity determines optimal subsidy policy choices and that preferences may hinder recycling efficiency. Taking a long-term view, Yu and Hou (2023) constructed a differential game model and found that the impact of cost subsidy policies on supply chain profits strengthens over time.

As policy research deepens, scholars increasingly focus on the differentiated effects of various policy tools, comparing reward-penalty mechanisms with traditional subsidies. Yang et al. (2022) contrasted no intervention, subsidies, and reward-penalty mechanisms, finding reward-penalty mechanisms more effective than subsidies at incentivising power battery recycling and increasing recovery rates. Tang et al. (2019) examined how reward-penalty and subsidy schemes affected the promotion of power battery recycling. Their findings show these mechanisms have more noticeable effects on raising recycling rates and social welfare. Zhang et al. (2022) constructed three models under government reward-penalty frameworks, demonstrating that automakers implementing reward-penalty policies can maximise both recycling rates and supply chain profits. Wei and Qi (2025) compared reward-penalty and subsidy mechanisms, finding that reward-penalty mechanisms are more effective in motivating recycling behaviour and improving recycling rates. Zhang et al. (2023b) built and compared several closed-loop supply chain models. Their research explores how government subsidies, deposit-refund systems, and reward-penalty policies affect recycling rates.

In summary, the power battery recycling in China has progressed, but the regulatory structure remains incomplete. Existing research mainly focuses on analysing the impact of governmental policies on the power battery recycling, and the reward-plenty mechanism designs are centred on static subsidies. Designing reward mechanisms within the supply chain under government subsidy frameworks becomes a key area for future research.

With this context, the main contents of this study are structured as follows: First, a tripartite evolutionary game model is created. It includes manufacturers, third-party recyclers, and cascade utilisation enterprises. Government subsidies for manufacturers are included in the model. By designing internal supply chain incentive mechanisms, the dynamic evolutionary patterns of each party's strategy selection are analysed. It investigates the equilibrium point stability conditions using Lyapunov stability theory. Simulate the evolutionary path of each game participant. It reveals cyclical fluctuations under the static reward-penalty mechanism. Second, optimisation techniques for dynamic reward and penalty mechanisms are devised. They address the limitations of static mechanisms. These methods investigate how different reward-penalty combinations influence system stability. Finally, the nonlinear reward-penalty mechanism is used to improve the dynamic reward-penalty control method. This approach supports the sustainable development of resource circulation systems under the dual carbon goals.

The recycling process for power batteries encompasses two key stages: cascade utilisation and dismantling recycling. This establishes a closed-loop supply chain system centred on three core nodes: manufacturers, third-party recyclers, and cascade utilisation enterprises. All these entities are dedicated to the recovery and reuse of power batteries. Under this framework, manufacturers employ differentiated pricing strategies to determine whether recycled battery materials should be utilised for remanufacturing, subsequently deploying newly produced power batteries.

According to the Industry Specification for Comprehensive Utilisation of Waste Power Batteries (Ministry of Industry and Information Technology et al., 2024) (hereinafter referred to as the Specification), recyclers achieving technical benchmarks such as lithium recovery rates of no less than 90% during smelting, electrode powder recovery rates of no less than 98% after crushing and separation, and impurity aluminium content below 1.5% shall be deemed to have implemented high-level processing. Third-party recyclers gather retired power batteries from the electric vehicle market. They consider whether to implement technological innovations to enhance the quality standards of processed remanufactured batteries. Third-party recyclers then sell remanufactured batteries of varying processing standards to cascade utilisation enterprises. They are responsible for collecting all waste batteries that have completed their cascade use cycle and delivering them to manufacturers for subsequent processing.

According to the Specification, cascade utilisation operators shall be deemed to engage in proactive cascade utilisation if the annual volume of cascade-utilised waste power batteries reaches no less than 60% (by weight) (Ministry of Industry and Information Technology et al., 2024) of the actual volume of waste power batteries recovered. Manufacturers, after receiving the waste batteries, hand them over to specialised dismantling facilities for processing.

To encourage coordination and long-term growth of this closed-loop supply chain, the government pays subsidies to manufacturers. Manufacturers design and implement reward and penalty mechanisms within the supply chain. These techniques are intended to stimulate active participation from third-party recyclers and cascade utilisation enterprises. Figure 1 illustrates the logical relationships among all three participants in the evolutionary game.

Figure 1
A flowchart illustrating the logic relationships of a power battery closed-loop supply chain.A flowchart illustrating the logic relationships of a power battery closed-loop supply chain. The government provides subsidies to manufacturers and implements the EPR system. The manufacturer is involved in the green disposal of discarded batteries and the payment of environmental treatment fees. The manufacturer also engages in recycled material remanufacturing and new material manufacturing, supplying the electric vehicle market. The cascade utilization enterprise processes recycling waste batteries and engages in active or passive cascade utilization. The cascade utilization enterprise sends recycled waste batteries back to the manufacturer and processes recycling batteries after cascade utilization to the third-party recycler. The third-party recycler recycles power batteries and sends them back to the manufacturer. The manufacturer reduces or does not affect disassembly costs and receives rewards or penalties.

Logic relationships of power battery closed-loop supply chain

Figure 1
A flowchart illustrating the logic relationships of a power battery closed-loop supply chain.A flowchart illustrating the logic relationships of a power battery closed-loop supply chain. The government provides subsidies to manufacturers and implements the EPR system. The manufacturer is involved in the green disposal of discarded batteries and the payment of environmental treatment fees. The manufacturer also engages in recycled material remanufacturing and new material manufacturing, supplying the electric vehicle market. The cascade utilization enterprise processes recycling waste batteries and engages in active or passive cascade utilization. The cascade utilization enterprise sends recycled waste batteries back to the manufacturer and processes recycling batteries after cascade utilization to the third-party recycler. The third-party recycler recycles power batteries and sends them back to the manufacturer. The manufacturer reduces or does not affect disassembly costs and receives rewards or penalties.

Logic relationships of power battery closed-loop supply chain

Close modal

To build the game model, each participant's options for strategy and the stability of the system must be considered. This will aid in determining the optimal set of circumstances when the three participants can use socially stable methods. The following assumptions are made based on practical contexts and with reference to the relevant literature (Liu and Ma, 2021; Guan et al., 2023):

Assumption 1.

Three participants are involved: manufacturers, third-party recyclers, and cascade utilisation enterprises. All three are bounded rational individuals. Their strategy choices shift and become stable as time passes. They adjust their own strategies by observing each other's strategy choices and payoff outcomes, focusing on the dynamic evolutionary process of their group behaviour.

Assumption 2.

Manufacturers, third-party recyclers, and cascade utilisation enterprises each have two strategy options. The manufacturer's strategy set for power battery production materials is {recycled material remanufacturing, new material manufacturing}. The probability of selecting remanufacturing of recycled materials is x and the probability of selecting manufacturing of new materials is 1x; Third-party recyclers have the following strategy set for retired battery recycling: {high-level processing, low-level processing}. The probability of selecting high-level processing is y, while the probability of selecting low-level processing is 1y. Cascade utilisation enterprises have the following strategy set for actively developing cascade utilisation products: {active cascade utilisation, passive cascade utilisation}. The probability of selecting active cascade utilisation is z, while the probability of selecting passive cascade utilisation is 1z, x,y,z[0,1].

Assumption 3.

The manufacturer's sales revenue from remanufacturing recycled materials is M11, while the sales revenue from manufacturing new materials is M12(M12>M11). If the manufacturer chooses to remanufacture recycled materials, it needs to gather all wasted power batteries from the cascade utilisation enterprise and third-party recycler, and chemically disassemble the recycled batteries for a cost C11. If the manufacturer decides to manufacture additional materials, the procurement cost is C14(C11>C14).

Assumption 4.

When manufacturers opt for new material production, the government mandates battery recycling under the EPR policy, requiring handover to specialised dismantling enterprises (Xu et al., 2023b). Manufacturers bear the associated recycling and processing costs, calculated as C13(C11>C13). Should third-party recyclers perform high-level processing on recovered batteries, this partially offsets the green processing fees paid by manufacturers, with an offset ratio of ϕ(ϕ(0,1)).

Assumption 5.

Third-party recyclers experience profit variations influenced by cascade utilisation strategies when processing waste power batteries. If third-party recyclers implement high-level processing and cascade utilisation enterprises take an active role in cascade utilisation, recyclers earn profits R21. If cascade utilisation enterprises passively utilise batteries, recyclers lose additional profits due to reduced sales volume T. High-level processing by recyclers causes additional costs C21, including technological improvements and equipment maintenance. When third-party recyclers perform low-level processing, if cascade utilisation enterprises actively engage in cascade utilisation, their profit is R22(R22>R21). If cascade utilisation enterprises passively engage in cascade utilisation, third-party recyclers lose profit D(T>D). If manufacturers remanufacture recycled materials, high-level processing by recyclers partially reduces dismantling costs by factor θ(θ(0,1)).

Assumption 6.

The cascade utilisation market remains in its early development stage, with relatively low demand for spent battery reuse. The volume of spent batteries recovered from the electric vehicle market can satisfy cascade utilisation enterprises for active cascade utilisation. If cascade utilisation operators engage in active cascade utilisation, the profit from selling high-level processed and remanufactured batteries is π31, while selling low-level processed batteries results in a profit loss of e. If they engage in passive cascade utilisation, the profit from selling high-level processed and remanufactured batteries is π32(π32>π31), while selling low-level processed batteries results in a profit loss of L(e>L). Given poor consistency and lifespan estimation difficulties, active cascade utilisation requires additional R&D costs C31 (Yuan and Zhang, 2025).

Assumption 7.

If manufacturers choose to use recycled materials for remanufacturing, the government provides manufacturers with reward subsidies S. Manufacturers use these rewards to guide enterprises toward technological upgrades and shared environmental responsibility, avoiding low-cost, low-quality competition and promoting the greening of power batteries across the whole life cycle. Consequently, manufacturers incentivise recyclers to achieve high-level processing and share subsidies Q, and active cascade utilisation enterprises to actively engage in cascade utilisation and share subsidies P. According to relevant regulations (Ministry of Industry and Information Technology et al., 2026), manufacturing enterprises must establish recycling service outlets within their sales regions and assume corresponding responsibility for the flow of power batteries. In line with the EPR requirements, manufacturers should translate external regulatory pressures into internal supply chain management requirements, effectively constraining non-compliant partners. Therefore, if recyclers fail to meet processing standards, manufacturers impose penalties n; if cascade utilisation enterprises demonstrate passive utilisation, they face sanctions f.

A game matrix is generated based on the model assumptions stated above. It includes manufacturers, third-party recyclers, and cascade utilisation enterprises. The matrix is detailed in Table 1.

Table 1

Tripartite game matrix for cascaded utilisation of retired power batteries

Third-party recyclerCascade utilisation enterprise
Active cascade
Utilisation (z)
Passive cascade utilisation (1z)
Battery ManufacturerRecycled Material Remanufacturing (x)High-Level Processing (y)M11θC11+SPQ, R21C21+Q, π31+PC31M11θC11+SQ+f, R21C21T+Q, π32f
Low-Level Processing (1y)M11C11+SP+n,
R22n, π31e+PC31
M11C11+S+n+f, R22Dn, π32Lf
New Material Manufacturing (1x)High-level Processing (y)M12ϕC13C14,
R21C21, π31C31
M12ϕC13C14, R21C21T, π32
Low-level processing (1y)M12C13C14,
R22, π31eC31
M12C13C14,
R22D, π32L

Construct replication dynamic equation functions based on the game matrix involving manufacturers, third-party recyclers, and cascade utilisation enterprises.

2.2.1 Manufacturer dynamic equation function construction

The expected benefits for manufacturers using recycled materials for remanufacturing E11, the expected returns for using new materials in manufacturing E12 and the average expected returns E1̅ are, respectively:

(2)
(3)
(4)

The dynamic replication equation for the manufacturer's chosen strategy is produced by the system of equations:

(5)

2.2.2 Third-party recyclers' dynamic equation function construction

The expected benefits for third-party recyclers using high-level processing E21, the expected returns for low-level processing by third-party recyclers E22 and the average expected returns E2̅ are, respectively:

(6)
(7)
(8)

For third-party recyclers, the dynamics replication equation is:

(9)

2.2.3 Cascade utilisation enterprise dynamic equation function construction

The expected benefits for active and passive cascade utilisation by cascade utilisation enterprises are E31 and E32, and the average expected benefits E3̅ are, respectively:

(10)
(11)
(12)

The dynamic replication equation for cascade utilisation enterprise strategy selection is given below:

(13)

From F(x)=0, F(y)=0, and F(z)=0, the equilibrium points of the system are obtained as: E1(0,0,0), E2(1,0,0), E3(0,1,0), E4(0,0,1), E5(1,1,0), E6(1,0,1), E7(0,1,1), and E8(1,1,1). The Jacobian matrix is expressed by the equation J below:

According to Lyapunov's indirect method, the stability of an equilibrium point is determined by the sign of the eigenvalues of its Jacobian matrix: if all eigenvalues are negative, the equilibrium is asymptotically stable; conversely, if at least one eigenvalue is positive, it is unstable. For the purpose of facilitating Stability analysis, set =M12M11+C11C13C14S , B=M12M11C14SϕC13C11+θC11 , H=R22R21+C21 , K=π32π31+C31. The eigenvalues of the Jacobian matrix corresponding to each equilibrium point are listed in Table 2.

Table 2

Equilibrium point stability evaluation

Equilibrium pointEigenvalueSignStability condition
E1(0,0,0)λ1=A+n+fλ2=H+DT<0λ3=Ke+L<0(,,)Uncertain point; Stable for A>n+f
E2(1,0,0)λ1=Anfλ2=H+DT+Q+nλ3=Ke+L+P+f(,,)Uncertain point; Stable for A<n+f,H>DT+Q+n,K>Le+P+f
E3(0,1,0)λ1=B+fQλ2=HD+T>0λ3=K<0(,+,)Saddle point
E4(0,0,1)λ1=A+nPλ2=H<0λ3=K+eL>0(,,+)Saddle point
E5(1,1,0)λ1=B+Qfλ2=HD+TQnλ3=K+P+f(,,)Uncertain point; Stable for B<fQ,H<DT+Q+n,K>P+f
E6(1,0,1)λ1=A+Pnλ2=H+Q+nλ3=K+eLPf(,,)Uncertain point; Stable for A<nP,H>Q+n,K<Le+P+f
E7(0,1,1)λ1=BPQλ2=H>0λ3=K>0(,+,+)Saddle point
E8(1,1,1)λ1=B+Q+Pλ2=HQnλ3=KPf(,,)Uncertain point; Stable for B<QP,H<Q+n,K<P+f

Note(s): + for positive sign; − for negative sign; unknown signs marked

To verify the validity of the evolutionary stability analysis, the model is parameterised with numerical values based on realistic conditions and relevant literature (Wu and Zhang, 2025; Zhang et al., 2024a, b). Table 3 Initial Parameter Settings presents the initial parameter settings.

Table 3

Initial parameter settings

ParameterM11M12C11C13C14R21R22TDC21S
Initial Value17322181245944641172030
Parametersπ31π32eLC31θϕnfQP
Initial Value14203250.80.92010510

Using MATLAB R2024a software, data simulations were conducted on the evolutionary track of each game entity. Figure 2 displays the tripartite evolutionary game's initial results.

Figure 2
A three-dimensional line graph with multiple colored lines forming a spiral pattern.A three-dimensional line graph with multiple colored lines forming a spiral pattern. The graph features three axes labeled x, y, and z, each ranging from 0 to 1. The lines converge towards the center of the graph, creating a dense, colorful spiral. The graph appears to illustrate a complex, dynamic system with multiple trajectories.

Initial tripartite evolutionary game path

Figure 2
A three-dimensional line graph with multiple colored lines forming a spiral pattern.A three-dimensional line graph with multiple colored lines forming a spiral pattern. The graph features three axes labeled x, y, and z, each ranging from 0 to 1. The lines converge towards the center of the graph, creating a dense, colorful spiral. The graph appears to illustrate a complex, dynamic system with multiple trajectories.

Initial tripartite evolutionary game path

Close modal

The system's evolutionary game path forms a closed curve with periodic motion around a stable centre point, lacking any stable equilibrium points. When initial values for x, y, z are set to (0.5, 0.5, 0.5) and (0.2, 0.5, 0.8) respectively, the system's evolutionary paths are shown in Figures 3 and 4. The strategic choices of third-party recyclers have gradually trended towards low-level treatment over time, and the strategic choices of manufacturers and cascade utilisation enterprises exhibit a pattern of continuous fluctuation. This demonstrates that under the current model settings, the interplay of interests among various entities is rather complex. Further optimisation of the reward and penalty mechanisms or adjustment of other relevant parameters is required to improve the system's evolutionary trajectory.

Figure 3
A line graph showing evolutionary paths with different initial values over time.A line graph displays the evolutionary path with initial values of 0.5, 0.5, and 0.5. The horizontal axis represents time (t) ranging from 0 to 10, and the vertical axis represents the variable p ranging from 0 to 1. There are three distinct lines: a blue line with star markers representing x equals 0.2, a green line with square markers representing y equals 0.5, and a pink line with diamond markers representing z equals 0.8. The blue and pink lines show oscillatory behavior with peaks and troughs, while the green line starts high and quickly drops to near zero, remaining flat for the rest of the time period. All values are approximated.

Evolutionary path with initial values (0.5, 0.5, 0.5)

Figure 3
A line graph showing evolutionary paths with different initial values over time.A line graph displays the evolutionary path with initial values of 0.5, 0.5, and 0.5. The horizontal axis represents time (t) ranging from 0 to 10, and the vertical axis represents the variable p ranging from 0 to 1. There are three distinct lines: a blue line with star markers representing x equals 0.2, a green line with square markers representing y equals 0.5, and a pink line with diamond markers representing z equals 0.8. The blue and pink lines show oscillatory behavior with peaks and troughs, while the green line starts high and quickly drops to near zero, remaining flat for the rest of the time period. All values are approximated.

Evolutionary path with initial values (0.5, 0.5, 0.5)

Close modal
Figure 4
A line graph showing the evolutionary path with initial values of 0.2, 0.5, and 0.8.A line graph displays the evolutionary path with initial values of 0.2, 0.5, and 0.8. The horizontal axis represents time (t) ranging from 0 to 10, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines: a blue line with star markers for x equals 0.2, a green line with square markers for y equals 0.5, and a magenta line with circle markers for z equals 0.8. The blue line shows a periodic pattern with peaks around 0.9 and troughs around 0.1. The magenta line also shows a periodic pattern but with slightly lower peaks around 0.85 and higher troughs around 0.3. The green line remains close to zero throughout the time range.

Evolutionary path with initial values (0.2, 0.5, 0.8)

Figure 4
A line graph showing the evolutionary path with initial values of 0.2, 0.5, and 0.8.A line graph displays the evolutionary path with initial values of 0.2, 0.5, and 0.8. The horizontal axis represents time (t) ranging from 0 to 10, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines: a blue line with star markers for x equals 0.2, a green line with square markers for y equals 0.5, and a magenta line with circle markers for z equals 0.8. The blue line shows a periodic pattern with peaks around 0.9 and troughs around 0.1. The magenta line also shows a periodic pattern but with slightly lower peaks around 0.85 and higher troughs around 0.3. The green line remains close to zero throughout the time range.

Evolutionary path with initial values (0.2, 0.5, 0.8)

Close modal

We conducted a sensitivity analysis on key parameters. By setting the penalty parameters for manufacturers towards third-party recyclers n to n=10,20,40 respectively, and those for cascade utilisation f to f=5,10,20 respectively, the effects of these parameter changes on the system's evolution are illustrated in Figures 5 and 6. The impact of changes in manufacturers' penalty on third-party recyclers. When the penalty level is low, the volatility of the system increases. When the penalty level imposed on third-party recyclers is high, cascade utilisation enterprises reach a stable state, but manufacturers and third-party recyclers remain in a volatile state. When the penalty level imposed on cascade utilisation enterprises is high, third-party recyclers remain in a stable state characterised by a tendency towards low-level recycling, whilst the volatility of manufacturers and cascade utilisation enterprises intensifies.

Figure 5
Three line graphs showing the impact of changes in manufacturers' penalty on third-party recyclers.The image contains three line graphs side by side, each representing different values of n (10, 20, and 40) and their respective x, y, and z coordinates over time t. The x-axis represents time t ranging from 0 to 10, and the y-axis represents the probability p ranging from 0 to 1. Each graph shows three lines: blue for x, green for y, and magenta for z. The lines exhibit periodic behavior with varying amplitudes and frequencies. For n equals 10, the blue line (x) shows the highest amplitude, followed by the magenta line (z), while the green line (y) remains close to zero. For n equals 20, the blue and magenta lines (x and z) show similar amplitudes with more frequent oscillations, while the green line (y) remains close to zero. For n equals 40, the blue line (x) shows the highest amplitude, followed by the magenta line (z), while the green line (y) shows more pronounced oscillations compared to the previous graphs. All values are approximated.

The impact of changes in manufacturers' penalty on third-party recyclers

Figure 5
Three line graphs showing the impact of changes in manufacturers' penalty on third-party recyclers.The image contains three line graphs side by side, each representing different values of n (10, 20, and 40) and their respective x, y, and z coordinates over time t. The x-axis represents time t ranging from 0 to 10, and the y-axis represents the probability p ranging from 0 to 1. Each graph shows three lines: blue for x, green for y, and magenta for z. The lines exhibit periodic behavior with varying amplitudes and frequencies. For n equals 10, the blue line (x) shows the highest amplitude, followed by the magenta line (z), while the green line (y) remains close to zero. For n equals 20, the blue and magenta lines (x and z) show similar amplitudes with more frequent oscillations, while the green line (y) remains close to zero. For n equals 40, the blue line (x) shows the highest amplitude, followed by the magenta line (z), while the green line (y) shows more pronounced oscillations compared to the previous graphs. All values are approximated.

The impact of changes in manufacturers' penalty on third-party recyclers

Close modal
Figure 6
Three line graphs depict the impact of changes in manufacturers' penalty on cascade utilisation enterprises.Three line graphs depict the impact of changes in manufacturers' penalty on cascade utilisation enterprises. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different frequencies (f) and variables (x, y, z). Panel A shows the impact for f=5, Panel B for f=10, and Panel C for f=20. Each panel contains three lines representing different variables: x in blue, y in green, and z in magenta. The lines show oscillatory behavior with varying amplitudes and frequencies. In Panel A, the blue line (f=5, x) has the highest amplitude, followed by the magenta line (f=5, z), while the green line (f=5, y) remains constant at zero. In Panel B, the blue line (f=10, x) and the magenta line (f=10, z) show similar amplitudes, while the green line (f=10, y) remains constant at zero. The graphs illustrate how the penalty changes affect the utilization enterprises over time.

The impact of changes in manufacturers' penalty on cascade utilisation enterprises

Figure 6
Three line graphs depict the impact of changes in manufacturers' penalty on cascade utilisation enterprises.Three line graphs depict the impact of changes in manufacturers' penalty on cascade utilisation enterprises. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different frequencies (f) and variables (x, y, z). Panel A shows the impact for f=5, Panel B for f=10, and Panel C for f=20. Each panel contains three lines representing different variables: x in blue, y in green, and z in magenta. The lines show oscillatory behavior with varying amplitudes and frequencies. In Panel A, the blue line (f=5, x) has the highest amplitude, followed by the magenta line (f=5, z), while the green line (f=5, y) remains constant at zero. In Panel B, the blue line (f=10, x) and the magenta line (f=10, z) show similar amplitudes, while the green line (f=10, y) remains constant at zero. The graphs illustrate how the penalty changes affect the utilization enterprises over time.

The impact of changes in manufacturers' penalty on cascade utilisation enterprises

Close modal

Varying Q=2,5,10 and P=5,10,20 gives the impact shown in Figures 7 and 8. When the reward parameters for third-party recyclers are altered, there is no significant change in the system's volatility. When the incentive level for cascade utilisation enterprises is low, the system's volatility increases. When the incentive level is high, the volatility of both manufacturers and cascade utilisation enterprises is significantly reduced.

Figure 7
Three line graphs showing the impact of changes in manufacturers' reward on third-party recyclers.The image contains three line graphs side by side, each depicting the impact of changes in manufacturers' reward on third-party recyclers. The x-axis represents time (t) ranging from 0 to 10, and the y-axis represents the probability (p) ranging from 0 to 1. Each graph shows three different scenarios labeled as Q equals 2x, Q equals 2y, and Q equals 2z for the first graph; Q equals 5x, Q equals 5y, and Q equals 5z for the second graph; and Q equals 10x, Q equals 10y, and Q equals 10z for the third graph. The lines are color-coded: blue for x, green for y, and magenta for z. The blue and magenta lines exhibit oscillatory behavior with varying amplitudes, while the green lines start at a high probability and quickly drop to zero, remaining constant thereafter. The blue and magenta lines show periodic peaks and troughs, indicating fluctuating probabilities over time. All values are approximated.

The impact of changes in manufacturers' reward on third-party recyclers

Figure 7
Three line graphs showing the impact of changes in manufacturers' reward on third-party recyclers.The image contains three line graphs side by side, each depicting the impact of changes in manufacturers' reward on third-party recyclers. The x-axis represents time (t) ranging from 0 to 10, and the y-axis represents the probability (p) ranging from 0 to 1. Each graph shows three different scenarios labeled as Q equals 2x, Q equals 2y, and Q equals 2z for the first graph; Q equals 5x, Q equals 5y, and Q equals 5z for the second graph; and Q equals 10x, Q equals 10y, and Q equals 10z for the third graph. The lines are color-coded: blue for x, green for y, and magenta for z. The blue and magenta lines exhibit oscillatory behavior with varying amplitudes, while the green lines start at a high probability and quickly drop to zero, remaining constant thereafter. The blue and magenta lines show periodic peaks and troughs, indicating fluctuating probabilities over time. All values are approximated.

The impact of changes in manufacturers' reward on third-party recyclers

Close modal
Figure 8
Three line graphs depict the impact of changes in manufacturers' reward on cascade utilisation enterprises.Three line graphs depict the impact of changes in manufacturers' reward on cascade utilisation enterprises. Each graph shows the probability (p) on the vertical axis and time (t) on the horizontal axis. The graphs are labeled with different parameters: P equals 5, 10, and 20, with corresponding x, y, and z values. Panel A: The first line graph shows the probability (p) over time (t) for P equals 5. The blue line represents P equals 5, x, the green line represents P equals 5, y, and the magenta line represents P equals 5, z. The blue and magenta lines show a rapid increase and then a decrease, while the green line remains relatively constant. Panel B: The second line graph shows the probability (p) over time (t) for P equals 10. The blue line represents P equals 10, x, the green line represents P equals 10, y, and the magenta line represents P equals 10, z. The blue and magenta lines exhibit oscillatory behavior, while the green line remains relatively constant. Panel C:

The impact of changes in manufacturers' reward on cascade utilisation enterprises

Figure 8
Three line graphs depict the impact of changes in manufacturers' reward on cascade utilisation enterprises.Three line graphs depict the impact of changes in manufacturers' reward on cascade utilisation enterprises. Each graph shows the probability (p) on the vertical axis and time (t) on the horizontal axis. The graphs are labeled with different parameters: P equals 5, 10, and 20, with corresponding x, y, and z values. Panel A: The first line graph shows the probability (p) over time (t) for P equals 5. The blue line represents P equals 5, x, the green line represents P equals 5, y, and the magenta line represents P equals 5, z. The blue and magenta lines show a rapid increase and then a decrease, while the green line remains relatively constant. Panel B: The second line graph shows the probability (p) over time (t) for P equals 10. The blue line represents P equals 10, x, the green line represents P equals 10, y, and the magenta line represents P equals 10, z. The blue and magenta lines exhibit oscillatory behavior, while the green line remains relatively constant. Panel C:

The impact of changes in manufacturers' reward on cascade utilisation enterprises

Close modal

Varying C21=10,20,30 and C31=2,5,10 gives the impact shown in Figures 9 and 10. When the additional investment costs for high-level recycling by third-party recyclers vary, there is no significant change in the system's volatility. When the additional investment costs associated with high-level recycling by third-party recyclers are high, the system's volatility increases. When these costs are low, the volatility of both manufacturers and cascade utilisation enterprises decreases.

Figure 9
Three line graphs showing the impact of cost changes among third-party recyclers.Three line graphs depict the impact of cost changes among third-party recyclers. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different cost change scenarios: c21=10, c21=20, and c21=30. Each scenario is represented by three lines: x, y, and z. The lines are color‑coded: blue for x, green for y, and magenta for z. In the three sets of subgraphs, the blue line (x) and the magenta line (z) exhibit sinusoidal oscillation patterns with peaks and troughs, while the green line (y) remains close to zero. The amplitude and phase show slight differences among the three groups of graphs.

The impact of cost changes among third-party recyclers

Figure 9
Three line graphs showing the impact of cost changes among third-party recyclers.Three line graphs depict the impact of cost changes among third-party recyclers. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different cost change scenarios: c21=10, c21=20, and c21=30. Each scenario is represented by three lines: x, y, and z. The lines are color‑coded: blue for x, green for y, and magenta for z. In the three sets of subgraphs, the blue line (x) and the magenta line (z) exhibit sinusoidal oscillation patterns with peaks and troughs, while the green line (y) remains close to zero. The amplitude and phase show slight differences among the three groups of graphs.

The impact of cost changes among third-party recyclers

Close modal
Figure 10
Three line graphs depict the impact of cost changes among cascade utilisation enterprises.Three line graphs depict the impact of cost changes among cascade utilisation enterprises. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different cost change scenarios: c31=2, c31=5, and c31=10. Each scenario is represented by three lines: x, y, and z. The lines are color-coded: blue for x, green for y, and magenta for z. Panel A shows the impact for c31=2. The blue line (x) and the magenta line (z) exhibit sinusoidal patterns with peaks and troughs, while the green line (y) remains close to zero. Panel B shows the impact for c31=5. The blue line (x) and the magenta line (z) again exhibit sinusoidal patterns, but with different amplitudes and phases compared to Panel A. The green line (y) remains close to zero. Panel C shows the impact for c31=10. The graphs illustrate how the variable p changes over time under different cost change scenarios.

The impact of cost changes among cascade utilisation enterprises

Figure 10
Three line graphs depict the impact of cost changes among cascade utilisation enterprises.Three line graphs depict the impact of cost changes among cascade utilisation enterprises. Each graph shows the relationship between time (t) on the horizontal axis and a variable (p) on the vertical axis. The graphs are labeled with different cost change scenarios: c31=2, c31=5, and c31=10. Each scenario is represented by three lines: x, y, and z. The lines are color-coded: blue for x, green for y, and magenta for z. Panel A shows the impact for c31=2. The blue line (x) and the magenta line (z) exhibit sinusoidal patterns with peaks and troughs, while the green line (y) remains close to zero. Panel B shows the impact for c31=5. The blue line (x) and the magenta line (z) again exhibit sinusoidal patterns, but with different amplitudes and phases compared to Panel A. The green line (y) remains close to zero. Panel C shows the impact for c31=10. The graphs illustrate how the variable p changes over time under different cost change scenarios.

The impact of cost changes among cascade utilisation enterprises

Close modal

The static system's evolutionary strategy exhibits fluctuating patterns. In order to develop an evolutionarily stable strategy, the fundamental model is enhanced. Dynamic reward and penalty mechanism is implemented in internal incentives. Manufacturers may regulate the optimum penalties and rewards for cascade utilisation enterprises and third-party recyclers according to their strategic decisions due to the dynamic system. The reward and penalty amounts are optimised from fixed constants n, f, Q, P to dynamic linear functions. Based on different combinations of reward and penalty mechanisms, the three mechanisms are as follows: dynamic linear penalty, dynamic linear reward, and dynamic linear reward-penalty.

3.1.1 Construction of the linear dynamic penalty mechanism

Assume manufacturers' penalty strategies toward third-party recyclers and cascade utilisation enterprises dynamically correlate with their respective strategy choices. Specifically, the manufacturer's penalty towards third-party recyclers is n1=αn(1y), and towards cascade utilisation enterprises is f1=βf(1z). α and β are the linear dynamic penalty coefficients, α>0, β>0, and 1y reflect the chance of third-party recyclers using low-level recycling methods, and 1z represents the probability of cascade utilisation enterprises using passive cascade utilisation strategies. The reward for cascade utilisation enterprises and third-party recyclers remains a fixed constant P and Q.

In the expected payoffs of the game's participants, the dynamic penalty term appears as a lump sum. Third-party recyclers and cascade utilisation enterprises have (1y)n1 and (1z)f1 deducted respectively for their non-compliant behaviour. Meanwhile, manufacturers see a corresponding increase in their payoffs as a result of collecting the fines. The pure strategy payoff expression E11, E12, E21, E22, E31 and E32, under the static mechanism remains unchanged. the dynamic reward and penalty terms are simply superimposed on the overall expected payoffs of each participant. The expected returns for each participant are obtained:

(13)
(14)
(15)

The dynamic equation is derived based on the expected return described above, as shown below:

(14)
(15)
(16)

3.1.2 Construction of the linear dynamic reward mechanism

Assume the manufacturer's reward strategies for the three participants are dynamically linked to their respective strategy choices. Specifically, the manufacturer's reward for third-party recyclers is Q1=μQy, and for cascade utilisation enterprises is P1=νPz, where μ and ν are linear dynamic reward coefficients, μ>0 and ν>0. Penalties for both third-party recyclers and cascade utilisation enterprises remain fixed constants n and f. In terms of expected returns, third-party recyclers and cascade utilisation enterprises receive additional returns of yQ1 and zP1 respectively as a result of their proactive behaviour. The expected return expressions are:

(17)
(18)
(19)

The replicator dynamics equations for the three participants are, respectively:

(20)
(21)
(22)

3.1.3 Construction of the linear dynamic reward-penalty mechanism

The penalty and reward mechanisms are combined. The expected return expressions are as follows:

(23)
(24)
(25)

The replicator dynamic equations are shown below:

(26)
(27)
(28)

3.2.1 Stability analysis of the linear dynamic penalty mechanism

According to differential equation stability theory, a strategy combination achieves evolutionary equilibrium when the replication dynamics equation equals zero and its derivative is less than zero. For manufacturer, the derivative of the replication dynamics equation with respect to strategy probability is: F(x)=(2x1)[Anf+y((1ϕ)C13(1θ)C11+Q+n)+z(P+f)]. Let F(x)=0, then x=0, x=1 and z=z0=A+f+n+y[(1ϕ)C13(1θ)C11+Q+n)]P+f. When z=z0, F(x)0 . At this point, regardless of the strategy chosen by the manufacturer, the system is in an evolutionarily stable state, and the strategy selection remains unchanged over time. When z<z0, x=1 is the evolutionarily stable strategy for the manufacturer, choosing to recycle materials for remanufacturing. When z>z0, x=0 is the evolutionarily stable strategy for the manufacturer choosing to manufacture with new materials.

For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: F(y)=(2y1)[HD+T+z(DT)x(Q+n)]+αn(y1)(3y2). Let F(y)=0, then y=0, y=1 and x=x0=[HD+T+z(DT)αn(1y)]/(Q+n). When x=x0, F(y)0. At this stage, regardless of the approach chosen by the third-party recycler, it is in an evolutionary stable state, with strategy selection remaining constant throughout time. When x<x0, y=1 is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when y=0, this is the third-party recycler's evolutionary stable method, who prefers low-level recycling.

The derivative of the cascade utilisation enterprise's replicator dynamic equation with respect to strategy probability is: F(z)=(2z1)[KL+e+y(Le)x(P+f)]+βf(z1)(3z2). Let F(z)=0, then z=0, z=1 and x=x0=[KL+e+y(Le)βf(1z)]/(P+f). When x=x0, F(z)0. At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system is evolutionarily stable, and strategy selection remains unchanged over time. When x<x0, z=1 is the evolutionarily stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, when z=0, this is the evolutionarily stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation.

Following Lyapunov's indirect method, the stability analysis is conducted on the equilibrium points of the system for the linear dynamic penalty mechanism. From F(x)=0, F(y)=0, and F(z)=0, the equilibrium points of the system are obtained as: E1(0,0,0), E2(1,0,0), E3(0,1,0), E4(0,0,1), E5(1,1,0), E6(1,0,1), E7(0,1,1), and E8(1,1,1). The equilibrium points and eigenvalues of the linear dynamic penalty mechanism are shown in Table 4.

Table 4

Equilibrium point of the linear dynamic penalty mechanism

Equilibrium pointλ1λ2λ3
E1(0,0,0)A+n+fH+DT+αnKe+L+βf
E2(1,0,0)AnfH+DT+Q+n+αnK+Le+P+f+βf
E3(0,1,0)B+βfQHD+T(>0)K+βf
E4(0,0,1)A+αnPH+αnKL+e(>0)
E5(1,1,0)Bf+QHD+TQnK+P+f+βf
E6(1,0,1)An+PH+Q+n+αnKL+ePf
E7(0,1,1)BQPH(>0)K(>0)
E8(1,1,1)B+Q+PHQnKPf

3.2.2 Stability analysis of the linear dynamic reward mechanism

For the manufacturer, the dynamic equations are identical to those in Section 3.2.1. Consequently, the conclusions of the stability analysis for the manufacturer are exactly the same. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: F(y)=(2y1)[HD+T+z(DT)x(Q+n)]+y2(3y2)μQ. Let F(y)=0, then y=0, y=1 and x=x0=[HD+T+z(DT)+μQy]/(Q+n). When x=x0, F(y)0. At this point, regardless of the strategy chosen by the third-party recycler, it is in an evolutionary stable state, with strategy selection remaining stable throughout time. When x<x0, y=1 is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when y=0, this is the third-party recycler's evolutionary stable method, who prefers low-level recycling.

The derivative of the cascade utilisation enterprise's replicator dynamic equation with respect to strategy probability is: F(z)=(2z1)[KL+e+y(Le)x(P+f)]+z2(3z2)νP. Let F(z)=0, then z=0, z=1 and x=x0=[KL+e+y(Le)+νPz]/(P+f). When x=x0, F(z)0. At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionary stable state, and the strategy selection remains unchanged over time. When x<x0, z=1 is the evolutionary stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, z=0 is the evolutionary stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation. The equilibrium points and eigenvalues of the linear dynamic reward mechanism are shown in Table 5.

Table 5

Equilibrium point of the linear dynamic reward mechanism

Equilibrium pointλ1λ2λ3
E1(0,0,0)A+n+fH+DTKe+L
E2(1,0,0)AnfH+DT+Q+nK+Le+P+f
E3(0,1,0)B+βfQHD+T(>0)K
E4(0,0,1)A+αnPHKL+e(>0)
E5(1,1,0)Bf+QHD+TQn+μQK+P+f
E6(1,0,1)An+PH+Q+n+αnKL+ePf+νP
E7(0,1,1)BQPH(>0)K(>0)
E8(1,1,1)B+Q+PHQn+μQKPf+νP

3.2.3 Stability analysis of the linear dynamic reward-penalty mechanism

Manufacturer's stability analysis is identical to Section 3.2.1. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: F(y)=(2y1)[HD+T+z(DT)x(Q+n)]+αn(y1)(3y2)+y2(3y2)μQ. Let F(y)=0, then y=0, y=1 and x=x0=[HD+T+z(DT)+μQyαn(1y)]/(Q+n). When x=x0, F(y)0. The third-party recycler remains in an evolutionary stable state regardless of the strategy it chooses, and this state remains unchanged over time. When x<x0, y=1 is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, y=0 is the third-party recycler's evolutionary stable strategy choosing low-level recycling.

For cascade utilisation enterprise, the derivative of the replicator dynamic equation with respect to the strategy probability is: F(z)=(2z1)[KL+e+y(Le)x(P+f)]+βf(z1)(3z2)+z2(3z2)νP. Let F(z)=0, then z=0, z=1 and x=x0=[KL+e+y(Le)+νPzβf(1z)]/(P+f). When x=x0, F(z)0. At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionarily stable state, and the strategy selection remains unchanged over time. When x<x0, z=1 is the evolutionarily stable strategy for the cascade utilisation enterprise, choosing active cascade utilisation. Conversely, z=0 is the evolutionarily stable strategy for the cascade utilisation enterprise, indicating the selection of passive cascade utilisation. The equilibrium points and eigenvalues of the linear dynamic penalty-reward mechanism are shown in Table 6.

Table 6

Equilibrium point of the linear dynamic reward-penalty mechanism

Equilibrium pointλ1λ2λ3
E1(0,0,0)A+n+fH+DT+αnKe+L+βf
E2(1,0,0)AnfH+DT+Q+n+αnK+Le+P+f+βf
E3(0,1,0)B+βfQHD+T(>0)K+βf
E4(0,0,1)A+αnPH+αnKL+e(>0)
E5(1,1,0)Bf+QHD+TQn+μQK+P+f+βf
E6(1,0,1)An+PH+Q+n+αnKL+ePf+νP
E7(0,1,1)BQPH(>0)K(>0)
E8(1,1,1)B+Q+PHQn+μQKPf+νP

The initial parameters of the linear dynamic reward and penalty system are the same as those of the static reward-penalty system. The dynamic penalty coefficient is set to α=1, β=2, and the dynamic reward coefficient is set to μ=2, ν=1. The simulation results of the system's evolution over time under the three reward and penalty mechanisms are shown in Figures 11–13, respectively.

Figure 11
A line graph showing the evolutionary path under a linear dynamic penalty mechanism.A line graph displays the evolutionary path under a linear dynamic penalty mechanism. The x-axis represents time (t) ranging from 0 to 3, and the y-axis represents the variable p ranging from 0 to 1. The graph includes three data lines: one in blue for x equals 0.5, one in green for y equals 0.5, and one in magenta for z equals 0.5. The blue line starts at a p value of approximately 0.4 and decreases to around 0.3, stabilizing around 0.3 at t equals 2.69839. The magenta line starts at a p value of approximately 0.4 and increases to around 0.7, stabilizing around 0.7 at t equals 2.55938. The green line starts at a p value of approximately 0.4 and decreases to around 0.0, stabilizing around 0.0 at t equals 1.5. All values are approximated.

Evolutionary path under linear dynamic penalty mechanism

Figure 11
A line graph showing the evolutionary path under a linear dynamic penalty mechanism.A line graph displays the evolutionary path under a linear dynamic penalty mechanism. The x-axis represents time (t) ranging from 0 to 3, and the y-axis represents the variable p ranging from 0 to 1. The graph includes three data lines: one in blue for x equals 0.5, one in green for y equals 0.5, and one in magenta for z equals 0.5. The blue line starts at a p value of approximately 0.4 and decreases to around 0.3, stabilizing around 0.3 at t equals 2.69839. The magenta line starts at a p value of approximately 0.4 and increases to around 0.7, stabilizing around 0.7 at t equals 2.55938. The green line starts at a p value of approximately 0.4 and decreases to around 0.0, stabilizing around 0.0 at t equals 1.5. All values are approximated.

Evolutionary path under linear dynamic penalty mechanism

Close modal
Figure 12
A line graph showing the evolutionary path under a linear dynamic reward mechanism.A line graph depicts the evolutionary path under a linear dynamic reward mechanism. The horizontal axis represents time (t) ranging from 0 to 2, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines labeled x equals 0.5, y equals 0.5, and z equals 0.5. The blue line labeled x equals 0.5 starts at the origin and rapidly increases, reaching a value close to 1 around t equals 1.5. The green line labeled y equals 0.5 starts at the origin and quickly decreases to near 0, remaining flat thereafter. The magenta line labeled z equals 0.5 starts at the origin, decreases slightly, and then gradually increases to a value around 0.4, remaining relatively flat after t equals 1. A data point is marked at approximately t equals 1.57577 and p equals 0.399139 on the magenta line.

Evolutionary path under linear dynamic reward mechanism

Figure 12
A line graph showing the evolutionary path under a linear dynamic reward mechanism.A line graph depicts the evolutionary path under a linear dynamic reward mechanism. The horizontal axis represents time (t) ranging from 0 to 2, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines labeled x equals 0.5, y equals 0.5, and z equals 0.5. The blue line labeled x equals 0.5 starts at the origin and rapidly increases, reaching a value close to 1 around t equals 1.5. The green line labeled y equals 0.5 starts at the origin and quickly decreases to near 0, remaining flat thereafter. The magenta line labeled z equals 0.5 starts at the origin, decreases slightly, and then gradually increases to a value around 0.4, remaining relatively flat after t equals 1. A data point is marked at approximately t equals 1.57577 and p equals 0.399139 on the magenta line.

Evolutionary path under linear dynamic reward mechanism

Close modal
Figure 13
A line graph showing the evolutionary path under a linear dynamic reward-penalty mechanism.A line graph displays the evolutionary path under a linear dynamic reward-penalty mechanism. The x-axis represents time (t) ranging from 0 to 4, and the y-axis represents the variable p ranging from 0 to 1. The graph includes three data lines: one in blue for x equals 0.5, one in green for y equals 0.5, and one in magenta for z equals 0.5. The blue line starts at the origin, rises steeply, and then levels off around 0.926679 at t equals 3.5875. The magenta line starts at 0.4, rises gradually, and levels off around 0.665992 at t equals 3.5875. The green line starts at 0.4, drops steeply, and then levels off around 0.0742487 at t equals 3.77748. All values are approximated.

Evolutionary path under linear dynamic reward-penalty mechanism

Figure 13
A line graph showing the evolutionary path under a linear dynamic reward-penalty mechanism.A line graph displays the evolutionary path under a linear dynamic reward-penalty mechanism. The x-axis represents time (t) ranging from 0 to 4, and the y-axis represents the variable p ranging from 0 to 1. The graph includes three data lines: one in blue for x equals 0.5, one in green for y equals 0.5, and one in magenta for z equals 0.5. The blue line starts at the origin, rises steeply, and then levels off around 0.926679 at t equals 3.5875. The magenta line starts at 0.4, rises gradually, and levels off around 0.665992 at t equals 3.5875. The green line starts at 0.4, drops steeply, and then levels off around 0.0742487 at t equals 3.77748. All values are approximated.

Evolutionary path under linear dynamic reward-penalty mechanism

Close modal

Under the linear dynamic penalty mechanism, the probability of manufacturers adopting a strategy involving the use of recycled materials stabilises at 0.3, whilst the probability of third-party recyclers opting for low-level processing rapidly converges to 0. This indicates that the penalty mechanism has limited effectiveness in incentivising manufacturers to recycle and fails to effectively curb recyclers' low-level processing behaviour. The probability of cascade utilisation enterprises engaging in active cascade utilisation stabilises at 0.7, reflecting that the dynamic penalty has a relatively significant deterrent effect on the negative behaviour of such enterprises. Overall, this mechanism can only partially improve the willingness to engage in cascade utilisation and lacks effective constraints on recyclers.

Under the linear dynamic reward mechanism, the probability of manufacturers using recycled materials stabilises at 1, indicating that this dynamic reward effectively encourages manufacturers to adopt recycled materials. The probability of third-party recyclers opting for high-level processing still converges to 0, suggesting that rewards alone cannot reverse the tendency of recyclers to choose low-level processing due to cost pressures. The probability of cascade utilisation enterprises engaging in active cascade utilisation stabilises at 0.4, indicating that the incentive effect of the rewards on these enterprises is limited and fails to fully stimulate their enthusiasm. Although this mechanism has enhanced manufacturers' sense of responsibility, it has failed to coordinate other links in the industrial chain.

Under the linear dynamic reward-penalty mechanism, the probability of manufacturers using recycled materials stabilised at 0.93, approaching full adoption of recycled materials. The probability of third-party recyclers employing high-standard processing rose slightly to 0.07, but remained at an extremely low level, indicating that whilst this combined reward-penalty mechanism brought about some improvement for recyclers, the effect was not significant. The probability of cascade utilisation enterprises actively engaging in cascade utilisation stabilised at 0.67, representing an improvement compared to a single reward mechanism.

Overall, the linear dynamic reward and penalty mechanisms can encourage manufacturers to fulfil their recycling responsibilities and exert a certain degree of incentive and constraint on cascade utilisation enterprises. However, they still lack sufficient driving force to promote high-level processing behaviour among third-party recyclers.

The linear dynamic reward-penalty mechanism, as previously indicated, can suppress fluctuations. However, under the three dynamic reward and penalty configurations described, third-party recyclers consistently prefer low-level processing of recovered power batteries. These are not optimal control solutions. Therefore, this study proposes optimising the dynamic reward and penalty control scheme by introducing nonlinear dynamic reward and penalty functions (Chang et al., 2017). The three mechanisms are as follows: nonlinear dynamic penalty, nonlinear dynamic reward, and nonlinear dynamic reward-penalty.

4.1.1 Construction of the nonlinear dynamic penalty mechanism

The manufacturer's nonlinear dynamic penalty function for third-party recyclers is defined as n2=α1n(1y)2+α2y/C21, while the function for cascade utilisation enterprises is f2=β1f(1z)2+β2z/C31, where α1, α2, β1, β2 are the dynamic penalty coefficients, α1>0, α2>0, and β1>0, β2>0. This nonlinear function signifies that the severity of penalties imposed on manufacturers is proportional to the additional investment costs paid for high-level processing of retired power batteries and active cascade utilisation. Specifically, the higher the probability that manufacturers choose to recycle materials for remanufacturing, the greater the penalty. Conversely, the lower the additional investment costs for third-party recyclers and cascade utilisation enterprises, the greater the penalty. The total penalty is reflected in the expected return as (1y)n2 and (1z)f2. The pure strategy payoff expression E11 , E12, E21 , E22, E31 and E32, under the static mechanism remains unchanged. The reward for cascade utilisation enterprises and third-party recyclers remains a fixed constant P, Q. The expected return expressions are as follows:

(29)
(30)
(31)

The replicator dynamics equations for the three participants are, respectively:

(32)
(33)
(34)

4.1.2 Construction of the nonlinear dynamic reward mechanism

The manufacturer's nonlinear dynamic reward function for third-party recyclers is set as Q2=μ1Qy2+μ2y/(1C21), and for cascade utilisation enterprises as P2=ν1Pz2+ν2z/(1C31), where μ1, μ2, ν1, ν2 are the dynamic reward coefficients, μ1>0, μ2>0, ν1>0, and ν2>0. This nonlinear function signifies that manufacturers' reward levels correlate with the additional investment costs incurred by high-level processing of the retired power batteries and active cascade utilisation. Specifically, manufacturers receive higher rewards when they choose to recycle materials for remanufacturing, while third-party recyclers and cascade utilisation enterprises receive higher rewards when their additional investment costs increase. Third-party recyclers and cascade utilisation enterprises receive additional returns of yQ2 and zP2 respectively as a result of their proactive behaviour. The penalty for third-party recyclers and cascade utilisation enterprise remains a fixed constant n, f. The expected return expressions are:

(35)
(36)
(37)

The replicator dynamics equations for the three participants are, respectively:

(38)
(39)
(40)

4.1.3 Construction of the nonlinear dynamic reward-penalty mechanism

The penalty and reward mechanisms are combined. The expected return expressions are as follows:

(41)
(42)
(43)

The replicator dynamics equations for the three participants are, respectively:

(44)
(45)
(45)

4.2.1 Stability analysis of the nonlinear dynamic penalty mechanism

For manufacturer, stability analysis is consistent with Section 3.2.1. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: F(y)=(2y1)[HD+T+z(DT)x(Q+n)]+α1n(1y)2(1+2y)+α2y(23y)/C21. Let F(y)=0, then y=0, y=1 and x=x0=[HD+T+z(DT)α1n(1y)2α2y/C21]/(Q+n). When x=x0, F(y)0. At this point, regardless of the strategy chosen by the third-party recycler, it remains in an evolutionary stable state, and the strategy selection remains unchanged all the time. When x<x0, y=0, F(y)>0, y=1, and F(y)<0. y=1 is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when y=0, this is the third-party recycler's evolutionary stable strategy choosing low-level recycling.

The derivative of the cascade utilisation enterprise's replicator dynamic equation with respect to strategy probability is: F(z)=(2z1)[KL+e+y(Le)x(f+P)]+β1f(1z)2(1+2z)+β2z(23z)/C31. Let F(z)=0, then z=0, z=1 and x=x0=[KL+e+y(Le)β1f(1z)2β2z/C31]/(f+P). When x=x0, F(z)0. At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionarily stable state, and the strategy selection remains unchanged over time. When x<x0, z=1 is the evolutionarily stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, when z=0, this is the evolutionarily stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation. Following Lyapunov's indirect method, the equilibrium points and eigenvalues of the nonlinear dynamic penalty mechanism are shown in Table 7.

Table 7

Equilibrium point of the nonlinear dynamic penalty mechanism

Equilibrium pointλ1λ2λ3
E1(0,0,0)A+n+fH+DT+α1nKe+L+β1f
E2(1,0,0)AnfH+DT+Q+n+α1nK+Le+P+f+β1f
E3(0,1,0)B+βfQHD+Tα2/c21K+β1f
E4(0,0,1)A+αnPH+α1nKL+eβ2/c31
E5(1,1,0)Bf+QHD+TQnα2/c21K+P+f+β1f
E6(1,0,1)An+PH+Q+n+α1nKL+ePfβ2/c31
E7(0,1,1)BQPHα2/c21Kβ2/c31
E8(1,1,1)B+Q+PHQnα2/c21KPfβ2/c31

4.2.2 Stability analysis of the nonlinear dynamic reward mechanism

The manufacturer's stability analysis is consistent with Section 3.2.1. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: F(y)=(2y1)[HD+T+z(DT)x(Q+n)]+μ1Q(4y3)y2+μ2(3y22y)/(1C21). Let F(y)=0, then y=0, y=1 and z=z0=[HD+Tx(Q+n)+μ1Qy2+μ2y/(1C21)]/(TD). When z=z0, F(y)0. At this point, regardless of the strategy chosen by the third-party recycler, it remains in an evolutionary stable state, and the strategy selection remains unchanged over time. When z<z0, y=0, F(y)>0, y=1, and F(y)<0. y=1 is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when y=0, this is the third-party recycler's evolutionary stable strategy, choosing low-level recycling.

For cascade utilisation enterprise, the derivative of the replicator dynamic equation with respect to the strategy probability is: F(z)=(2z1)[KL+e+y(Le)x(P+f)]+ν1P(4z3)z2+ν2(3z22z)/(1C31). Let F(z)=0, then z=0, z=1 and x=x0=[KL+e+y(Le)+ν1Pz2+ν2z/(1C31)]/(f+P). When x=x0, F(z)0. At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionarily stable state, and the strategy selection remains unchanged over time. When x<x0, z=0, F(z)>0, z=1, F(z)<0. At this point, z=1 is the evolutionarily stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, when z=0, this is the evolutionarily stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation. The equilibrium points and eigenvalues are shown in Table 8.

Table 8

Equilibrium point of the nonlinear dynamic reward mechanism

Equilibrium pointλ1λ2λ3
E1(0,0,0)A+n+fH+DTKe+L
E2(1,0,0)AnfH+DT+Q+nK+Le+P+f
E3(0,1,0)B+βfQHD+T+μ1Q+μ21C21(>0)K(<0)
E4(0,0,1)A+αnPH(<0)KL+e+ν1P+ν21C31(>0)
E5(1,1,0)Bf+QHD+TQn+μ1Q+μ21C21K+P+f
E6(1,0,1)An+PH+Q+nKL+ePf+ν1P+ν21C31
E7(0,1,1)BQPH+μ1Q+μ21C21(>0)K+ν1P+ν21C31(>0)
E8(1,1,1)B+Q+PHQn+μ1Q+μ21C21KPf+ν1P+ν21C31

4.2.3 Stability analysis of the nonlinear dynamic reward-penalty mechanism

The stability analysis for the manufacturer is consistent with Section 3.2.1. For third-party recycler, the derivative of the replicator dynamic equation with respect to the strategy probability is: F(y)=(2y1)[HD+T+z(DT)x(Q+n)]+α1n(1y)2(1+2y)+α2y(23y)C21+μ1Q(4y3)y2+μ2(3y22y)1C21. Let F(y)=0, then y=0, y=1 and z=z0=[HD+Tx(Q+n)α1n(1y)2α2y/C21+μ1Qy2+μ2y/(1C21)]/(TD). When z=z0, F(y)0. At this point, regardless of the strategy chosen by the third-party recycler, it remains in an evolutionary stable state, and the strategy selection remains unchanged over time. When z<z0, y=0, F(y)>0, y=1, and F(y)<0. y=1 is the third-party recycler's evolutionary stable strategy choosing high-level recycling. Conversely, when y=0, this is the third-party recycler's evolutionary stable strategy, choosing low-level recycling.

The derivative of the cascade utilisation enterprise's replicator dynamic equation with respect to strategy probability is: F(y)=(2y1)[HD+T+z(DT)x(Q+n)]+β1f(1z)2(1+2z)+β2z(23z)C31+ν1P(4z3)z2+ν2(3z22z)1C31. Let F(z)=0, then z=0, z=1 and x=x0=[KL+e+y(Le)β1f(1z)2β2z/C31+ν1Pz2+ν2z/(1C31)]/(f+P). When x=x0, F(z)0. At this point, regardless of the strategy chosen by the cascade utilisation enterprise, the system remains in an evolutionarily stable state, and the strategy selection remains stable all the time. When x<x0, z=1 is the evolutionarily stable strategy for the cascade utilisation enterprise choosing active cascade utilisation. Conversely, z=0 is the evolutionarily stable strategy for the cascade utilisation enterprise choosing passive cascade utilisation. The equilibrium points and eigenvalues of the nonlinear dynamic reward-penalty mechanism are shown in Table 9.

Table 9

Equilibrium point of the nonlinear dynamic reward-penalty mechanism

Equilibrium pointλ1λ2λ3
E1(0,0,0)A+n+fH+DT+α1nKe+L+β1f
E2(1,0,0)AnfH+DT+Q+n+α1nK+Le+P+f+β1f
E3(0,1,0)B+βfQHD+Tα2c21+μ1Q+μ21C21K+β1f
E4(0,0,1)A+αnPH+α1nKL+eβ2c31+ν1P+ν21C31
E5(1,1,0)Bf+QHD+TQnα2c21+μ1Q+μ21C21K+P+f+β1f
E6(1,0,1)An+PH+Q+n+α1nKL+ePfβ2c31+ν1P+ν21C31
E7(0,1,1)BQPHα2c21+μ1Q+μ21C21Kβ2c31+ν1P+ν21C31
E8(1,1,1)B+Q+PHQnα2c21+μ1Q+μ21C21KPfβ2c31+ν1P+ν21C31

4.3.1 Model simulation analysis of nonlinear dynamic penalty mechanism

Initial parameters are consistent with those used in the static mechanism. The dynamic penalty coefficients are set as follows: α1=10, α2=1, β1=10, and β2=1. Figure 14 shows the simulation results of the nonlinear dynamic penalty mechanism model.

Figure 14
A line graph showing the evolutionary path under a nonlinear dynamic penalty mechanism.A line graph displays the evolutionary path under a nonlinear dynamic penalty mechanism. The horizontal axis represents time (t) ranging from 0 to 1, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines labeled x equals 0.5, y equals 0.5, and z equals 0.5. The blue line labeled x equals 0.5 shows a decreasing trend, starting at around 0.5 and approaching 0 as time increases. The green line labeled y equals 0.5 shows a slight decrease, starting at around 0.5 and stabilizing just below 0.5. The magenta line labeled z equals 0.5 shows an increasing trend, starting at around 0.5 and stabilizing just above 0.6. Two data points are highlighted at t equals 0.971857, with p values of approximately 0.66482 and 0.546261 for the green and blue lines, respectively.

Evolutionary path under nonlinear dynamic penalty mechanism

Figure 14
A line graph showing the evolutionary path under a nonlinear dynamic penalty mechanism.A line graph displays the evolutionary path under a nonlinear dynamic penalty mechanism. The horizontal axis represents time (t) ranging from 0 to 1, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines labeled x equals 0.5, y equals 0.5, and z equals 0.5. The blue line labeled x equals 0.5 shows a decreasing trend, starting at around 0.5 and approaching 0 as time increases. The green line labeled y equals 0.5 shows a slight decrease, starting at around 0.5 and stabilizing just below 0.5. The magenta line labeled z equals 0.5 shows an increasing trend, starting at around 0.5 and stabilizing just above 0.6. Two data points are highlighted at t equals 0.971857, with p values of approximately 0.66482 and 0.546261 for the green and blue lines, respectively.

Evolutionary path under nonlinear dynamic penalty mechanism

Close modal

According to simulation results, under the non-linear dynamic penalty mechanism, the probability of manufacturers using recycled materials stabilised at 0, indicating that this penalty alone is insufficient to motivate manufacturers to proactively adopt recycled materials. The probability of high-level processing by third-party recyclers stabilises at 0.54, a significant improvement over the linear mechanisms, indicating that the non-linear penalty design effectively curbs low-level processing behaviour and encourages more than half of recyclers to opt for technological upgrades. The probability of proactive cascade utilisation by cascade utilisation enterprises stabilises at 0.66, comparable to the level under linear mechanisms. This mechanism significantly improves the constraint effect on recyclers, but provides insufficient incentives for manufacturers.

4.3.2 Model simulation analysis of nonlinear dynamic reward mechanism

Initial parameters were set consistently with those under the static mechanism, with dynamic reward coefficients defined as μ1=10, μ2=1, ν1=10, and ν2=1. Figure 15 shows the simulation results of the nonlinear dynamic reward mechanism model.

Figure 15
A line graph showing the evolutionary path under a nonlinear dynamic reward mechanism.A line graph with three data lines representing different variables over time. The x-axis is labeled 't' and ranges from 0 to 1. The y-axis is labeled 'p' and ranges from 0 to 1. The blue line with diamond markers represents 'x equals 0.5', the green line with square markers represents 'y equals 0.5', and the magenta line with cross markers represents 'z equals 0.5'. The blue line starts at approximately 0.45 on the y-axis and gradually decreases to near 0 as time progresses. The green line starts at approximately 0.45 on the y-axis and rapidly decreases to near 0 within the first 0.2 units of time. The magenta line starts at approximately 0.45 on the y-axis and quickly increases to near 1 within the first 0.1 units of time, remaining constant thereafter. All values are approximated.

Evolutionary path under nonlinear dynamic reward mechanism

Figure 15
A line graph showing the evolutionary path under a nonlinear dynamic reward mechanism.A line graph with three data lines representing different variables over time. The x-axis is labeled 't' and ranges from 0 to 1. The y-axis is labeled 'p' and ranges from 0 to 1. The blue line with diamond markers represents 'x equals 0.5', the green line with square markers represents 'y equals 0.5', and the magenta line with cross markers represents 'z equals 0.5'. The blue line starts at approximately 0.45 on the y-axis and gradually decreases to near 0 as time progresses. The green line starts at approximately 0.45 on the y-axis and rapidly decreases to near 0 within the first 0.2 units of time. The magenta line starts at approximately 0.45 on the y-axis and quickly increases to near 1 within the first 0.1 units of time, remaining constant thereafter. All values are approximated.

Evolutionary path under nonlinear dynamic reward mechanism

Close modal

In this mechanism, the probability of manufacturers using recycled materials and the probability of third-party recyclers achieving high-level processing both stabilise at 0, indicating that the non-linear rewards have virtually no incentive effect on manufacturers or recyclers. Conversely, the probability of cascade utilisation enterprises actively engaging in cascade utilisation stabilises at 1, suggesting that the non-linear rewards generate a strong positive incentive for these enterprises, causing them to be fully inclined towards active cascade utilisation. This mechanism is capable of driving the cascade utilisation market to an optimal state on its own, but it fails to improve manufacturers' willingness to recycle or recyclers' processing standards.

4.3.3 Model simulation analysis of nonlinear dynamic reward-penalty mechanism

Initial parameters were set consistently with those under the static mechanism. The dynamic penalty coefficients were defined as α1=10, α2=1, β1=1, and β2=1, while dynamic reward coefficients were set as μ1=10, μ2=1, ν1=1, and ν2=1. Figure 16 shows the simulation results of the nonlinear dynamic reward-penalty mechanism model.

Figure 16
A line graph showing the evolutionary path under a nonlinear dynamic reward-penalty mechanism.A line graph depicts the evolutionary path under a nonlinear dynamic reward-penalty mechanism. The horizontal axis represents time (t) ranging from 0 to 3, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines, each representing different initial conditions: x equals 0.5, y equals 0.5, and z equals 0.5. The blue line with diamond markers represents x equals 0.5, showing an upward trend starting from approximately 0.5 and stabilizing around 0.86. The green line with square markers represents y equals 0.5, showing a slight upward trend starting from approximately 0.5 and stabilizing around 0.65. The magenta line with cross markers represents z equals 0.5, showing a downward trend starting from approximately 0.5 and stabilizing around 0.33. The graph includes labels for specific data points, such as X 2.87805 and Y 0.864839 for the blue line, X 1.68639 and Y 0.65443 for the green line, and X 2.74513 and Y 0.325632 for the magenta line.

Evolutionary path under nonlinear dynamic reward-penalty mechanism

Figure 16
A line graph showing the evolutionary path under a nonlinear dynamic reward-penalty mechanism.A line graph depicts the evolutionary path under a nonlinear dynamic reward-penalty mechanism. The horizontal axis represents time (t) ranging from 0 to 3, and the vertical axis represents the variable p ranging from 0 to 1. The graph includes three data lines, each representing different initial conditions: x equals 0.5, y equals 0.5, and z equals 0.5. The blue line with diamond markers represents x equals 0.5, showing an upward trend starting from approximately 0.5 and stabilizing around 0.86. The green line with square markers represents y equals 0.5, showing a slight upward trend starting from approximately 0.5 and stabilizing around 0.65. The magenta line with cross markers represents z equals 0.5, showing a downward trend starting from approximately 0.5 and stabilizing around 0.33. The graph includes labels for specific data points, such as X 2.87805 and Y 0.864839 for the blue line, X 1.68639 and Y 0.65443 for the green line, and X 2.74513 and Y 0.325632 for the magenta line.

Evolutionary path under nonlinear dynamic reward-penalty mechanism

Close modal

Under the non-linear dynamic reward-penalty mechanism, the probability of manufacturers producing recycled materials stabilises at 0.86, a significant improvement on the previous two single mechanisms. This demonstrates that combining rewards and penalties can effectively incentivise manufacturers to fulfil their recycling responsibilities, as they face both revenue from penalties imposed on non-compliant parties and expenditure on rewards for compliant parties; their net revenue is closely linked to the behaviour of recyclers and cascade utilisation enterprises. The probability of high-level processing by third-party recyclers stabilises at 0.65, the highest level across all mechanisms, indicating that the combination of non-linear rewards and penalties simultaneously discourages low-level processing and incentivises high-level processing, thereby creating a powerful driving force for recyclers. The probability of proactive cascade utilisation by cascade utilisation enterprises stabilises at 0.33, which is relatively low. Nevertheless, this mechanism performs well in balancing the interests of the three parties, enhancing the technical capabilities of recyclers and the participation of manufacturers, thereby achieving a relatively ideal state of systemic equilibrium.

As the central entity, manufacturers should design internal reward and penalty systems to translate government subsidies into dynamic incentives and constraints for upstream and downstream players. For recyclers, a combination of rewards and penalties is more effective than a single approach in driving technological upgrades; for cascade utilisation enterprises, a medium-level equilibrium may arise when the intensity of rewards and penalties is mismatched. Regulators need to fine-tune the reward and penalty coefficients to avoid unduly dampening enthusiasm. In practice, enterprises can dynamically adjust reward and penalty parameters according to the cost structures of different stages: offering higher rewards for costly technological upgrades and imposing progressively stricter penalties on behaviour prone to “free-riding”, thereby achieving overall optimisation at the supply chain level. For policymakers, they should encourage the adoption of dynamic internal incentive and penalty schemes by providing technical guidance and promoting information sharing across the supply chain, and validate the practical effectiveness of different reward and penalty functions through pilot projects, thereby providing empirical evidence to support industry-wide adoption.

This research centres on the reward and penalty methods in the power battery recycling. A tripartite evolutionary game model involving manufacturers, third-party recyclers, and cascaded utilisation enterprises is developed to examine the revenues, costs, and strategic stability of each party under different strategy choices. The model's validity was evaluated using numerical simulation analysis. This research investigates how various reward and penalty mechanisms influence the evolutionary trajectory of the system. The main conclusions include:

The system exhibits cyclical fluctuations and struggles to achieve stable cooperation under the static reward-penalty mechanism. The linear dynamic mechanism provides some motivation for manufacturers and cascade utilisation enterprises to recycle materials for remanufacturing and to actively engage in cascade utilisation, but it fails to provide effective incentives or constraints for third-party recyclers. Nonlinear dynamic reward-penalty mechanisms, by dynamically linking reward and penalty intensity to strategy selection, is more effective in guiding third-party recyclers to improve their processing standards, whilst significantly enhancing manufacturers' willingness to recycle, thereby achieving a relatively ideal systemic equilibrium.

Based on these findings, manufacturers should proactively assume primary responsibility for recycling. By thoroughly considering stakeholders' interests and behavioural patterns, they should design internal reward and penalty mechanisms. Concurrently, manufacturers can establish long-term, stable partnerships with third-party recyclers and cascade utilisation enterprises, and jointly invest resources in technological research and development to enhance the quality of recycled materials and the market competitiveness of cascade utilisation products. Furthermore, governments and relevant institutions should provide the technology and platforms to facilitate information sharing within the supply chain.

This paper has limitations: First, the research focuses only on manufacturers, recyclers, and utilisation enterprises, excluding other stakeholders like governments and consumers. Second, the model assumes bounded rationality with static parameters, ignoring dynamic market fluctuations.

Chang
,
J.
,
Zhao
,
L.
and
Du
,
J.
(
2017
), “
Regulatory evolutionary game analysis and stability control of corporate environmental behaviour: a system dynamics approach
”,
Systems Engineering
, Vol. 
35
No. 
10
, pp. 
79
-
87
.
Guan
,
Y.
,
He
,
T.H.
and
Hou
,
Q.
(
2023
), “
Tripartite evolutionary game analysis of power battery cascade utilisation under government subsidies
”,
IEEE Access
, Vol. 
11
, pp. 
66382
-
66399
, doi: .
Hu
,
K.
,
Hou
,
Q.
and
Yu
,
S.
(
2025
), “
Research on closed-loop supply chain pricing and profit distribution of power battery echelon utilisation based on noncooperative–cooperative biform game
”,
International Game Theory Review
, Vol. 
28
No. 
02
, 2550012,
(pre-publish)
, doi: .
Jiang
,
S.
,
Zhang
,
L.
,
Hua
,
H.
,
Liu
,
X.
,
Wu
,
H.
and
Yuan
,
Z.
(
2021
), “
Assessment of end-of-life electric vehicle batteries in China: future scenarios and economic benefits
”,
Waste Management
, Vol. 
135
, pp. 
70
-
78
, doi: .
Liu
,
J.
and
Ma
,
J.
(
2021
), “
Research on reverse subsidy mechanisms in closed-loop supply chains for power batteries considering tiered utilisation
”,
Industrial Engineering and Management
, Vol. 
26
No. 
3
, pp. 
80
-
88
, doi: .
Liu
,
J.
and
Zhu
,
L.
(
2024
), “
Research on power battery recycling mode selection considering dual behavioural preferences under different government subsidies
”,
International Journal of Low-Carbon Technologies
, Vol. 
19
, pp. 
1579
-
1595
, doi: .
Ma
,
L.
,
Wang
,
J.
and
Yang
,
J.
(
2026
), “
Research on recycling decisions in closed-loop supply chains of power batteries based on blockchain technology
”,
Journal of Energy Storage
, Vol. 
141
,
PA
, 118869, doi: .
Ministry of Industry and Information Technology, et al.
(
2024
), “
Announcement No. 42 of 2024: industry standard conditions for comprehensive utilisation of waste power batteries from new energy vehicles (2024 Edition)
”,
Surface Engineering and Remanufacturing
, Vol. 
24
No. 
6
, pp. 
51
-
55
.
Ministry of Industry and Information Technology, et al.
(
2026
), “
Interim measures for the management of recycling and comprehensive utilisation of waste power batteries from new energy vehicles
”,
Resource Regeneration
, No. 
1
, pp. 
40
-
45
.
Tang
,
Y.
,
Zhang
,
Q.
,
Li
,
Y.
,
Li
,
H.
,
Pan
,
X.
and
Mclellan
,
B.
(
2019
), “
The social-economic-environmental impacts of recycling retired EV batteries under reward-penalty mechanism
”,
Applied Energy
, Vol. 
251
, 113313, doi: .
Tian
,
T.
,
Zheng
,
C.
,
Yang
,
L.
,
Luo
,
X.
and
Lu
,
L.
(
2022
), “
Optimal recycling channel selection of power battery closed-loop supply chain considering corporate social responsibility in China
”,
Sustainability
, Vol. 
14
No. 
24
, 16712, doi: .
Wei
,
H.
and
Qi
,
Z.
(
2025
), “
Pricing strategy for sustainable recycling of power batteries considering recycling competition under the reward–penalty mechanism
”,
Sustainability
, Vol. 
17
No. 
16
, p.
7224
, doi: .
Wen
,
D.
,
Wang
,
Z.
,
Wang
,
M.
,
Appolloni
,
A.
and
Xiao
,
T.
(
2025
), “
Online collection channel and government subsidy strategy of retired power batteries considering cascade utilisation
”,
Computers and Industrial Engineering
, Vol. 
209
, 111417, doi: .
Wu
,
W.
and
Zhang
,
M.
(
2025
), “
Decision analysis of closed-loop supply chains for power batteries under carbon allowance trading and subsidy policies
”,
Chinese Journal of Management Science
, Vol. 
33
No. 
8
, pp. 
340
-
354
, doi: .
Wu
,
W.
,
Li
,
M.
and
Huang
,
G.Q.
(
2025
), “
Optimal recovery mode for new energy vehicle battery recycling under government policies
”,
Managerial and Decision Economics
, Vol. 
46
No. 
4
, pp. 
2629
-
2642
, doi: .
Xu
,
N.
,
Xu
,
Y.
and
Zhong
,
H.
(
2023a
), “
Pricing decisions for power battery closed-loop supply chains with low-carbon Input by echelon utilisation enterprises
”,
Sustainability
, Vol. 
15
No. 
23
, 16544, doi: .
Xu
,
Y.
,
Ma
,
N.
,
Feng
,
Y.
,
Gao
,
M.
,
Chen
,
L.
and
Wang
,
G.
(
2023b
), “
Analysis of waste power battery Issues based on EPR system
”,
China Storage and Transport
, No. 
10
, pp. 
199
-
200
, doi: .
Yan
,
Y.
,
Cao
,
J.
,
Zhou
,
Y.
,
Zhou
,
G.
and
Chen
,
J.
(
2024
), “
Decisions for power battery closed-loop supply chain: cascade utilisation and extended producer responsibility
”,
Annals of Operations Research
, pp. 
1
-
41
, doi: .
Yang
,
K.
,
Zhang
,
W.
and
Zhang
,
Q.
(
2022
), “
Research on incentive policies for power battery recycling considering cascade utilisation
”,
Industrial Engineering and Management
, Vol. 
27
No. 
2
, pp. 
1
-
8
, doi: .
Yu
,
S.
and
Hou
,
Q.
(
2023
), “
A closed-loop power battery supply chain differential game model considering echelon utilisation under a cost subsidy
”,
Kybernetes
, Vol. 
52
No. 
8
, pp. 
2826
-
2846
, doi: .
Yu
,
H.
and
Wang
,
S.
(
2025
), “
Blockchain-enabled closed-loop supply chain optimisation for power battery recycling and cascading utilisation
”,
Sustainability
, Vol. 
17
No. 
9
, p.
4192
, doi: .
Yuan
,
W.
and
Zhang
,
X.
(
2025
), “
Research on policy and technology development of retired power battery recycling system
”,
Nonferrous Metals
, Vol. 
15
No. 
6
, pp. 
1081
-
1086
, doi: .
Zhang
,
Z.
and
Liang
,
H.
(
2023
), “
Research on coordination of the NEV battery closed-loop supply chain considering CSR and fairness concerns in third-party recycling models
”,
Scientific Reports
, Vol. 
13
No. 
1
, 22172, doi: .
Zhang
,
Z.
,
Guo
,
M.
and
Yang
,
W.
(
2022
), “
Analysis of NEV power battery recycling under different government reward-penalty mechanisms
”,
Sustainability
, Vol. 
14
No. 
17
, 10538, doi: .
Zhang
,
Z.
,
Wang
,
Y.
,
Guo
,
Y.
and
Song
,
H.
(
2023a
), “
Power battery closed-loop supply chain decision of green investment under different subsidy objects
”,
Procedia Computer Science
, Vol. 
221
, pp. 
1162
-
1169
, doi: .
Zhang
,
M.
,
Wu
,
W.
and
Song
,
Y.
(
2023b
), “
Study on the impact of government policies on power battery recycling under different recycling models
”,
Journal of Cleaner Production
, Vol. 
413
, 137492, doi: .
Zhang
,
C.
,
Tian
,
Y.
and
Cui
,
M.
(
2024a
), “
Hybrid Channel recycling model selection and carbon reduction decision-making for electric vehicle power battery manufacturers
”,
Chinese Journal of Management Science
, Vol. 
32
No. 
6
, pp. 
184
-
195
, doi: .
Zhang
,
W.
,
Zhu
,
L.
,
Liu
,
X.
,
Wang
,
W.
and
Song
,
H.
(
2024b
), “
Optimal strategies in electric vehicle battery closed-loop supply chain considering government subsidies and echelon utilisation
”,
Journal of Energy Storage
, Vol. 
99
,
PB
, 113341, doi: .
Published in Modern Supply Chain Research and Applications. Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at Link to the terms of the CC BY 4.0 licence.

or Create an Account

Close Modal
Close Modal