Figure 6.7.
Line chart showing cumulative rewards over 500,000 timesteps, with multiple curves comparing baseline vs. reinforcement-learning performance for total, energy, and thermal-comfort rewards, all plotted against a shared cumulative reward axis.
Cumulative reward (baseline vs. RL approach).

or Create an Account

Close Modal
Close Modal