Abstract:Under the dual pressures of the global energy crisis and environmental pollution, electric vehicles have emerged as an important development direction in the transportation sector, thanks to their zero-emission and high energy-efficiency characteristics. However, electric vehicles still face challenges, including high energy consumption, limited driving range, and reduced comfort in dynamic, complex environments. Traditional energy efficiency optimization methods, such as rule-based control strategies, are computationally simple but lack flexibility. Optimization-based approaches can approximate global optimality but struggle with real-time adaptability due to their high computational complexity. Existing deep reinforcement learning methods can handle high-dimensional state spaces. However, these methods often suffer from energy waste and driving instability due to inadequate environmental perception or abrupt changes in action. This study proposes an innovative framework that integrates environmental perception with deep reinforcement learning to achieve efficient energy conservation and enhanced driving comfort in complex road conditions. To address the challenge of processing high-dimensional traffic data, a city road-condition perception module based on a World Model is proposed. The encoder is trained using both contrastive learning and coupled reinforcement learning, enabling the agent to perceive real-time traffic changes and generate control strategies aligned with energy-saving objectives. In addition, the study incorporates an active learning framework that uses a KL-divergence-based strategy to select more difficult training samples. Periodic offline training is then employed to optimize the policy. This intelligent approach significantly enhances the energy efficiency of electric vehicles under urban road conditions. The proposed method has been validated on the CARLA simulation platform. The results demonstrate that it can effectively extract key environmental features and accelerate model convergence. Under real-world driving conditions, vehicle energy consumption, SOC variation, and battery degradation rate are improved. Meanwhile, statistical analyses of real-time acceleration and jerk data show that the proposed approach yields smooth control behavior and substantially reduces the frequency of sudden acceleration and deceleration events. In conclusion, this study proposes an efficient, smooth, and adaptive energy efficiency optimization framework for electric vehicles. By deeply integrating environmental perception with active reinforcement learning, the framework demonstrates innovation in low-dimensional traffic representation and achieves breakthroughs in optimizing multi-objective reward functions. The results confirm that the method achieves significant energy savings and improves driving comfort in dynamic, complex environments, providing strong theoretical support for the practical application of intelligent driving systems. Future research will focus on optimizing perception models for multimodal traffic scenarios, thereby enabling fully intelligent, energy-efficient driving technologies. Additionally, drivers' personalized driving habits significantly impact vehicle energy consumption and battery lifespan. Future work will integrate driving behavior modeling with energy-saving strategies by learning from drivers' operational preferences, aiming to achieve intelligent control strategies that balance energy efficiency with personalized adaptability. Meanwhile, hardware-in-the-loop(HIL)platforms will be used to systematically evaluate the algorithm's real-time performance, stability, and reliability. By deploying the policy network on embedded controllers and testing its response latency and fault-recovery capabilities in dynamic traffic environments, the method can be ensured to provide sufficient safety and robustness for engineering implementation.
彭自然, 范泽宇. 融合环境感知强化学习的电动汽车能效优化控制[J]. 电工技术学报, 2026, 41(16): 5712-5728.
Peng Ziran, Fan Zeyu. Energy-Efficient Control for Electric Vehicles Incorporating Environment-Aware Reinforcement Learning. Transactions of China Electrotechnical Society, 2026, 41(16): 5712-5728.
[1] Garau M, Torsæter B N.A methodology for optimal placement of energy hubs with electric vehicle charging stations and renewable generation[J]. Energy, 2024, 304: 132068. [2] Zhang Hengwei, Zhang Yisheng, Wang Zhigang, et al.A novel knowledge-driven flexible human-robot hybrid disassembly line and its key technologies for electric vehicle batteries[J]. Journal of Manufacturing Systems, 2023, 68: 338-353. [3] 侯文博, 杨平, 陈可, 等. 电动汽车驱动复用升压技术[J]. 电工技术学报, 2024, 39(增刊1): 95-105. Hou Wenbo, Yang Ping, Chen Ke, et al.Multiplexing booster technology for electric vehicle drive[J]. Transactions of China Electrotechnical Society, 2024, 39(S1): 95-105. [4] Ferloni A.Transitions as a coevolutionary process: the urban emergence of electric vehicle inventions[J]. Environmental Innovation and Societal Transitions, 2022, 44: 205-225. [5] 彭自然, 杨肖阳, 肖伸平. 基于EKF-HInformer模型估计汽车动力电池的SOC&SOH[J]. 电子测量与仪器学报, 2025, 39(3): 21-33. Peng Ziran, Yang Xiaoyang, Xiao Shenping.SOC and SOH of the battery are estimated based on the EKF-HInformer model[J]. Journal of Electronic Measurement and Instrumentation, 2025, 39(3): 21-33. [6] Wu Yue, Huang Zhiwu, Zheng Yusheng, et al.Spatial-temporal data-driven full driving cycle predi-ction for optimal energy management of battery/supercapacitor electric vehicles[J]. Energy Conversion and Management, 2023, 277: 116619. [7] 陈中, 万玲玲, 张梓麒. 面向共享电动汽车的用户助推与充电协同调度[J]. 电工技术学报, 2025, 40(11): 3572-3590. Chen Zhong, Wan Lingling, Zhang Ziqi.Nudging users and charging optimization for electric car-sharing system scheduling[J]. Transactions of China Electrotechnical Society, 2025, 40(11): 3572-3590. [8] Yang Yiqin, Hu Hao, Li Wenzhe, et al.Flow to control: offline reinforcement learning with lossless primitive discovery[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2023, 37(9): 10843-10851. [9] Ada S E, Oztop E, Ugur E.Diffusion policies for out-of-distribution generalization in offline reinfor-cement learning[J]. IEEE Robotics and Automation Letters, 2024, 9(4): 3116-3123. [10] Kim T, Park Y, Park Y, et al.Acceleration of actor-critic deep reinforcement learning for visual grasping by state representation learning based on a preprocessed input image[C]//2021 IEEE/RSJ Inter-national Conference on Intelligent Robots and Systems(IROS), Prague, Czech Republic, 2021: 198-205. [11] Nair A, Pong V, Dalal M, et al.Visual reinforcement learning with imagined goals[C]//Proceedings of the 32nd International Conference on Neural Information Processing Systems, Montréal, Canada, 2018: 9209-9220. [12] Yang J, Lee G, Chang S, et al.Towards governing agent's efficacy: action-conditional β-VAE for deep transparent reinforcement learning[C]//Asian Confer-ence on Machine Learning, PMLR, 2019: 32-47. [13] Fujimoto S, Chang W D, Smith E J, et al.For SALE: state-action representation learning for deep reinfor-cement learning[C]//Proceedings of the 37th Inter-national Conference on Neural Information Pro-cessing Systems, Red Hook, NY, 2023: 61573-61624. [14] Zheng R, Wang X, Sun Y, et al.Temporal latent action-driven contrastive loss for visual reinforcement learning[J]. Advances in Neural Information Pro-cessing Systems, 2023, 36: 48203-48225. [15] Gelada C, Kumar S, Buckman J, et al.Deepmdp: learning continuous latent space models for repre-sentation learning[C]//International Conference on Machine Learning, PMLR, 2019: 2170-2179. [16] Peng Ziran, He Zhenyu.Optimization of regenerative braking control strategy for dual-motor electric vehicles based on deep reinforcement learning[J]. IEEE Transactions on Intelligent Transportation Systems, 2025, 26(7): 10954-10967. [17] Peng Ziran, Yang Xiaoyang.Short-and medium-term power load forecasting model based on a hybrid attention mechanism in the time and frequency domains[J]. Expert Systems with Applications, 2025, 278: 127329. [18] Macias Fernandez A, Kandidayeni M, Boulon L, et al.An adaptive state machine based energy management strategy for a multi-stack fuel cell hybrid electric vehicle[J]. IEEE Transactions on Vehicular Tech-nology, 2020, 69(1): 220-234. [19] 彭自然, 贺振宇, 肖伸平, 等. 基于深度强化学习模型TD3优化和改进的电动汽车制动能量回收策略[J]. 控制与决策, 2025, 40(8): 2361-2372. Peng Ziran, He Zhenyu, Xiao Shenping, et al.Electric vehicle brake energy recovery strategy based on deep reinforcement learning model TD3 optimization and improvement[J]. Control and Decision, 2025, 40(8): 2361-2372. [20] 陈剑, 杜文娟, 王海风. 采用深度迁移学习定位含直驱风机次同步振荡源机组的方法[J]. 电工技术学报, 2021, 36(1): 179-190. Chen Jian, Du Wenjuan, Wang Haifeng.A method of locating the power system subsynchronous oscillation source unit with grid-connected PMSG using deep transfer learning[J]. Transactions of China Electro-technical Society, 2021, 36(1): 179-190. [21] 吴延波, 韩志伟, 王惠, 等. 基于双延迟深度确定性策略梯度的受电弓主动控制[J]. 电工技术学报, 2024, 39(14): 4547-4556. Wu Yanbo, Han Zhiwei, Wang Hui, et al.Active pantograph control of deep reinforcement learning based on double delay depth deterministic strategy gradient[J]. Transactions of China Electrotechnical Society, 2024, 39(14): 4547-4556. [22] 彭自然, 王顺豪, 肖伸平. 基于SDAE-DCPInformer的电动汽车电池SOC和SOH估算方法[J]. 智能系统学报, 2025, 20(4): 969-983. Peng Ziran, Wang Shunhao, Xiao Shenping.SOC and SOH estimation method of electric vehicle battery based on SDAE-DCPInformer[J]. CAAI Transactions on Intelligent Systems, 2025, 20(4): 969-983. [23] 张薇, 王浚宇, 杨茂, 等. 基于分布式双层强化学习的区域综合能源系统多时间尺度优化调度[J]. 电工技术学报, 2025, 40(11): 3529-3544. Zhang Wei, Wang Junyu, Yang Mao, et al.The multi-time-scale optimal scheduling for regional integrated energy system based on the distributed bi-layer reinforcement learning[J]. Transactions of China Electrotechnical Society, 2025, 40(11): 3529-3544. [24] 杜伟, 王圣, 李健, 等. 基于CNN-LSTM-AM模型的储能锂离子电池荷电状态预测[J]. 电工技术学报, 2025, 40(9): 2982-2995. Du Wei, Wang Sheng, Li Jian, et al.Prediction of state of charge for energy storage lithium-ion batteries based on CNN-LSTM-AM model[J]. Trans-actions of China Electrotechnical Society, 2025, 40(9): 2982-2995. [25] Liu T, Wang B, Yang C.Online Markov Chain-based energy management for a hybrid tracked vehicle with speedy Q-learning[J]. Energy, 2018, 160: 544-555. [26] 陈泽宇, 方志远, 杨瑞鑫, 等. 基于深度强化学习的混合动力汽车能量管理策略[J]. 电工技术学报, 2022, 37(23): 6157-6168. Chen Zeyu, Fang Zhiyuan, Yang Ruixin, et al.Energy management strategy for hybrid electric vehicle based on the deep reinforcement learning method[J]. Transactions of China Electrotechnical Society, 2022, 37(23): 6157-6168. [27] Chen Bin, Wang Miaoben, Hu Lin, et al.A hierarchical cooperative eco-driving and energy management strategy of hybrid electric vehicle based on improved TD3 with multi-experience[J]. Energy Conversion and Management, 2025, 326: 119508. [28] Wu Yitao, Liu Yonggang, Peng Jiang, et al.Enhanced hierarchical eco-driving control for electric vehicles via global velocity optimization with road slope adaptation[J]. IEEE Transactions on Transportation Electrification, 2025, 11(4): 10322-10335. [29] Yan Su, Fang Jiayi, Yang Chao, et al.Eco-driving for connected automated hybrid electric vehicles in learning-enabled layered transportation systems[J]. Transportation Research Part D: Transport and Environment, 2025, 142: 104677. [30] Haarnoja T, Zhou A, Abbeel P, et al.Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor[C]//International Con-ference on Machine Learning, Pmlr, 2018: 1861-1870. [31] Duan Jingliang, Guan Yang, Li S E, et al.Distri-butional soft actor-critic: off-policy reinforcement learning for addressing value estimation errors[J]. IEEE Transactions on Neural Networks and Learning Systems, 2022, 33(11): 6584-6598. [32] 陈嘉琛, 陈中, 李冰融, 等. 基于二阶随机动力学的多虚拟电厂自趋优能量管理策略[J]. 中国电机工程学报, 2024, 44(16): 6294-6306. Chen Jiachen, Chen Zhong, Li Bingrong, et al.Energy management strategy for multi-virtual power plants with self-optimization based on second-order stochastic dynamics[J]. Proceedings of the CSEE, 2024, 44(16): 6294-6306.