LLM-Enhanced Multi-Agent Transfer Reinforcement Learning for Sensing, Communication, Computing, and Control Co-Optimization in Cyber-Physical Systems
Junyuan Zhang, Chi Xu, Haibin Yu
Shenyang Institute of Automation
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Dear Editor, With the rapid development of new-generation information and communication technology, the fusion of sensing, communication, computing, and control (S3C) is becoming increasingly significant for cyber-physical systems (CPS). However, due to the non-convexity, the curse of dimensionality, and the partial observability faced by CPS, traditional convex optimization algorithms are challenging to deal with S3C co-optimization. Thus, this letter establishes the S3C problem as a partially observable Markov decision process (POMDP) and proposes a large language model (LLM)-enhanced policy transfer (PT) framework for multi-agent deep reinforcement learning (MADRL), denoted as LLMPT-MADRL. Herein, MADRL learns a co-optimization policy while LLM-enhanced policy transfer improves the exploration efficiency, as LLM utilizes large-scale dynamic neural neurons like human brain to form powerful intelligent reasoning capabilities. However, MADRL and LLM exhibit heterogeneity in neural network scale. Thus, we further design an adaptive policy transfer loss function to achieve dynamic adaptation between MADRL and LLM, avoiding the mismatch of policy transfer information. Furthermore, we conduct a case study on the co-optimization of sampling coefficient, communication bandwidth, computing frequency, and control input to minimize the system delay and control error of CPS. Experimental results show that the reward of LLMPT-MADRL is increased by 16.8%, the system delay is reduced by 24.5%, and the control error is reduced by 10.4% compared with the benchmark MADRL-based algorithms.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIReinforcement Learning in Robotics
Adaptive Dynamic Programming Control · Neural Networks and Reservoir Computing
参考文献 8
此处列出前 3 条