摘要: The China initiative Accelerator Driven System (CiADS) is a major research facility for nuclear waste transmutation in China, whose stable operation relies on precise control of beam transport. To overcome the limitations of conventional beam-tuning methods in dynamic adaptability and intelligent control, a data-driven multi-agent accelerator control approach is investigated. A digital simulation environment for the medium-energy beam transport section is first established using a cascaded back-propagation neural network to accurately reproduce the beam transport dynamics. On this basis, a multi-agent reinforcement learning control system based on the multi-agent soft actor-critic (MASAC) algorithm is developed to achieve online adaptive tuning and optimization of the beam trajectory. Experimental results demonstrate that the proposed reinforcement learning model can autonomously learn effective control policies and significantly improve beam trajectory regulation. For 50 sets of test data, an average beam-position tuning gain of 12.54\% is achieved, resulting in a beam trajectory that is more closely concentrated around the central transport axis. The proposed method provides a promising approach for intelligent operation and control of large-scale accelerator facilities, offering high accuracy, real-time adaptability, and robust control performance.