The single-agent DQN strategy outperforms the non-RL setting in both Top1-ASR and Top5-ASR across both LLMs. Additionally, the multi-agent MADDPG strategy achieves even greater improvements, demonstrating the effectiveness of reinforcement learning in enhancing attack success rates.