The attack time cost increases as the model parameters grow in size. Given the substantial time commitment, we explored whether the top templates identified from one model or the weights of reinforcement learning agents could be transferred effectively to attack another model.
For template transfer attack, it is partially successful across different Llama models. The Top1- and Top5- ASR success rates vary, with Llama-3.1-8b showing the highest transfer effectiveness (97% Top1 and 98% Top5) and Llama-3.1-405b exhibiting the lowest (64% Top1 and 81% Top5).
Specifically, when agents transfer knowledge from other tasks to the current one, they exhibit enhanced resilience against attacks. This advantage likely stems from the agents' ability to leverage previous experiences, making them more adaptable and effective in new scenarios.