We compare our approach with five SOTA jailbreak attack framework. The results are as follow:
As we select the 15 most effective templates to calculate the final ASR, the strategy for choosing these templates is based on the scores provided by our original judgment model, GPT-4o-mini. When using other models to assess the success of an attack, the ASR may decrease compared to the assessment made by GPT-4o-mini. However, our method still demonstrates strong performance compared to other frameworks, particularly in commercial models.Â