We want to wonder whether our approach is available to attack the model using a new harmful question. We use Deepseek-chat as the target model and select 15 harmful question from AdvBench. And then we calculate the average IQ and reward of each iteration which is successful attack.
We notice that the average accumulated IQ and reward increase as the epoch increases.