We chose four state-of-the-art adversarial attack methods, i.e., FGSM (Fast Gradient Sign Method), BIM (Basic Iterative Method), Deepfool and C&W attacks, to generate adversarial examples. We used the existing Python toolkit foolbox (Foolbox: A Python toolbox to benchmark the robustness of machine learning models) to perform these attacks.
Each attack is configured with the default setting as follows:
FGSM: 1000 number of steps towards the direction of the sign of the gradient, with step size between 0 and maximum-step-size 1.
BIM: limit on the perturbation size as 0.3; step size for gradient descent as 0.05; number of iterations for each gradient descent run as 10.
Deepfool: maximum number of steps to perform as 100; limit on the number of the most likely classes considered as 10.
C&W: maximum iterations as 1000; confidence of adversarial examples as 0; learning rate for the attack algorithm as 5e-3; initial tradeoff-constant to use to tune the relative importance of distance and confidence as 1e-2.
In the defense performance evaluation, 9,000 benign data from test data set and 9,000 adversarial data generated through attacking methods as well as benign/adversarial data generated through KuK are evaluated.
We first provide the overall success rates towards each defense technique for existing data (column Comm) and data generated by KuK (column Unco), which are also shown in Section 5.2 of the paper.
The table above shows the success rates of defense techniques on the existing data and generated data for NIN, ResNet-20 and LeNet5. The average reduction rates among the defense techniques for these models are 0.433, 0.448 and 0.395, respectively. We can see that the defense techniques are very effective in identifying BEs and AEs with common uncertainty patterns, while they perform poorly on the generated data with uncommon patterns.
The table above shows the success rates of the defense techniques on the existing data and generated data for MobileNet. We can see that the success rate is relatively low for MobileNet as it is more challenging to perform the defense for more complex models. The highest reduction rate is 0.316, which occurs at mutation-based adversarial example detection, while the average reduction rate across the defense techniques for MobileNet is 0.15. It is mainly because the label change ratio used as threshold in mutation-based detection technique is similar to VRO metric. Data with uncommon patterns of VRO could bypass this defense technique more easily.
Next we provide more detailed evaluation results about the success rates of defense w.r.t each specific uncertainty type towards data generated through KuK. The row #data indicates the number of generated data for each type which satisfy the objectives (constraints) we set in Section 5.1 of the paper. The success rates of each techinique for benign/adversarial data with different types are shown separately.
Success rate of the defense techniques on each type of the generated data for NIN
Success rate of the defense techniques on each type of the generated data for ResNet-20
We could find that different kinds of defense techniques show uneven defense capacity for data with different types. For example, binary classifier perform well on LL AEs in NIN, ResNet-20 and LeNet5 models. However, the success rates for other data types are pretty low, especially for NIN and ResNet-20. In ResNet-20 model, the success rates of data with other types are all below 0.276, where the success rates for HL AEs and HH/LH/LL BEs are less than 0.065. In NIN model, the success rates on HL AEs and LH/LL BEs are all below 0.17 and the success rates of HH AEs and BEs are about o.5. Meanwhile, defensive distillation and feature squeezing present the worst defense performance towards LL AEs in both NIN and ResNet-20 models. In LeNet5 model, the success rates for HL AEs and LH/LL BEs are all less than 0.012.
Success rate of the defense techniques on each type of the generated data for LeNet5
Success rate of the defense techniques on each type of the generated data for MobileNet
From the experiment results, we could see that data with certain types are more effective in bypassing defense techniques than other types. For example, HH and HL AEs for MobileNet model are more effective than data with other types in bypassing the defense techniques. The success rates towards HH and HL AEs are less than 0.2 for all defense techniques. LL AEs show a fairly good performance with the success rates of all defense less than 0.4. Meanwhile, HH BEs perform poorly on bypassing the defense techniques, where the minimum success rate is 0.929.