We show the distribution of different types of generated data by KuK in the following figures (only the distribution of generated data on NIN is shown in the paper).
From the figures, you could see that we did not generate BEs with HL type ( High PCS and Low VRO) or AEs with LH type (Low PCS and High VRO), which are both data with common uncertainty patterns.
Based on the metric performance of existing AEs/BEs, the objectives we set for the lower bound of High PCS is 0.7, the lower bound of High VRO is 0.6, the upper bound of Low PCS is 0.3, the upper bound of Low VRO is 0.4.
different types of data generated by KuK
different types of data generated by KuK
different types of data generated by KuK
different types of data generated by KuK
Despite the objective setting, from the results, we could see that KuK can generate HH and HL data with very high PCS, which is comparable to existing benign data, as well as LH and LL data with pretty low PCS. Same observations can also be obtained w.r.t metric VRO. (For data that disobey the objectives, we remove those in the following applications and evaluation. )
For the generated data with LL type, the overall VRO metric values for AEs (with orange color) are relatively greater than BEs (with green color), especially for NIN and ResNet20 models, which shows the usefulness of uncertainty estimates for capturing the intrinsic characteristic of adversarial defects of DL softwares. For the same reason, AEs with LL type are much harder to generate than BEs with LL type.
For HH type data, there is no clear tendency that shows the difference between the difficulty of generating HH BEs and HH AEs. For NIN and ResNet20 model, HH AEs is much easier to generate, while for LeNet5 and MobileNet, HH BEs is easier to generate. To a certain extent, the combination of single-shot and multi-shot execution-based metrics could reveal the difference of behaviors among different models.
Based on the study towards uncertainty patterns of existing AEs/BEs, we use KuK to generate data with uncommon patterns.
The following figures show the generated benign/adversarial data through KuK and their seeds towards HH/HL/LH/LL types, where only adversarial data are generated for HL type and only benign data are generated for LH type.
For each model, we show the seed in the first line and the four types of generated data in the following four lines. In each column, four types of data are generated with the same seed in the first line.
seed
HH type
HL type (AEs)
LH type (BEs)
LL type
seed
HH type
HL type (AEs)
LH type (BEs)
LL type
seed
HH type
HL type (AEs)
LH type (BEs)
LL type
seed
HH type
HL type (AEs)
LH type (BEs)
LL type
For the task of hand-written digit classification, HL AEs are the most difficult ones to generate. It is mainly because the relative simplicity of MNIST dataset. Moreover, we forbid affine transformations and set the upper bound of perturbation. As a result, tasks with higher complexity provide larger space for manipulation. From the figures of seeds and generated data shown above, we could see that the perturbations are all imperceptible to human eyes, especially for the most complex task: ImageNet classification.