We want to see the impact of the setting of our MARL framework. One important part is the dual-pool of both questions and templates. The other is the reward function, which combines IQ and judgement score.
We compare our results for the above settings, the result is as follow: