Style Description based Text-to-Speech with Conditional Prosodic Style Normalization based Diffusion GAN