Efficient and accurate facial landmark estimation is crucial for various embedded systems applications. The proposed approach in this paper achieves improved performance by iteratively refining the training data labels and reducing the model size while minimizing the computational resources required for deployment. Experimental results demonstrate the effectiveness of the presented method in optimizing facial landmark estimation for embedded systems, paving the way for more efficient and accurate facial analysis applications in resource-constrained environments. In particular, these strategies notably propelled us to secure the top place and second position in the facial landmark detection qualification and final competitions.Â