For various applications, AI models sometimes need to be fine tuned for their specific use cases. Because AI models are general models, they may not always give the best responses to a specialized application.
I used the unsloth library to download the models. The specific model I used in this case was Alibaba's Qwen3 model having 14 billion parameters. It's performance is similar or sometimes better than Deepseek.
In order to set up the fine tuning processes, a formal prompt template had to be defined. This was the general structure of the response. The task that the model was fine tuned on was competitive programming, as the model was given a problem and it had to give an accurate answer to such a problem.
This is the test question I used to evaluate the model's response
This was the model's response pre fine tuning:
thinking process:
Actual response:
The thinking process was way too long for this response, which is the main issue of the model. With fine tuning, we can condense the model's thinking process to be much shorter.
Now when fine tuning the model, I obviously couldn't adjust all 14 billion parameters as that would be way too computationally expensive. So I had to apply some optimization methods in order to fine tune the model in my computationally limited environment.
I first applied quantization, which essentially reduced the percision of the model's weights. Quantization essentially reduces the percision of the weights of each number
In my project, I reduced the precision of the weights to FP16, meaning that all numbers could only be represented through 16 bits instead of the standard 32 bit in most LLMs.
Another optimization technique I used was LoRA or Low Rank Adaptation. What LoRA does is that it uses a matrix multiplication trick in order to reduce the number of parameters needed to adjust a larger number of the model's parameters
Here the main model weights are stored in a matrix W, whose applied change is decomposed into the smaller matrices A and B. We only have to adjust the A and B matrices elements in order to change the model's overall weights.
When LoRA is applied to an LLM, the entire model's weights are frozen and only a few modules have the A and B matrices. These are the only matrices that can be changed during training.
The training overall was very short and the model's losses declined to a stable state after some fine tuning.
thinking part:
Actual response:
The model's overall response length became much more shorter both in thinking and actual response, making it much more suitable for competitive programming. Hence as a result of fine tuning, the model was able to learn to respond more concisely.