github link: https://github.com/Thomasche69/School_rag
For this project, I wanted to create a fully local system where students could utilize the chatbot without any risks of data or privacy breaches. I also wanted to give the students the ability to upload a PDF, so that the LLM would be able read from it. This project was more software engineering than actual Machine Learning training like the other projects.
Back then, LLMs like ChatGPT were more or less toys. Nowadays, they have become useful because they now have different tools. One of the first tools made for LLMs was Retrieval Augmented Generation (RAG). This tool essentially allowed LLMs to read entire PDFs no matter their size. So how does it work?
You as the user upload a PDF. The information of this PDF is put into the document store. Your PDF is first broken down into chunks, each chunk is then turned into a vector. You then give the chatbot a prompt, the document store then uses cosine similarity search in order to find which chunk is most relevant to your prompt. It then takes the most relevant chunks and passes it to the LLM chatbot. Using the chunks, the model then gives an accurate response based on the PDF's content.
I set up seperate APIs hosting the LLM and the embedding model used to turn the PDFs chunks into vectors. The main programe whenever it needs to use the LLM, gives the LLM endpoint some text and it returns the LLM response. When the main programe needs to turn the chunks into vectors, it calls the embedding endpoint, passes it the texts and the endpoint returns the vectors.
LLM endpoint:
Embeddings endpoint:
There are some other minor details in the implementation of the website, but those are the essential components needed to intuitively understand how the website works.Â
Here is a demonstration of its functionality: