Run a Colab notebook that demonstrates how malicious files can inject hidden instructions into an LLM’s context. The lab simulates an AI resume checker and shows how hidden payloads in DOCX/PDF files or metadata can override the model’s intended task. For the final section, you’ll see how a Moderation LLM (QwenGuard) can act as a specialized defense:
Model: Qwen/Qwen2.5-3B-Chat (Hugging Face) — "Resume-review" chat model. Notebook auto-selects GPU if available (Colab runtime: GPU recommended but optional).
Qwen/Qwen3Guard-Gen-4B (Hugging Face) — Moderation guard model
Python runtime: Google Colab (Python 3) — no local install required.
Core libraries:
transformers, torch (model inference)
python-docx, pdfplumer, PyMuPDF, pypdf (safe .docx parsing)
re, unicodedata (sanitization / regex checks)
pytesseract / pdf2image (optional OCR for image/PDF text)
ipywidgets (UI controls)
UI elements (recommended): File uploader, token sliders, Generate buttons, metadata injection text box, output panels.
Data: Prepare a small set of test files to upload during the lab: one benign and one or more malicious files containing hidden payloads embedded in body text, comments, or with invisible characters.
Colab Link (M2): https://colab.research.google.com/drive/1WyJYTB6goLQCqYWiZC6qrT9GOIHZVqdi?usp=sharing
If prompted, Run all to install/import dependencies.
What this section does
Extracts all text from a DOCX file (body, headers, footers, tables, textboxes, comments, footnotes, endnotes).
Feeds it into the acting resume checker with no filtering
Steps
Upload a benign DOCX resume → Generate → observe normal summary
Upload a malicious DOCX resume with hidden instructions → Generate → observe if the model follows them.
Observe the model’s response. Record whether the indirect injection was effective and feel free to make another attempt.
What this section does
Extracts text from PDF files using multiple methods (pdfplumber, PyMuPDF, OCR).
Feeds it directly into the resume checker.
Steps
Upload benign PDF → Generate → observe normal summary.
Upload malicious PDF with hidden instructions → Generate → observe model behavior.
Observe the model’s response. Record whether the indirect injection was effective and feel free to make another attempt.
What this section does
Allows injection of malicious instructions into metadata fields (DOCX comments or PDF subject).
Extracts metadata + file text and feeds both into the chat model
Steps
Use the metadata injection widget to add a hidden instruction.
Download the injected file and re‑upload it
Generate → observe how metadata influences the model’s response
Expected Observation
Metadata injection is especially dangerous because metadata is rarely reviewed by humans but still ingested by the model.
What this section does
Uses QwenGuard to classify file input as Safe / Unsafe / Controversial and assign categories (e.g., Jailbreak, PII, Violent).
If unsafe, the chat model is blocked from responding.
If safe, the chat model generates a response.
Steps
Upload benign file → Guard labels input Safe → chat model responds normally.
Upload malicious file → Guard labels input Unsafe → response blocked
Expected Observation
Moderation LLMs act as gatekeepers, preventing unsafe inputs from reaching the responder model. They are efficient and purpose‑built, but can still miss cleverly disguised payloads.
👉 Click here to see the result of each scenario (Post-Lab)