After completing this learning module, you will be able to:
Explain how generative AI can be used to create deepfake phishing emails and why this is dangerous.
Describe how a guard / moderation LLM (Qwen-Guard) can classify messages as DEEPFAKE or LEGIT using few-shot examples.
Use a controlled UI to generate synthetic phishing emails and classify both generated and real emails.
Recognize the limitations of LLM-based phishing detection and how prompt design and examples influence its behavior.
Understand why all phishing content used in the lab must be synthetic, non-actionable, and safely constrained.
Python 3 via Google Colab (no local install required).
Basic familiarity with:
Large Language Models and prompting (prompts, sampling, temperature).
The idea of a “generator” model vs. a “guard” or “moderator” model.
Access to:
This Case 3 notebook / script, which loads Qwen/Qwen2.5-1.5B-Instruct as the generator model.
Internet connection to download Hugging Face models and use transformers.
Optionally, you may bring a few legitimate emails (e.g., meeting reminders, invoices, newsletters) whose text you can safely paste into the classification box for testing.
Deepfake Phishing Content
Synthetic, AI-generated text that imitates a phishing email (password reset, urgent invoice, money transfer, etc.). In this case, it is not used against real targets, but only for testing defenses.
Synthetic Phishing Email
A phishing-style message produced by the generator LLM using predefined templates such as urgent_transfer, password_reset, and fake_invoice. These include fake placeholders like <link> instead of real URLs.
Attack Mode
The part of the notebook that uses Qwen to generate phishing-like emails via generate_text, prompt templates, a sentinel token, and the synthesize_phish helper.
Defense Mode / Guard LLM
The part of the notebook that uses qwen_guard (built on Qwen again) to classify a given email as DEEPFAKE, LEGIT, or UNCERTAIN, using a small set of few-shot examples and simple heuristics.
Few-Shot Examples
A small labeled set of short emails (some legit, some phishing) that are embedded in the guard’s prompt to guide its behavior toward consistent labels. These are stored in FEW_SHOT_EXAMPLES.
Sentinel Token (<END_EMAIL>)
A special marker instructing the generator to stop at a clean boundary. The lab uses GEN_SENTINEL = "<END_EMAIL>" plus build_email_prompt and finalize_generated_email to cut off stray text and keep output as a clean email.
Unsanitized vs. Sanitized Input (in this case)
Unsanitized: Any email text (generated or pasted) the guard sees without checks.
Sanitized: Email text that you have ensured is synthetic or safe to use in testing (e.g., no real credentials, no real links).
The objective of Case 3 is:
“To demonstrate how generative AI can be used to create and detect deepfake phishing content, highlighting the risks and defenses surrounding AI-generated social-engineering attacks.”
Modern LLMs can generate emails that look and feel like real corporate communication: password resets, billing issues, delivery notices, and more. This is powerful but risky:
Attackers could use these models to mass-produce convincing phishing messages.
Defenders, however, can also use LLMs to generate synthetic phishing emails for testing and to classify messages as potentially malicious.
Case 3 focuses on this duality. You will work with:
An Attack Model (Qwen-Instruct) that generates synthetic phishing emails from templates.
A Defense Model (Qwen-Guard-style classifier) that labels each email as DEEPFAKE, LEGIT, or UNCERTAIN.
The important constraint is that all phishing content in the lab is synthetic and non-actionable: no real names, no real credentials, and no real company infrastructure are involved.
The notebook defines several prompt templates in TEMPLATES (e.g., "urgent_transfer", "password_reset", "fake_invoice"). Each template describes a type of phishing scenario and instructs the model to output a Subject line and body, including a placeholder like <link>.
The core generator function is:
generate_text(...):
Takes a prompt, calls Qwen using model.generate, and returns the generated text after removing the echoed prompt and trimming at optional stop sequences.
Uses sampling parameters like temperature and repetition_penalty to control variability.
To cleanly shape emails, the notebook then uses:
GEN_SENTINEL = "<END_EMAIL>"
build_email_prompt(template_text) to wrap the instruction and demand a Subject + body, ending with <END_EMAIL>.
finalize_generated_email(text) to:
Cut off at the sentinel.
Ensure the email ends on a sentence boundary.
Remove obvious instruction leakage above the Subject: line.
The helper synthesize_phish(template_key, n, max_new_tokens) uses these tools to generate multiple synthetic phishing emails from the chosen template with small variations (e.g., adding “This is time-sensitive” or “Please respond ASAP”).
On the defense side, the notebook defines:
FEW_SHOT_EXAMPLES: A small list of emails labeled LEGIT or DEEPFAKE. These provide pattern examples for Qwen-Guard.
qwen_guard(candidate_text, few_shot=True):
Builds a prompt telling the model: “You are a cybersecurity classifier. Answer with exactly ONE WORD: DEEPFAKE or LEGIT.”
Includes the few-shot examples, followed by the candidate email and Label:.
Calls generate_text with do_sample=False (greedy, deterministic) to get a short answer.
If it can’t parse DEEPFAKE or LEGIT from the answer, it falls back to simple keyword heuristics (e.g., detecting CLICK, RESET, PASSWORD, <LINK>) and either returns DEEPFAKE or UNCERTAIN.
This function is the moderation/guard LLM for Case 3. It doesn’t block model generation itself but evaluates the safety/legitimacy of emails after they’re written.
The notebook exposes this behavior through a simple widget-based UI:
A Prompt text box where you type a generation request (or a template key like "password_reset").
A Generate button and token slider that produce a synthetic phishing email using generate_text or synthesize_phish.
An output area that shows the generated email.
A second Paste box where you can paste any email (generated or real text).
A Classify Email button that runs qwen_guard on the pasted text and prints a verdict (e.g., Verdict: DEEPFAKE).
This lets you experiment with both synthetic phishing emails and legit emails to see how the guard behaves.
Case 3 doesn’t deploy a full production system, but it demonstrates important defensive ideas:
Use synthetic data for testing
You never need real victims or real accounts to study phishing behavior.
Synthetic phishing + synthetic placeholders for names and links are sufficient to test classifiers.
Combine templates with a guard model
Templates ensure generated phishing emails stay within known, controlled patterns (password resets, invoices, transfers).
The guard LLM (with few-shot examples) gives a realistic demonstration of how content filters can flag risky messages.
Acknowledge limitations
The guard’s decisions depend heavily on the prompt and the few-shot examples.
Some phishing emails may be misclassified as LEGIT or UNCERTAIN, showing why additional layers (keyword filters, URL checks, reputation systems) are needed in real deployments.
This pre-lab prepares you to think about both how AI can assist attackers