Research proposal: detecting phishing emails with transformer models
A focused research proposal with a narrow question, a gap drawn from the literature, a feasible method and a realistic project plan.
- Paper type
- Research Proposal
- Subject
- Computer Science
- Level
- Master's
- Length
- 2,500 words
- Pages
- 9 pages
- Referencing
- IEEE
The brief
Write a research proposal for your MSc project. Include background, research questions, a brief literature review, methodology, ethical considerations and a timeline. 2,500 words, IEEE referencing.
Why this sample works
What a marker would single out, and what to look for as you read.
- A research question narrow enough to answer in one project
- Gap identified from specific limitations in prior work
- Evaluation plan with named metrics and baselines
- Ethics and data handling addressed concretely
Contents
- 01Background and motivationIn preview
- 02Research questionsIn preview
- 03Related work
- 04Methodology and evaluation
- 05Ethical considerations
- 06Project plan and risks
Preview
Research proposal: detecting phishing emails with transformer models
1. Background and motivation
Phishing remains one of the most common initial routes into organisational networks. Rule-based and classical machine learning filters catch much of it, but they rely on features such as known malicious URLs and suspicious keywords that attackers adapt to quickly. Transformer-based language models, which learn contextual representations of text [1], [2], offer a way to detect phishing from the persuasive structure of a message rather than from surface features.
Most published work fine-tunes these models on public corpora that are several years old and dominated by crude, mass-mailed phishing. Far less is known about how well they detect targeted messages that imitate internal communications. This project addresses that gap.
2. Research questions
- RQ1: How does a fine-tuned transformer classifier compare with a TF-IDF and logistic regression baseline on targeted phishing emails?
- RQ2: How much does performance degrade when a model trained on public corpora is tested on recent, organisation-style phishing?
- RQ3: Which features of a message most influence the model's classification, as indicated by attribution methods?
The rest of this sample is sent on request
Free, privately, usually within the hour. Kept off the web so it never turns up in a similarity check.
Request the full sampleReferences (extract, IEEE)
- [1] A. Vaswani et al., "Attention is all you need," in Advances in Neural Information Processing Systems, vol. 30, 2017.
- [2] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. NAACL-HLT, 2019, pp. 4171–4186.
Samples are reference material written for a different brief. Use them to understand structure and argument; submitting one, in whole or in part, would be an academic integrity breach.
All samples