Use the existing dataset(s) and apply DL techniques to classify data into benign and malicious. Also do a multi-classification using either DL or ruled based analysis (kind of semi auto, you can run Python scripts to do this part as well) to perform deep analysis about the email sources, contents, nature of engineering techniques, formation of email messages, crafts etc..
Dataset: https://github.com/rokibulroni/Phishing-Email-Dataset.git https://research.utwente.nl/en/datasets/phishing-validation-emails-dataset/