Article navigation
Purpose

Customers often communicate with a company by sending it emails, filling out the “contact us” form on the company’s website, and posting messages on social media platforms. Unfortunately, a substantial part of these incoming customer messages can be spam. Whether an email is authentic (i.e. ham) or spam is specific to the company. Consequently, commercial spam classifiers (e.g. the one Microsoft Outlook uses) do not work well. A machine learning-based spam classifier can solve this problem.

Design/methodology/approach

I used a large dataset from a publicly traded retailer in the United States to test 27 TF-IDF-based and six word embedding-based binary classifiers.

Findings

I found that RoBERTa – a sophisticated embedding-based classifier – provided the lowest false positive rate of 5.31%. However, its rival classifier – XLNet – offered marginally superior (92%) balanced accuracy, relative to RoBERTa’s 91%. To enhance the use of the models by other organizations and academics, I offer supplemental models trained on only external datasets and a combination of internal and external datasets.

Research limitations/implications

The manuscript contributes by applying a broad set of existing machine learning models to solve a real problem for a real company.

Practical implications

I publish my code (via the journal), which other companies and researchers can use.

Originality/value

The manuscript is original in comparing a broad range of machine learning classifiers on real data from a real company.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$39.00
Rental

or Create an Account

Close Modal
Close Modal