Article navigation
Purpose

Spam messages on social networking sites (SNS) cause financial loss and social unrest. While spam classification has been studied widely, scholars have not studied their diffusibility. Diffusibility of spam messages is their ability to spread quickly and reach a large audience. Diffusibility is measured by incorporating the virality and popularity of messages.

Design/methodology/approach

Deep learning (DL) techniques are used to extract novel features from 84,810 spam tweets. Diffusibility is computed from retweets and likes. SHAP and permutation importance analysis are used to determine the features that impact diffusibility. The Mann–Kendall test is performed to test the drift. A machine learning (ML) model for classifying diffusibility is developed through stages of improvement.

Findings

The statistical analysis of the features shows a significant difference between the high and low diffusibility classes of spam tweets. The novel features that are extracted from the tweet content and the user profile predict the diffusibility of spam tweets with an accuracy of 89.5%. Expressions in the profile image (PI) and the pronounceability of screen name (SN) rank among the top predictors of spam. A moving window approach is able to address the problem of drift and improve the prediction of diffusibility.

Originality/value

The paper presents the application of deep learning to extract novel features, such as the profile images' blurriness and the SN's pronounceability. It pioneers the study of the diffusibility of spam tweets that will help diminish the spread of spam tweets. The study aims to reduce risks, such as financial loss and social unrest, that is fueled by fast-spreading spam tweets.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$41.00
Rental

or Create an Account

Close Modal
Close Modal