Table 3

Machine learning predictive models of bot likelihood given speech topics

Likelihood of bot speaker – original tweets sampleLikelihood of bot speaker – all tweets sample
(1)(2)
Relative strength of relationship with bot likelihood (0 → 1)Direction of relationship with bot likelihood (+/−)Relative strength of relationship with bot likelihood (0 → 1)Direction of relationship with bot likelihood (+/−)
Topic 80.02+0.23+
Topic 170.06–0.23+
Topic 40.13+0.05+
Topic 180.02+0.11+
Topic 10.45+0.4+
Topic 120.03–0.24+
Topic 31+1+
Topic 210.17–0.17+
Topic 60.01+0.02+
Topic 230.04–0.39+
Topic 190.32–0.09–
Topic 130.03–0.04–
Topic 70.02–0.06+
Topic 240.07+0.07+
Topic 90.22+0.12+
Topic 140.05+0.03+
Topic 220+0.01+
Topic 100.01+0.02+
Topic 110.01+0.03+
Topic 160–0.59+
Topic 150.01+0–
Topic 200.34+0.2+
Topic 50.03+0.02+
Topic 20.02–0–
Topic 250.02–0.01–
N3,978,1039,260,886
Model accuracy91.3%83.6%

Note(s): Table 3 shows results from a support vector machine (SVM) classification model that predicts speaker type (bot or non-bot) based on the speakers' use of the 25 topics. Model Accuracy shows the accuracy of the model in predicting bots on a 20% “hold-out” sample employing 5-fold cross-validation. The Relative Strength column shows the SVM coefficients “normalized” on a 0 → 1 scale, such that the variable with the strongest predictive values is given a score of “1,” the variable with the weakest effect is given a score of “0,” and the remaining variables a score between 0 and 1 based on their strength relative to the weakest and strongest variables. The column thus permits a comparison of the relative predictive effect of the 25 topic models. The Direction of Relationship column shows the positive or negative direction of each topic's association with bots

Source(s): Authors' own work

or Create an Account

Close subscription notice
Close access options