Machine learning predictive models of bot likelihood given speech topics
| Likelihood of bot speaker – original tweets sample | Likelihood of bot speaker – all tweets sample | |||
|---|---|---|---|---|
| (1) | (2) | |||
| Relative strength of relationship with bot likelihood (0 → 1) | Direction of relationship with bot likelihood (+/−) | Relative strength of relationship with bot likelihood (0 → 1) | Direction of relationship with bot likelihood (+/−) | |
| Topic 8 | 0.02 | + | 0.23 | + |
| Topic 17 | 0.06 | – | 0.23 | + |
| Topic 4 | 0.13 | + | 0.05 | + |
| Topic 18 | 0.02 | + | 0.11 | + |
| Topic 1 | 0.45 | + | 0.4 | + |
| Topic 12 | 0.03 | – | 0.24 | + |
| Topic 3 | 1 | + | 1 | + |
| Topic 21 | 0.17 | – | 0.17 | + |
| Topic 6 | 0.01 | + | 0.02 | + |
| Topic 23 | 0.04 | – | 0.39 | + |
| Topic 19 | 0.32 | – | 0.09 | – |
| Topic 13 | 0.03 | – | 0.04 | – |
| Topic 7 | 0.02 | – | 0.06 | + |
| Topic 24 | 0.07 | + | 0.07 | + |
| Topic 9 | 0.22 | + | 0.12 | + |
| Topic 14 | 0.05 | + | 0.03 | + |
| Topic 22 | 0 | + | 0.01 | + |
| Topic 10 | 0.01 | + | 0.02 | + |
| Topic 11 | 0.01 | + | 0.03 | + |
| Topic 16 | 0 | – | 0.59 | + |
| Topic 15 | 0.01 | + | 0 | – |
| Topic 20 | 0.34 | + | 0.2 | + |
| Topic 5 | 0.03 | + | 0.02 | + |
| Topic 2 | 0.02 | – | 0 | – |
| Topic 25 | 0.02 | – | 0.01 | – |
| N | 3,978,103 | 9,260,886 | ||
| Model accuracy | 91.3% | 83.6% | ||
| Likelihood of bot speaker – original tweets sample | Likelihood of bot speaker – all tweets sample | |||
|---|---|---|---|---|
| (1) | (2) | |||
| Relative strength of relationship with bot likelihood (0 → 1) | Direction of relationship with bot likelihood (+/−) | Relative strength of relationship with bot likelihood (0 → 1) | Direction of relationship with bot likelihood (+/−) | |
| Topic 8 | 0.02 | + | 0.23 | + |
| Topic 17 | 0.06 | – | 0.23 | + |
| Topic 4 | 0.13 | + | 0.05 | + |
| Topic 18 | 0.02 | + | 0.11 | + |
| Topic 1 | 0.45 | + | 0.4 | + |
| Topic 12 | 0.03 | – | 0.24 | + |
| Topic 3 | 1 | + | 1 | + |
| Topic 21 | 0.17 | – | 0.17 | + |
| Topic 6 | 0.01 | + | 0.02 | + |
| Topic 23 | 0.04 | – | 0.39 | + |
| Topic 19 | 0.32 | – | 0.09 | – |
| Topic 13 | 0.03 | – | 0.04 | – |
| Topic 7 | 0.02 | – | 0.06 | + |
| Topic 24 | 0.07 | + | 0.07 | + |
| Topic 9 | 0.22 | + | 0.12 | + |
| Topic 14 | 0.05 | + | 0.03 | + |
| Topic 22 | 0 | + | 0.01 | + |
| Topic 10 | 0.01 | + | 0.02 | + |
| Topic 11 | 0.01 | + | 0.03 | + |
| Topic 16 | 0 | – | 0.59 | + |
| Topic 15 | 0.01 | + | 0 | – |
| Topic 20 | 0.34 | + | 0.2 | + |
| Topic 5 | 0.03 | + | 0.02 | + |
| Topic 2 | 0.02 | – | 0 | – |
| Topic 25 | 0.02 | – | 0.01 | – |
| 3,978,103 | 9,260,886 | |||
| 91.3% | 83.6% | |||
Note(s): Table 3 shows results from a support vector machine (SVM) classification model that predicts speaker type (bot or non-bot) based on the speakers' use of the 25 topics. Model Accuracy shows the accuracy of the model in predicting bots on a 20% “hold-out” sample employing 5-fold cross-validation. The Relative Strength column shows the SVM coefficients “normalized” on a 0 → 1 scale, such that the variable with the strongest predictive values is given a score of “1,” the variable with the weakest effect is given a score of “0,” and the remaining variables a score between 0 and 1 based on their strength relative to the weakest and strongest variables. The column thus permits a comparison of the relative predictive effect of the 25 topic models. The Direction of Relationship column shows the positive or negative direction of each topic's association with bots
Source(s): Authors' own work
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.