Word-level (XGBoost) and topic model (LDA) speech patterns
| Topic | Overlap of LDA topic words and XGBoost top 200 features |
|---|---|
| Topic 8 | |
| Topic 17 | theother (10), polic (162) |
| Topic 4 | post (5), photo (2), wall (56), movement (91) |
| Topic 18 | tax (37), million (79) |
| Topic 1 | teaparti (4), tcot (14) |
| Topic 12 | govern (26), liberti (54), peopl (29) |
| Topic 3 | temp (30), forecast (7) |
| Topic 21 | |
| Topic 6 | bring (124), peace (90), bridg (61) |
| Topic 23 | bank (178) |
| Topic 19 | gop (35), job (73), liber (54) |
| Topic 13 | |
| Topic 7 | arrest (18), happi (118) |
| Topic 24 | contact (48) |
| Topic 9 | kpop (22), chang (94) |
| Topic 14 | money (123), politician (71) |
| Topic 22 | |
| Topic 10 | |
| Topic 11 | |
| Topic 16 | |
| Topic 15 | court (171) |
| Topic 20 | livestream (25), viewer (9) |
| Topic 5 | monsanto (69) |
| Topic 2 | |
| Topic 25 |
| Topic | Overlap of LDA topic words and XGBoost top 200 features |
|---|---|
| Topic 8 | |
| Topic 17 | theother (10), polic (162) |
| Topic 4 | post (5), photo (2), wall (56), movement (91) |
| Topic 18 | tax (37), million (79) |
| Topic 1 | teaparti (4), tcot (14) |
| Topic 12 | govern (26), liberti (54), peopl (29) |
| Topic 3 | temp (30), forecast (7) |
| Topic 21 | |
| Topic 6 | bring (124), peace (90), bridg (61) |
| Topic 23 | bank (178) |
| Topic 19 | gop (35), job (73), liber (54) |
| Topic 13 | |
| Topic 7 | arrest (18), happi (118) |
| Topic 24 | contact (48) |
| Topic 9 | kpop (22), chang (94) |
| Topic 14 | money (123), politician (71) |
| Topic 22 | |
| Topic 10 | |
| Topic 11 | |
| Topic 16 | |
| Topic 15 | court (171) |
| Topic 20 | livestream (25), viewer (9) |
| Topic 5 | monsanto (69) |
| Topic 2 | |
| Topic 25 |
Note(s): Table compares words in our LDA-based topics to words identified through an XGBoost random forest machine learning model (Chen et al., 2016). The XGBoost model tells us which words in the entire tweet corpus distinguish bots from other speakers, and Table 6 shows which of the top 200 words from the XGBoost model are also found in our 25 LDA-based topic words (numbers in parentheses indicate XGBoost feature rank). The results allow a comparison of the intersection of the word-level and topic-level analysis of what distinguishes bot from non-bot speaking patterns. For the topics that do not contain a word-level influential feature (e.g. Topic 8), we can assume that it is the more holistic level topic that drives Table 3 results rather than a specific word(s)
Source(s): Authors' own work
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.