Table 6

Word-level (XGBoost) and topic model (LDA) speech patterns

TopicOverlap of LDA topic words and XGBoost top 200 features
Topic 8 
Topic 17theother (10), polic (162)
Topic 4post (5), photo (2), wall (56), movement (91)
Topic 18tax (37), million (79)
Topic 1teaparti (4), tcot (14)
Topic 12govern (26), liberti (54), peopl (29)
Topic 3temp (30), forecast (7)
Topic 21 
Topic 6bring (124), peace (90), bridg (61)
Topic 23bank (178)
Topic 19gop (35), job (73), liber (54)
Topic 13 
Topic 7arrest (18), happi (118)
Topic 24contact (48)
Topic 9kpop (22), chang (94)
Topic 14money (123), politician (71)
Topic 22 
Topic 10 
Topic 11 
Topic 16 
Topic 15court (171)
Topic 20livestream (25), viewer (9)
Topic 5monsanto (69)
Topic 2 
Topic 25 

Note(s): Table compares words in our LDA-based topics to words identified through an XGBoost random forest machine learning model (Chen et al., 2016). The XGBoost model tells us which words in the entire tweet corpus distinguish bots from other speakers, and Table 6 shows which of the top 200 words from the XGBoost model are also found in our 25 LDA-based topic words (numbers in parentheses indicate XGBoost feature rank). The results allow a comparison of the intersection of the word-level and topic-level analysis of what distinguishes bot from non-bot speaking patterns. For the topics that do not contain a word-level influential feature (e.g. Topic 8), we can assume that it is the more holistic level topic that drives Table 3 results rather than a specific word(s)

Source(s): Authors' own work

or Create an Account

Close subscription notice
Close access options