Skip to article sections

This book describes some of the recent work that has come from applying advanced computer methods to traditional arts and humanities subjects. It is a collection of research articles centred on the concept that, given enough monkeys and enough typewriters, the monkeys would eventually reproduce Shakespeare. The research disproves that, but does prove that you can learn something about texts by counting words that you cannot learn any other way. What the articles show is that computer‐aided word frequency/keyword analysis of large corpus texts can provide an objective means of uncovering patterns that stimulate theory formulation and can suggest areas for further qualitative methods.

The science of word frequency analysis has been transformed by availability of almost limitless numbers of texts to analyse. Computers have enabled researchers to build text databases of unprecedented size, upwards of four million words, and provided the tools to mine them in new ways. In almost every case computerised repetition of manual text analysis has thrown up new insights and challenged old ideas.

Modern researchers do more than just count words in the corpus. The chapters in the book describe not only how words can be counted in context, by linguistic origin and many other ways; how words and phrases can be tagged to allow the creation and analysis of group membership; but also how the texts themselves can be revised by standardising spelling or substituting words where the meaning has changed. However, the message of the collection is that computerised analysis should not be thought of as replacing human researchers but as ways of augmenting and extending their work. The book aims to make the benefits to be gained by numerical techniques in the humanities known to a wider audience and to suggest where they might be applied next.

The chapters all focus on the use of keywords and word count, but each researcher has his, or her own, distinct, approach. One consideration that researchers all have to address is the issue of what to count in a text, that is, what to define as a word. This leads to a discussion of the theory of word distribution and its use and abuse in identifying authorship. Computerised analysis has allowed the use of many more keywords than can be analysed by hand. Comparisons of up to 6,000 word frequencies rather than the 100 most common are allowing researchers to investigate the differences between characters in Shakespeare's plays or between different genres of poetry. Computer analysis is shown to be able to identify bad reference corpus, so as to be able to determine optimum corpus size and the problems of distinguishing between differences due to different genres and different authors. Not all of the work is in classical texts. One chapter applies the tools to analysing the “moral panic” writing of the reformer Mary Whitehouse and testing how well these fit into previously defined sociological theories. Another compares the texts of pro‐ and anti‐fox hunting advocacy.

Each chapter is complete in itself, and together the chapters provide a fascinating insight into the sometimes quirky and arcane world of text analysis. The result is an easy to read, engaging and always interesting book.

Data & Figures

Supplements

References

Languages

or Create an Account

Close Modal
Close Modal