Parallel methods for the generation of partitioned inverted files

MacFarlane, A.; McCann, J.A.; Robertson, S.E.

doi:10.1108/00012530510621888

Article navigation

Research Article| October 01 2005

Parallel methods for the generation of partitioned inverted files

A. MacFarlane;

A. MacFarlane

Centre for Interactive Systems Research, City University, London, UK

Search for other works by this author on:

This Site

PubMed

Google Scholar

J.A. McCann;

J.A. McCann

Department of Computing, Imperial College, London, UK

Search for other works by this author on:

This Site

PubMed

Google Scholar

S.E. Robertson

Centre for Interactive Systems Research, City University, London, UK Microsoft Research Ltd, Cambridge, UK

Search for other works by this author on:

This Site

PubMed

Google Scholar

Author & Article Information

Publisher: Emerald Publishing

Online ISSN: 1758-3748

Print ISSN: 0001-253X

2005

Aslib Proceedings (2005) 57 (5): 434–459.

https://doi.org/10.1108/00012530510621888

Purpose

The generation of inverted indexes is one of the most computationally intensive activities for information retrieval systems: indexing large multi‐gigabyte text databases can take many hours or even days to complete. We examine the generation of partitioned inverted files in order to speed up the process of indexing. Two types of index partitions are investigated: TermId and DocId.

Design/methodology/approach

We use standard measures used in parallel computing such as speedup and efficiency to examine the computing results and also the space costs of our trial indexing experiments.

Findings

The results from runs on both partitioning methods are compared and contrasted, concluding that DocId is the more efficient method.

Practical implications

The practical implications are that the DocId partitioning method would in most circumstances be used for distributing inverted file data in a parallel computer, particularly if indexing speed is the primary consideration.

Originality/value

The paper is of value to database administrators who manage large‐scale text collections, and who need to use parallel computing to implement their text retrieval services.

2005

You do not currently have access to this content.

Don't already have an account? Register

Parallel methods for the generation of partitioned inverted files

Email Alerts

Cited By

Parallel methods for the generation of partitioned inverted files

Sign in

Client Account

ICE Member Sign In

Email Alerts

Suggested Reading

Related Chapters

Recommended for you

Cited By

Sharing Unavailable