This paper will discuss the integration of document image processing and text retrieval principles in order to process and load existing paper documents automatically in an electronic document database that broadens the user's capability to retrieve relevant information more accurately, without going through costly processes to get paper documents into electronic text. The principles of document image processing systems, as well as the problems and shortcomings of most of today's document image processing systems, will be discussed. Then concept retrieval as the latest development in text retrieval will be discussed, with specific reference to the ability of the TOPIC intelligent text retrieval system to allow users to build up a knowledge base of search objects or concepts that can be used at any point in time by all users for the system. This paper will further specifically look at the automatic processing of paper documents by converting the scanned document image pages through to electronic text. The use of optical character recognition technology, the indexing and loading of the documents in a text database, the automatic linking of the documents to the related document images and the retrieval technology available in TOPIC, specifically the TYPO operator that was developed to handle so‐called dirty data such as the common misspellings, character transpositions and ‘dirty’ text received as output from the OCR process, will be discussed. A possible solution to load paper documents quickly and cost‐effectively into an electronic document database will be discussed and demonstrated in detail. The advantages and disadvantages of this approach will be discussed with specific reference to an electronic news clipping service application.
Article navigation
1 April 1993
Review Article|
April 01 1993
The integration of document image processing and text retrieval principles
Niël van der Merwe
Niël van der Merwe
Xcel, PO Box 20355, Alkantrant 0005, South Africa
Search for other works by this author on:
Publisher: Emerald Publishing
Online ISSN: 1758-616X
Print ISSN: 0264-0473
© MCB UP Limited
1993
The Electronic Library (1993) 11 (4-5): 273–278.
Citation
van der Merwe N (1993), "The integration of document image processing and text retrieval principles". The Electronic Library, Vol. 11 No. 4-5 pp. 273–278, doi: https://doi.org/10.1108/eb045245
Download citation file:
New and popular articles
Suggested Reading
RSS feeds behavior analysis, structure and vocabulary
International Journal of Web Information Systems (August,2014)
A Typo-Morphological Study: The Cmc Industrial Mass Housing District, Lefke, Northern Cyprus
Open House International (June,2013)
Effect of typo-morphological analysis and place understanding on the nature of intervention within historic settings: the case of Amman, Jordan
Archnet-IJAR: International Journal of Architectural Research (February,2023)
Turbo Lightning: Spelling correction as you type
The Electronic Library (May,1986)
The morphology of urban tourism space (case: Malioboro Main Street as cosmological Axis of Yogyakarta city, Indonesia)
International Journal of Tourism Cities (June,2024)
Related Chapters
Victory Through Vulnerability
Escape the Cape, From Existing to Evolving: Amplifying Voices of Black and Brown Women in the Mental Health Profession
Verba Bestiae: How Latin Conquered Heavy Metal
Multilingual Metal Music: Sociocultural, Linguistic and Literary Perspectives on Heavy Metal Lyrics
In the Eye of the Beholder
Enhancing Writing Skills
Recommended for you
These recommendations are informed by your reading behaviors and indicated interests.
