XML is becoming one of the most important structures for data exchange on the web. Despite having many advantages, XML structure imposes several major obstacles to large document processing. Inconsistency between the linear nature of the current algorithms (e.g. for caching and prefetch) used in operating systems and databases, and the non‐linear structure of XML data makes XML processing more costly. In addition to verbosity (e.g. tag redundancy), interpreting (i.e. parsing) depthfirst (DF) structure of XML documents is a significant overhead to processing applications (e.g. query engines). Recent research on XML query processing has learned that sibling clustering can improve performance significantly. However, the existing clustering methods are not able to avoid parsing overhead as they are limited by larger document sizes. In this research, We have developed a better data organization for native XML databases, named sibling‐first (SF) format that improves query performance significantly. SF uses an embedded index for fast accessing to child nodes. It also compresses documents by eliminating extra information from the original DF format. The converted SF documents can be processed for XPath query purposes without being parsed. We have implemented the SF storage in virtual memory as well as a format on disk. Experimental results with real data have showed that significantly higher performance can be achieved when XPath queries are conducted on very large SF documents.
Article navigation
27 September 2007
Research Article|
September 27 2007
Sibling‐First Data Organization for Parse‐Free XML Data Processing
Hooman Homayounfar;
Hooman Homayounfar
Department of Computing and Information Science, University of Guelph, Guelph, Ontario, Canada email: hhomayou@uoguelph.ca
Search for other works by this author on:
Fangju Wang
Fangju Wang
Department of Computing and Information Science, University of Guelph, Guelph, Ontario, Canada email: hhomayou@uoguelph.ca
Search for other works by this author on:
Publisher: Emerald Publishing
Online ISSN: 1744-0092
Print ISSN: 1744-0084
© Emerald Group Publishing Limited
2006
International Journal of Web Information Systems (2007) 2 (3-4): 176–186.
Citation
Homayounfar H, Wang F (2007), "Sibling‐First Data Organization for Parse‐Free XML Data Processing". International Journal of Web Information Systems, Vol. 2 No. 3-4 pp. 176–186, doi: https://doi.org/10.1108/17440080780000298
Download citation file:
New and popular articles
Suggested Reading
Digital Libraries and the Challenges of Digital Humanities
Program (May,2007)
Using XML: A How‐to‐do‐it Manual and CD‐ROM for Librarians
Library Hi Tech (June,2008)
An XSketch-based spelling suggestion approach for XML keyword search
International Journal of Web Information Systems (August,2014)
A mediation layer for heterogeneous XML schemas
International Journal of Web Information Systems (February,2005)
The Index of Middle English Prose: Index to Volumes I-XX
Reference Reviews (September,2015)
Related Chapters
Unlocking the Potential of Data Lakes: Organizing and Storing Marketing Data for Analysis
Data Engineering for Data-driven Marketing
Assessing Progress Made by Indian States and UTs for the Attainment of Sustainable Development Goals
Modeling Economic Growth in Contemporary India
A Very, Very Particular Library. A Conceptual Kaleidoscope: Between the Creation Laboratory and the Research Labyrinth
Poetry as Knowledge in Librarianship: Inquiry, Identity, and Praxis
Recommended for you
These recommendations are informed by your reading behaviors and indicated interests.
