Article navigation

Fedora Version 2.0

Released

The Fedora Project (www.fedora.info) has announced the release of version 2.0 of the Fedora open-source digital repository software. This release represents a significant increase in features and functionality over previous releases. New features include the ability to represent and query relationships among digital objects, a simple XML encoding for Fedora digital objects, enhanced ingest and export interfaces for interoperability with other repository systems, enhanced administrative features, and improved documentation. More than ever, Fedora is capable of serving as the foundation for many types of information management applications,including institutional repositories, digital libraries, records management systems, archives, and educational software.

As with prior versions of the software, all Fedora functionality is exposed through web service interfaces. At the core of this functionality is the Fedora object model that enables the aggregation of multiple content items into digital objects. This allows objects to have several accessible representations. For example, a digital object can represent an electronic document in multiple formats, a digital image with its descriptive metadata, or a complex science publication containing text, data, and video. Services can be associated with digital objects, allowing dynamically-produced views, or virtual representations of the objects. Historical views of digital objects are preserved through a powerful content versioning system.The new Fedora 2.0 introduces the Resource Index which is a module that allows a Fedora repository to be viewed as a graph of inter-related objects. Using the Resource Description Framework (RDF),relationships among objects can be declared, and queries against these relationships are supported by an RDF-based triple store. Fedora 2.0 also introduces Fedora Object XML (FOXML) which is a simple XML format for encoding Fedora digital objects. To support multiple XML standards, Fedoras ingest/export interface has been enhanced, permitting digital objects to be encoded in different formats. Currently, there is support for METS and FOXML. In future releases other XML formats will be supported, including MPEG21-DIDL. Other new features include a mass-update utility for modifying objects, a new administrative reporting interface, improved documentation, and tutorials.The Fedora open-source software is jointly developed by Cornell University and the University of Virginia with generous funding from the Andrew W. Mellon Foundation. Fedora 2.0 marks the final milestone in Phase I, a three year project to develop the core Fedora Repository system. Now underway, Fedora Phase II is a three year development project that will focus on advanced features including workflow, digital preservation, policy enforcement, information networks, and federated repositories.

Fedora Project: www.fedora.info

OCKHAM Initiative

Releases Version 0.5.3 of Harvest-to-Query (H2Q) Software

The OCKHAM initiative has released version 0.5.3 of its Harvest-to-Query(H2Q) software. This is the first widely-publicized release of H2Q. H2Q is an end-to-end solution for providing standard querying capabilities (such as Z39.50) for OAI-PMH available metadata collections.

Currently, H2Q has the following major features:

  • easy to install;

  • able to harvest metadata from any OAI-PMH available collection which provides its records in Dublin Core; and

  • provides Z39.50 querying to harvested collections.

Once H2Q achieves its 1.0 status, it will have the following major features;it will:

  • be as easy to install as your toaster;

  • be able to harvest metadata from any OAI-PMH available collection;

  • provide Z39.50 and SRU/W querying to harvested collections; and

  • allow harvesting and indexing of any XML-based metadata scheme.

The OCKHAM H2Q software can be downloaded from www.ockham.org/services.php The OCKHAM Initiative seeks to promote the development of digital libraries via collaboration between librarians and digital library researchers. By promoting simple, open approaches and standards for digital library tools, services, and content, the gap between digital library development and the adoption of digital library systems by the traditional library community will hopefully be bridged. The initiative is sponsored by the Digital Library Federation (www.diglib.org).

OCKHAM Initiative: www.ockham.org

IMIRSEL, GSLIS and UIUC

M2K (Music-to-Knowledge) Toolkit Released

The International Music Information Retrieval Systems Evaluation Laboratory(IMIRSEL) at the Graduate School of Library and Information Science (GSLIS),University of Illinois at Urbana-Champaign (UIUC), has announced the official release of the M2K (Music-to-Knowledge) Alpha 1.0 toolkit.

M2K is an open-sourced Java-based framework designed to allow Music Information Retrieval (MIR) and Music Digital Library (MDL) researchers to rapidly prototype, share and scientifically evaluate their sophisticated MIR and MDL techniques. M2K builds on and extends the D2K/T2K data mining framework developed by the Automated Learning Group (ALG) at the National Center for Supercomputing Applications (NCSA).

M2K Alpha 1.0 is currently strongest in audio-based approaches to MIR/MDL tasks. All MIR/MDL researchers with interests in symbol-based and metadata-based techniques are encouraged to join in and help extend the functionality of M2K.

M2K Alpha 1.0 downloads, installation requirements and instructions, and a wide range of documentation, can be found via: http://music-ir.org/evaluation/m2k

Brainboost

Unveils Improved Answer Engine

Brainboost has unveiled a new version of its popular www.brainboost.com site, powered by an improved Brainboost Answer Engine utilizing Brainboost's patent pending AnswerRank™ Technology. For the task of answering questions posed in plain English, Brainboost renders other search engines obsolete by leapfrogging what can be achieved with natural language processing technologies.

Brainboost can answer questions posed in plain English as opposed to simply directing users to pages that mention the question, like existing engines do. Brainboost represents a fundamentally different way to find knowledge on the web. Whereas search engines force users to spend time scouring for an answer by tediously reading through search results, Brainboost does this work for the user by automatically reading hundreds of search results and extracting answers to the user's question.

While search engines like Ask Jeeves and MSN have recently made limited attempts at answering questions from finite structured databases such as encyclopedias or pre-canned editorially collected answers, Brainboost is the only commercial engine that leverages the full knowledge available on the web with its ability to understand even the most complex unstructured document. For Brainboost, any web page or document on the internet is a potential answer source.

A study conducted by Brainboost confirmed that while all major search engines barely reach a 50 percent success rate in providing the answer to a users question in the first page of search results, Brainboost clocks in at about 87 percent. When looking at how many of those actually had the answer as the first result, the gap is even wider at 23 percent for all search engines and 75 percent for Brainboost. Search Engines examined in this study include: Ask,Google, MSN and Yahoo!.

For the past year Brainboost, in stealth mode, has been operating the now popular BrainBoost.com search engine. Brainboost has become a popular destination for web savvy users wanting to get an answer to a question. The enthusiasm of millions of users has resulted in a highly accurate Answer Engine leveraging the millions of questions Brainboost has been asked over this time frame. With every day that goes by, the Brainboost Answer Engine becomes more intelligent, benefiting from the feedback gained from its users and the growing knowledge base it continues to amass.

www.brainboost.com/

OPAL

Offers Podcasts of Library Audio Programs

Online Programming for All Libraries (OPAL) has begun podcasting audio recordings of archived OPAL online events. Listeners can now hear OPAL events on a wide variety of portable MP3 players. Listeners also can link to an RSS feed to be notified whenever a new podcast becomes available.

OPAL is a collaborative effort by libraries of all types to provide cooperative web-based programming and training for library users and library staff members. These live, online events are held in an online auditorium where participants can interact via voice-over-IP, text chatting, and synchronized browsing. Examples of OPAL public online programs include book discussion programs, interviews, library training, memoir writing workshops, and virtual tours of special digital library collections.

Digital audio recordings of OPAL programs are placed in the OPAL Archive (www.opal-online.org/archive.htm)so that interested patrons who missed the live online event can listen at a convenient time.

In a related development, digital audio recordings of OPAL programs will become available in the popular MP3 format. Until now, audio recordings were available only in Windows Media Audio (WMA) format. Offering both formats will extend the reach and usability of OPAL programs.

OPAL utilizes software from Talking Communities (www.talkingcommunities.com/)featuring voice-over-IP, text chatting, and synchronized browsing. OPAL is administered by the Alliance Library System (www.alliancelibrarysystem.com/),the Mid-Illinois Talking Book Center (www.mitbc.org/),and the Illinois State Library Talking Book and Braille Service (www.cyberdriveillinois.com/departments/library/who_we_are/talking_book_and_braille_service/home.html).

To experience an OPAL podcast: http://feeds.feedburner.com/OpalPodcast

An RSS link has been added to the OPAL homepage: www.opal-online.org

NaCTeM

To Provide Text Mining Services to the Academic Community

The National Centre for Text Mining (NaCTeM) is a collaboration between the Universities of Manchester, Liverpool and Salford. Funding is provided by the Joint Information Systems Committee (JISC), the Biotechnology and Biological Research Council (BBSRC) and the Engineering and Physical Sciences Research Council (EPSRC).

Search engines return thousands of documents, but the difficulty for the user is to find those which are most personally relevant. Most of these searches have little concept of the meaning of words that is gained from the context of a sentence. By using natural language processing, text mining can discover this meaning and focus on specific needs of the user.

Detailed abstracts can then be compared and contrasted using data mining to discover patterns and associations that the human eye is more likely to miss. This has proved to be particularly useful in the fields of drug discovery and predictive toxicology.

Initially focusing on providing a service for the fields of biological and biomedical science, the Centre will also serve the broader needs of the academic community through the provision of text mining tools, advice and ongoing research.

NaCTeM will be run by an international consortium, including Manchester,Liverpool and Salford Universities. These core partners are extended by international partners: the University of California Berkeley, the University of Geneva, the San Diego Supercomputer Centre, and the University of Tokyo, with the European Bioinformatics Institute having presence on the Technical Directorate.

The first publicly funded text mining centre in the world, NaCTeM will be housed in the Manchester Interdisciplinary Biocentre at The University of Manchester, and will contribute to the associated national and international research agenda by establishing a service for the wider academic community. Strong contacts will be forged by the Centre with business and government sectors to achieve long term sustainability for the service.

http://www.nactem.ac.uk/

SUNCAT

Pilot Service Launched

The National Serials Union Catalogue (SUNCAT) for the UK research community was launched as a pilot service in February 2005. The contents of journals and other serials represent an immense and invaluable resource for researchers in all subjects. However, the task of identifying, locating and accessing these serials, held by institutions across the UK, has up to now presented a significant challenge. Meeting this challenge is the goal of the SUNCAT project.

Funded by Joint Information Systems Committee (JISC) and the Research Support Libraries Programme (RSLP) since 2003, and developed by EDINA at the University of Edinburgh in partnership with Ex-Libris, the catalogue has achieved a critical mass of some 3.7 million records. These are made up of records from national libraries, the largest UK academic library collections and international databases, such as the ISSN World Serials database and CONSER, the database of MARC21 serials records available from the Library of Congress.

As a centralized catalogue of high-quality bibliographic records, SUNCAT will also provide librarians with a means by which local records can be upgraded through access to standardized high-quality records.

The launch also marks the start of phase 2 of the programme, also funded by JISC. Representing an investment of £1 million, this phase will see coverage of the catalogue extend to up to 60 new libraries across the UK.

SUNCAT will also be developed to integrate fully into the emerging national information environment, supporting and contributing to work in the areas of e-theses, repositories, electronic subscription information, online access to journals and electronic document delivery, amongst others. As such, SUNCAT will further develop as a key tool for researchers, librarians and others within colleges, universities and beyond.

To access the pilot SUNCAT service and for further information: www.edina.ac.uk/suncat

Automated Personalized Assistance System

Improves Web Search Success

A Penn State researcher has developed software that improves web searching with a personalized system that offers automated assistance for structuring and refining queries, evaluating search results and finding more relevant information.

"Research shows 50 percent of all web results retrieved are not relevant,pointing to a need for improved searching techniques," said Jim Jansen,assistant professor of information sciences and technology. "This technology enabled a 20 percent performance increase."

The technology, designed to be integrated with a browser, monitors what searchers are looking for based on user-system interactions and then interjects help in finding needed information.

Other approaches to personalizing searches rely on "explicit feedback"where the system interrupts searchers as they hunt for information. Research has shown that only 1 percent to 2 percent of users are likely to use such systems because of the extra effort involved, Jansen said.

His technology uses "implicit feedback" as revealed through searchers' query patterns, so it does not place a burden on the user. Furthermore, because the application occurs on the client side and not the server side, it leverages the downtime during searches to complete its computations.

The technology is outlined in a paper, "Seeking and Implementing Automated Assistance During the Search Process" (available online at www.sciencedirect.com) that will appear in print in the Journal of Information Processing and Management's, July 2005 issue.

Jansen evaluated the automated personalized assistance system with 30 college students who were instructed to search for five minutes on one of two chosen topics. The students were told their web-search system had searching advice that could be accessed by clicking an assistance button on the browser. The searching assistance also could be ignored. Users were receptive to the assistance, taking advantage of it 54 percent of the time. All viewed it at least once and 27 of the 30 users implemented the assistance, suggesting that the help was of value.

The test results also showed at what point in the search process users could benefit from even more personalized assistance. This occurred after viewing the initial results and after viewing a relevant document.

www.psu.edu/ur/2005/automatedsearch.html

Nielsen//NetRatings

Online Searchers Using More than One Search Engine

Nielsen//NetRatings has reported that a minority of searchers exclusively use only one of the top three search engines – Google Search, Yahoo! Search and MSN Search. According to the latest custom research from Nielsen//NetRatings MegaView Search, 58 percent of Google searchers also visited at least one of the other top two search engines, MSN Search and Yahoo! Search, showing that even though Google's market share is dominant today, there is significant opportunity for its competitors to grow their share. The use of multiple search engines is not limited to Google's searchers. Nearly 71 percent of those who searched at Yahoo! also visited at least one of the other top two search engines, and 70 percent of those who searched at MSN also tried their luck at one or both of the other two.

Drilling down into the user overlap at the top three search engines,Nielsen//NetRatings found that Google users showed a higher degree of loyalty than its competitors. While Google shared 58 percent of its visitors: 26 percent with Yahoo!, 19 percent with MSN, and 14 percent with both Yahoo! and MSN, its competitors' overlap of searchers was higher. Yahoo! shared 71 percent of its traffic: 39 percent with Google, 11 percent with MSN, and 21 percent with both Google and MSN. MSN shared 70 percent of its traffic: 33 percent with Google, 13 percent with Yahoo!, and 24 percent with both Google and Yahoo!

Full Press Release with tables: http://biz.yahoo.com/prnews/050228/sfm063_1.html

Pew Internet & American Life Project

Music and Video Downloading Report

About 36 million Americans – or 27 percent of internet users – say they download either music or video files and about half of them have found ways outside of traditional peer-to-peer networks or paid online services to gather and swap their files, according to the most recentsurvey of the Pew Internet& American Life Project titled "Music and Video Downloading Moves Beyond P2P". The Project's national survey of 1,421 adult internet users conducted between January 13 and February 9, 2005 showed that 19 percent of current music and video downloaders, about 7 million adults, say they have downloaded files from someone else's iPod or MP3 player. About 28 percent, or 10 million people,say they get music and video files via email and instant messages. There is some overlap between these two groups; 9 percent of downloaders say they have used both of these sources.

In all, 48 percent of current downloaders have used sources other than peer-to-peer networks or paid music and movie services to get music or video files. Beyond MP3 players, email and instant messaging, these alternative sources include music and movie web sites, blogs and online review sites. The survey of internet users has a margin of error of plus or minus 3 percent.

Other highlights in the new Pew Internet Project survey include:

  • A total of 49 percent of all Americans and 53 percent of internet users believe that the firms that own and operate file-sharing networks should be deemed responsible for the pirating of music and movie files. Some 18 percent of all Americans think individual file traders should be held responsible and 12 percent say both companies and individuals should shoulder responsibility. Almost one in five Americans (18 percent) say they do not know who should be held responsible or refused to answer the question.

  • The public is sharply divided on the question of whether government enforcement against music and movie pirates will work, but broadband users strongly believe that a government crackdown will not succeed. Some 38 per cent of all Americans believe that government efforts would reduce file-sharing and 42 percent believe that government enforcement would not work very well. Broadband users are more skeptical about government anti-piracy efforts. Some 57 percent of broadband users believe there is not much the government can do to reduce illegal file-sharing, compared to 32 percent who believe that enforcement would help control piracy.

Report: www.pewinternet.org/pdfs/PIP_Filesharing_March05.pdf

Pew Internet & American Life Project web site: www.pewinternet.org

CLIR

Library as Place: New CLIR Report Released

The Council on Library and Information Resources (CLIR) has released a new report entitled: "Library as Place: Rethinking Roles, Rethinking Space". CLIR commissioned six experts to explore the questions of the role of the library in a networked world and effect of the role change for the creation and design of library space, in a series of essays. The authors – an architect,four librarians, and a humanities professor – provide diverse visions of the library, its services, and its space in the twenty-first century. The essays suggest that changes in approaches to teaching and learning, combined with the possibilities offered by technology, present rich new opportunities for libraries and for library design.

Report: www.clir.org/pubs/abstract/pub129abst.html

DigiCULT

Third DigiCULT Technology Watch Report Available

The third of DigiCULT's Technology Watch Report is now available online. Topics covered include: Open Source Software and Standards; Natural Language Processing; Information Retrieval; Location-Based Systems; Visualization of Data; Telepresence, Haptics, Robotics. Digital Culture (DigiCULT) is an IST Support Measure (IST-2001-34898) to establish a regular technology watch for cultural and scientific heritage.

Report: www.digicult.info/downloads/TWR3-highres.pdf

DigiCULT web site: www.digicult.info/pages/index.php

DOAR

DOAR – Directory of Open Access Repositories

The University of Nottingham, UK and University of Lund, Sweden are developing a new service to support the rapidly emerging movement towards Open Access to research information. The new service, called DOAR – the Directory of Open Access Repositories – will categorize and list the wide variety of Open Access research archives that have grown up around the world. Such repositories have mushroomed over the last two years in response to calls by scholars and researchers worldwide to provide open access to research information. Leading UK universities are establishing a national network of archives, where research papers are available to download for free. Internationally, universities and research institutes in Australia, India,Europe and the USA are doing the same, so that there are now large numbers of archives of different sizes, composition and scope.

DOAR will provide a comprehensive and authoritative list of institutional and subject-based repositories, as well as archives set up by funding agencies– like the National Institutes for Health in the USA or the Wellcome Trust in the UK and Europe. Users of the service will be able to analyze repositories by location, type, the material they hold and other measures. This will be of use both to users wishing to find original research papers and for third-party"service providers", like search engines or alert services, which need easy to use tools for developing tailored search services to suit specific user communities.

The project is a joint collaboration between the University of Nottingham in the UK and the University of Lund in Sweden. Both institutions have a track record in the Open Access area.

Lund operates the Directory of Open Access Journals (DOAJ), which is known throughout the world. Nottingham leads SHERPA, an institutional repository project that has helped establish Open Access archives in 20 of the leading UK research universities. Nottingham also runs the SHERPA/RoMEO database, which is used worldwide as a reference for publisher's copyright policies.

The importance and widespread support for the project can be seen in its funders, led by the international Open Society Institute (OSI), which is a major player in advocacy for the spread of open access to the world's research findings. The UK funding body Joint Information Systems Committee (JISC) has also backed the 18 month project, as part of a larger program of funding for repository development in UK institutions. There has been additional contributory funding from the Consortium of Research Libraries (CURL) and from SPARCEurope – an alliance of European research libraries, library organizations, and research institutions.

Much work has already been done in listing repositories in different countries and the project will build on this existing work. The project will last 18 months in its initial phase, during which time work will be done to survey and classify existing repositories, produce the database for use and to assist new repositories to register themselves to maximize their visibility and the use of their research content.

Project web site: www.opendoar.org/

Liberating Scholarly Literature

Open Access Bibliography Now Available

ARL and Charles W. Bailey of the University of Houston have announced the availability of a new bibliography of open access works. The Open Access Bibliography: Liberating Scholarly Literature with E-Prints and Open Access Journals presents over 1,300 selected English-language books, conference papers(including some digital video presentations), debates, editorials, e-prints,journal and magazine articles, news articles, technical reports, and other printed and electronic sources that are useful in understanding the open access movement's efforts to provide free access to and unfettered use of scholarly literature. Most sources have been published between 1999 and August 31, 2004;however, a limited number of key sources published prior to 1999 are also included. Where possible, links are provided to sources that are freely available on the Internet (approximately 78 percent of the bibliography's references have such links).

The bibliography is organized into the following categories: General Works,Open Access Statements, Copyright Arrangements for Self-Archiving and Use, Open Access Journals, E-Prints, Disciplinary Archives, Institutional Archives and Repositories, Open Archives Initiative and OAI-PMH, Conventional Publisher Perspectives, Government Inquiries and Legislation, and Open Access Arrangements for Developing Countries. The publication also includes a concise overview of key concepts that are central to the open access movement.This bibliography has been published as a printed book (ISBN 1-59407-670-7) by the Association of Research Libraries (ARL). ARL and the author have made the PDF version of the bibliography freely available. It is licensed under the Creative Commons Attribution-NonCommercial License.

Open Access Bibliography: http://info.lib.uh.edu/cwb/oab.pdf

ARL Ordering Information: www.arl.org/pubscat/pubs/openaccess/

or Create an Account

Close subscription notice
Close access options