Article navigation

David Yakimischak

Introduction

In the late summer of 2002, JSTOR experienced unusual and unprecedented high usage from a location on a participating campus. Over the course of about a month, over 10,000 articles had been downloaded. As is our normal process, we contacted the institution, but upon investigation, it could not find a valid reason for this level of usage. We then saw similar levels of high usage from other participating institutions. When we stitched together the separate events,we found that different issues of the same journal titles were being downloaded at the various campuses, resulting, in aggregate, in the downloading of a nearly contiguous run of a journal backfile. We began to suspect that the usage was related and that someone (or some group) was systematically trying to download the full runs of several journals, primarily from the fields of sociology and economics.

What happened

One of the first institutions we contacted responded and informed us that the computer in question was not a user workstation but instead had inadvertently been set up as a proxy server without any authentication controls. To offer some background, a proxy server is a computer that stands as an intermediary between an end user and a Web site. Ordinarily, when a user accesses JSTOR without a proxy server, the IP address that our servers see is the IP address of that (the origin) machine. However, if a user trying to access JSTOR configures his or her browser to use a proxy server, then his or her requests are not sent to us directly, but rather, they are first sent to the proxy server, and then the proxy server sends the request to JSTOR. The IP address that we see for a user using a proxy server is not the user's own IP address but the IP address of the proxy server. Since JSTOR primarily uses IP Authentication (see description below), we permit or deny access based on IP address. Using a proxy server does not present problems as long as there are authentication controls, which limit the use of the proxy server to authorized individuals within that institution. However, it is easy to configure a computer to act as a proxy server and to overlook putting these access controls in place. This is what was happening at these institutions with the unusually high usage. Anyone, anywhere, on the Internet who was aware of these proxy servers could connect through them and gain unauthorized access to JSTOR. From the JSTOR perspective, the usage would appear to have come from a legitimate computer within the institution.

The institution in question was very helpful and shared with us the logs from the proxy server that had been used to access JSTOR. When we closely analyzed the logs, we discovered that there were a large number of accesses from an IP address that belonged to a non-participating institution. In other words, it appeared that someone had hijacked the proxy server, without knowledge or approval of the participating institution, and used it to gain access to JSTOR.

As the rate of downloading intensified, we became even more convinced that this was an orchestrated effort to download a substantial portion of the content in the JSTOR archive. We also began to recognize a pattern being used by the downloader. As we began to take steps to block this access, the pattern changed,and thus began a cat-and-mouse game in which we blocked the access path in advance of a suspected attempt. Although we were eventually able to stop the highest volumes of downloading, the net result was that over 51,000 articles were taken before we were able to stem the flow of unauthorized use.

Action taken to date

As we began to appreciate the gravity of the situation, our staff looked outside of JSTOR for advice and suggestions, speaking with librarians,publishers, network managers, experts, and trusted colleagues. We found that within the scholarly community as a whole there was very little awareness of the risks that exist as a result of open proxies. Those who were aware of the potential for misuse arising in connection with open proxies generally had not previously seen them used for such an organized attack on licensed resources. As we continued to research this topic, we found Web sites with listings of open proxy servers and even detailed technical instructions on how the uninitiated could use open proxies to gain access to licensed resources available at third-party institutions, resources which such individuals ordinarily would not be able to use.

This particular open proxy exploit takes advantage of the relative insecurity of IP address authentication. IP authentication is a scheme that allows or denies access to a resource based on the IP (numerical) address of the machine on the Internet. Although IP authentication provides ease of access to licensed resources, it is fundamentally not secure. The corollary to IP addresses is street addresses. Street addresses do not inherently provide security. We still need locks on our doors. Most of the open proxy servers that have been discovered are not the centralized, managed proxies that many institutions employ. They are set up inadvertently by individuals or departments, often with no awareness of the open path they provide.

In reaction to this experience, we began two parallel tracks of action within JSTOR. The first track was to pursue those individuals engaged in this active downloading. We are taking steps to ensure that those involved in the illicit downloading of content in the JSTOR archive are aware that it is not allowed and do not distribute this content further.

The second track we are following began with our decision to "go public"with this situation. We feel it is important to inform the community about the risks of open proxies and to begin a community-wide discussion of possible solutions. From the beginning of this situation, we realized that this is an issue that affects libraries, scholars, and other resource providers in our community. We believe, as a result of our experiences, that we are able to inform others and help develop solutions.

In December 2002, we sent a message to our library and publisher participants, which was then carried on a widely read listserv (http://www.library.yale.edu/~llicense/ListArchives/0212/msg00022.html). We have presented the discovery of this problem at several conferences and have had discussions with various groups in regards to potential solutions.

Responses and next steps

We have received many communications from participating institutions and have entered into a broad and useful dialog with many people on these issues. Institutional responses have been serious and engaged. Almost everyone we speak with is concerned about the situation and wants to do what he or she can to control the exposure and improve the ability to address this issue in the future.

The range of immediate responses is wide. Some institutions are evaluating whether there are policies in place to deal with this situation or if policies need to be updated or changed. Several are taking steps to educate themselves and their communities about the exposure, risks, and actions that can be taken. Others are taking the very direct action of scanning their own networks for open proxies, and shutting them down when they are found. Still others are actively looking into new technologies such as digital certificates (http://www.diglib.org/architectures/digcert.htm)and Shibboleth (http://shibboleth.internet2.edu/)as more permanent solutions.

Recent feedback from the technical community includes the suggestion that JSTOR take a lead in providing information about tools that would allow an institution to route access through secure, managed points of access on campus. Another suggestion has been that, prior to providing access, JSTOR should ensure that any machine attempting to access JSTOR is tested to determine whether or not it is an open proxy. These concepts have the potential to offer short-term improvements to the current situation, and they need to be discussed and debated in the community as we consider them.

Conclusions

This experience has reaffirmed the importance of the collaboration needed among different constituents as technologies connect us more closely together. This problem cannot be addressed by any one organization or even groups of organizations, and we are still learning about its implications (for example,what impact would widespread abuse of electronic resources have on usage statistics and resulting licensing strategies and decisions?). In this case,JSTOR has needed to work closely with librarians and networking and technology staff both to diagnose the problems and to address them. The positive spirit that has characterized this effort has encouraged us at JSTOR. We are confident that we can continue to make steady progress to ensure that the relationships of trust upon which the productive distribution of information resources is based can be maintained and even enhanced as we go forward.

Note# JSTOR and David Yakimischak.

David Yakimischak (davidyak@jstor.org) is the Chief Technology Officer at JSTOR, New York, USA.

or Create an Account

Close subscription notice
Close access options