Perseus Digital Library

Gregory R. Crane, Editor-in-Chief

Tufts University

Research

The Mission of Perseus

Our larger mission is to make the full record of humanity - linguistic sources, physical artifacts, historical spaces - as intellectually accessible as possible to every human being, regardless of linguistic or cultural background. Of course, such a mission can never be fully realized any more than we can reach the stars by which we guide the twisting paths and blind allies though the world around us. Similar instincts motivated scholars at the library at Alexandria in the 3rd century BCE, the Arab translators of Greek at Baghdad in the 9th century CE, the entrepreneurial printers of Greek and Latin in 15th century Italy, and the 19th century German scholars who built the infrastructure on which 20th century scholarship depended. None of these groups of scholars realized, of course, the fullest vision of universal knowledge that moved them. But that idealized vision allowed each to change the worlds in which they lived and carry humanity a little farther. We do not know what form such fundamental instruments as editions, lexica, encyclopedias, atlases, diagrams, museum catalogues, and archaeological site reports will assume, but we know that the infrastructure that we design now will materially enable or constrict how the next generation will be able to read languages from the past, scrutinize ancient artifacts, and explore the historical spaces.

Perseus has a particular focus upon the Greco-Roman world and classical Greek and Latin, but the larger mission provides the distant, but fixed star by which we have charted our path for over two decades. Early modern English, the American Civil War, the History and Topography of London, the History of Mechanics, automatic identification and glossing of technical language in scientific documents, customized reading support for Arabic language, and other projects that we have undertaken allow us to maintain a broader focus and to demonstrate the commonalities between Classics and other disciplines in the humanities and beyond. At a deeper level, collaborations with colleagues outside of classical studies make good on the claim that a classical education generally provides those critical skills and that intellectual adaptability that we claim to instill in our students. We offer the combination of classical and non-classical projects that we pursue as one answer to those who worry that a classical education will leave them or their children with narrow, idiosyncratic skills.

Within this larger mission, we focus on three categories of access:

Human readable information: digitized images of objects, places, inscriptions, and printed pages, geographic information, and other digital representations of objects and spaces. This layer of functionality allows us to call up information relevant to a longitude and latitude coordinate or a library call number. In this stage digital representations provide direct access to the physical senses of actual people in particular places and times. In some cases (such as high resolution, multi-spectral imaging), digital sources already provide better physical access than has ever been feasible when human beings had direct contact with the physical artifact.

Machine actionable knowledge: catalogue records, encyclopedia articles, lexicon entries, and other structured information sources. Physical access can serve our senses but provides no information about what we are encountering - in effect, physical access is like visiting a historical site about which we may know nothing and where any visible documentation is in a language that we cannot understand. Machine actionable knowledge allows us to retrieve information about what we are viewing. Thus, if we encounter a page from a Greek manuscript of Homer, we could at this stage find cleanly printed modern editions of the Greek, modern language translations, commentaries and other background information about the passage on that manuscript page. If we moved through a virtual Acropolis, we could retrieve background information about the buildings and the sculpture.

Machine generated knowledge: By analyzing existing information automated systems can produce new knowledge. Machine actionable knowledge allows, for example, us to look up a dictionary entry (e.g., facio, "to do, make") in a dictionary or to find pre-existing translations for a passage in Latin or Greek. Machine generated knowledge allows a machine to recognize that fecisset is a pluperfect subjunctive form of facio and to provide reading support where there is no pre-existing human translation. Such reading support might include full machine translation but also finer grained services such as word and phrase translation (e.g., recognizing whether orationes in a given context more likely corresponds to English "speeches," "prayers" or some other term), syntactic analysis (e.g., recognizing that orationes in a given passage is the object of a given verb), named entity identification (e.g., identifying Antonium in a given passage as a personal name and then as a reference to Antonius the triumvir).

Background

When we began work on Perseus in 1985, support from the Annenberg/CPB Project allowed us to create a critical mass of information - textual, archaeological, and artistic - about the ancient Greek world. As the Greek collections in Perseus matured, we were able not only to include Roman civilization but to explore other areas in the humanities, such as the history of science and early modern English, for which our collections and infrastructure were useful.

This broader research agenda led in 1998 to a major grant from the Digital Library Initiative Phase 2, funded primarily by the National Endowment for the Humanities (NEH) and the National Science Foundation (NSF), which funded us to study the problems of creating a digital library for the humanities as a whole. With this as our foundational support, we were able to produce collections on topics such as the History and Topography of London, the American Civil War, and early modern English, and services such as historical named entity identification and a digital library environment that anticipated services offered by giants such as Yahoo and Google and whose full functionality still exceeds any system with which we are familiar, and a stream of publications on the methods involved.

We had already begun building what the NSF would in the early years of this century call Cyberinfrastructure: an aggregate of collections and services, automatically linked and analyzed, that begins to show emergent properties qualitatively distinct from the print world. After having identified a number of practices that distinguish digital from print infrastructure, we had begun to define the services and collections that this new digital infrastructure would require. We then began addressing the problem of extracting the sophisticated knowledge needed for these new services from very large collections with hundreds of thousands and millions of books, available only as scanned page images. Planning grants from the Mellon Foundation allowed us to run a series of seminars and preliminary research on the general problem of "What do you do with a million books?" and the more specific topic of "Classics in the Million Book Library."

Our research on a range of topics outside of Classics allowed us to see how the problems of classical studies related to those of other disciplines. In exploring the challenges and opportunities of very large collections, we looked at the problems of collections from the early modern period (which present the greatest challenges for automatic processing), from the 19th century (which provide a best case, with many documents in easily analyzed print and a wealth of detailed information about people, places, organizations and other topics in a modern format), and classical studies (with the complex layouts of its critical editions, lexica, and other reference works, and its need to manage materials in not only Greek and Latin but English, French, German, and Italian, and, if we wanted to cover the full classical tradition, Arabic as well). Since classical editions were among the first printed books and classical scholarship has not only continued ever since but also particularly flourished in the 19th century, we realized that classical studies raised a superset of the challenges we had set out to study. When we considered as well that classical studies covered not only literature and history but ancient science and medicine, and art and archaeology, we realized that classics covered a superset of the problems that many other fields within the humanities faced. Having worked on Shakespeare and Early Modern English, 19th century newspapers and the American civil war, the city of London and the history of early modern science, we decided that classical studies provided the best space within which to advance our work on a cyberinfrastructure for the humanities in general.