Workshop on Electronic Texts: Proceedings, 9-10 June 1992 — Story, Setting & Ideas
Edition facts
Workshop on Electronic Texts: Proceedings, 9-10 June 1992 — Story, Setting & Ideas can be approached with a clearer sense of reading commitment from its source measurements: 65,549 words, 4 hr 45 min estimated reading time, and 7 detected text sections.
The text analysis averages about 25.5 words per sentence, while the detected sections provide another way to judge how the source is divided.
Project Gutenberg metadata also associates the work with “Text processing (Computer science) -- Congresses,” connecting these edition facts with the source record’s subject description.
Calculated from edition completeness, EPUB availability, text structure and catalogue metadata. Not a user rating.
How this score is calculated
- Description quality20 pts
- Title & short description10 pts
- Source metadata20 pts
- Text length15 pts
- Chapters / structure15 pts
- EPUB file integrity20 pts
Total of 100 points, scaled to a 2.5-5.0 range. Editions with an empty description or a missing EPUB file are not scored.
Read the Text
Memory had decided from the start to contract out conversion to external service bureaus. The criteria used to select these contractors were cost and quality of results, as opposed to methods of conversion. ERWAY noted that historical documents or books often do not lend themselves to OCR. Bound materials represent a special problem. In her experience, quality control--inspecting incoming materials, counting errors in samples--posed the most time-consuming aspect of contracting out conversion. ERWAY reckoned American Memory's costs at $4 per page, but cautioned that fewer cost-elements had been included than in NAL's figure.
OPTIONS FOR DISSEMINATION
The topic of dissemination proper emerged at various points during the Workshop. At the session devoted to national and international computer networks, LYNCH, Howard BESSER, Ronald LARSEN, and Edwin BROWNRIGG highlighted the virtues of Internet today and of the network that will evolve from Internet. Listeners could discern in these narratives a vision of an information democracy in which millions of citizens freely find and use what they need. LYNCH noted that a lack of standards inhibits disseminating multimedia on the network, a topic also discussed by BESSER. LARSEN addressed the issues of network scalability and modularity and commented upon the difficulty of anticipating the effects of growth in orders of magnitude. BROWNRIGG talked about the ability of packet radio to provide certain links in a network without the need for wiring. However, the presenters also called attention to the shortcomings and incongruities of present-day computer networks. For example: 1) Network use is growing dramatically, but much network traffic consists of personal communication (E-mail). 2) Large bodies of information are available, but a user's ability to search across their entirety is limited. 3) There are significant resources for science and technology, but few network sources provide content in the humanities. 4) Machine-readable texts are commonplace, but the capability of the system to deal with images (let alone other media formats) lags behind. A glimpse of a multimedia future for networks, however, was provided by Maria LEBRON in her overview of the Online Journal of Current Clinical Trials (OJCCT), and the process of scholarly publishing on-line.
The contrasting form of the CD-ROM disk was never systematically analyzed, but attendees could glean an impression from several of the show-and-tell presentations. The Perseus and American Memory examples demonstrated recently published disks, while the descriptions of the IBYCUS version of the Papers of George Washington and Chadwyck-Healey's Patrologia Latina Database (PLD) told of disks to come. According to Eric CALALUCA, PLD's principal focus has been on converting Jacques-Paul Migne's definitive collection of Latin texts to machine-readable form. Although everyone could share the network advocates' enthusiasm for an on-line future, the possibility of rolling up one's sleeves for a session with a CD-ROM containing both textual materials and a powerful retrieval engine made the disk seem an appealing vessel indeed. The overall discussion suggested that the transition from CD-ROM to on-line networked access may prove far slower and more difficult than has been anticipated.
WHO ARE THE USERS AND WHAT DO THEY DO?
Although concerned with the technicalities of production, the Workshop never lost sight of the purposes and uses of electronic versions of textual materials. As noted above, those interested in imaging discussed the problematical matter of digital preservation, while the TEI proponents described how machine-readable texts can be used in research. This latter topic received thorough treatment in the paper read by Avra MICHELSON. She placed the phenomenon of electronic texts within the context of broader trends in information technology and scholarly communication.
Among other things, MICHELSON described on-line conferences that represent a vigorous and important intellectual forum for certain disciplines. Internet now carries more than 700 conferences, with about 80 percent of these devoted to topics in the social sciences and the humanities. Other scholars use on-line networks for "distance learning." Meanwhile, there has been a tremendous growth in end-user computing; professors today are less likely than their predecessors to ask the campus computer center to process their data. Electronic texts are one key to these sophisticated applications, MICHELSON reported, and more and more scholars in the humanities now work in an on-line environment. Toward the end of the Workshop, Michael LESK presented a corollary to MICHELSON's talk, reporting the results of an experiment that compared the work of one group of chemistry students using traditional printed texts and two groups using electronic sources. The experiment demonstrated that in the event one does not know what to read, one needs the electronic systems; the electronic systems hold no advantage at the moment if one knows what to read, but neither do they impose a penalty.
The Workshop on Electronic Texts, held at the Library of Congress in June 1992, convened a diverse group—librarians, technologists, publishers, and scholars—to compare methods for placing historical textual materials in computerized form. The proceedings, edited by James Daly, preserve the presentations and discussions across seven sessions, from user needs to image capture standards. The opening remarks by Prosser Gifford and Carl Fleischhauer set a cautious tone: the assembly did not form a new nation, as the introduction puts it, but attendees gained much in insight. This record is valuable for its unvarnished look at the technical and policy debates that shaped early digitization.
Session I: Who Will Use Electronic Texts?
The first session, moderated by James Daly, grapples with user expectations. Avra Michelson’s overview, Susan Veccia’s user evaluation, and Joanne Freeman’s talk on scholars’ needs reveal a central tension: digitization promises broad access, but actual use remains uncertain. Veccia’s evaluation likely drew on early studies of how researchers interact with electronic sources, though the excerpt does not detail her findings. Freeman’s focus on scholars suggests that the workshop prioritized academic audiences over general readers. The discussion that followed—unfortunately truncated in the excerpt—hints at disagreements about whether digitization should serve specialized research or public education. This session frames the entire proceedings: every technical decision later debated (resolution, OCR accuracy, distribution) is ultimately about who the audience is and what they will do with the texts.
Image Capture: The Cost of Legibility
Session IV, on image capture and storage formats, offers the most concrete technical detail. George Thoma of the National Library of Medicine illustrates the trade-offs in scanning: higher resolution (600 or 1200 dpi) improves legibility but increases equipment, time, media, and transmission costs. He demonstrates techniques like fixed thresholding for black-and-white text with uniform contrast, and dynamic thresholding for pages with variable lighting or stains. A striking example shows a page with a burn mark, coffee stains, and yellow marker—corrected not by a single method but by adaptive algorithms. The discussion reveals that image quality is not absolute; it depends on the document’s physical condition and the intended use. For librarians, this session is a primer on the practical compromises that digitization projects faced in 1992, many of which remain relevant.
Show and Tell: Early Digital Projects
Session II, moderated by Jacqueline Hess, functions as a showcase of early digital initiatives. Presenters include Elli Mylonas on the Perseus Project (a digital collection of classical texts), Eric Calaluca on the Patrologia Latina Database, Carl Fleischhauer and Ricky Erway on the American Memory project, Dorothy Twohig on the Papers of George Washington, Maria Lebron on the Online Journal of Current Clinical Trials, and Lynne Personius on Cornell mathematics books. Each project represents a different approach to digitization—from scholarly editions to journal publishing to archival collections. The discussion that followed each presentation likely compared methods and standards, though the excerpt does not capture those exchanges. This session demonstrates the diversity of early electronic text projects and the lack of consensus on best practices.
Copyright and the Future of Access
Session VI, a single presentation by Marybeth Peters on copyright issues, addresses the legal framework that would govern electronic texts. The excerpt does not include her remarks, but the topic is critical: copyright determines what can be digitized, how it can be distributed, and who controls access. The workshop’s emphasis on historical and public-domain materials (e.g., American Memory, the Papers of George Washington) suggests that copyright was a limiting factor. The concluding session, moderated by Prosser Gifford, includes a general discussion that likely revisited these legal and policy questions. The proceedings as a whole show that technical solutions cannot be separated from legal and institutional constraints—a lesson that remains central to digital libraries today.
These proceedings capture a moment when digitization was still experimental, and participants were openly debating fundamental choices: resolution standards, OCR versus rekeying, distribution networks, and copyright. Readers interested in the history of digital libraries will find a primary source that reveals the assumptions and disagreements of early practitioners. The record is incomplete—discussions are summarized, not transcribed—but its value lies in the raw, unresolved questions it preserves.
I remember standing in a library basement, holding a microfiche reader like a fragile time machine. These 1992 proceedings—the arguments over OCR, the fear of losing the artifact—reminded me of that strange guilt before touching old paper. I once found the same feeling in Makers of Many Things — A Reader’s Guide: hands making objects, objects making lives, all held together by a quiet attention that digitization still hasn’t quite replaced.
There are no reviews for this eBook.
How did this book work for you?
Your answers remain private and are stored only in this browser.