The Newberry French Revolution Collection at ARTFL

As we begin planning Digitizing Enlightenment IV, which will take place in the context of the ISECS Congress in Edinburgh in July 2019, we are keen to broaden the scope and breadth of the Digitizing Enlightenment community in order to highlight new, and existing, digital projects across the interdisciplinary spectrum of eighteenth-century studies. This post, based on work presented at the Digitizing Enlightenment III workshop held in Oxford in July 2018, demonstrates how to identify text reuse – citations, borrowings, plagiarisms – as well as other techniques for leveraging freely available large data-sets from the 18C.
– Glenn Roe, Voltaire Lab

The incredible richness of the Newberry Library’s French Revolution Collection (FRC) has been long known. It consists of more than 30,000 pamphlets and more than 23,000 issues of 180 periodicals published between 1780 and 1810, representing the opinions of all the factions that opposed and defended the monarchy during the turbulent period between 1789-1799 and also contains innumerable ephemeral publications of the early First Republic. The Newberry has released digital copies of more than 35,000 pamphlets totalling approximately 850,000 pages. Not only has the Newberry made the collection available to the public, but it has released a data feed of the entire collection, consisting of the Library’s exceptional metadata describing each object, the OCR text data, and links to the digital facsimiles accessible from the Internet Archive, encouraging researchers and instructors to incorporate the digital collection in new kinds of scholarship and engagement.

In order to facilitate experimental work at ARTFL on this unparalleled resource, we have loaded two versions of this collection – based on a download of the collection from the Newberry’s GitHub repository in November 2017 – into PhiloLogic4, the latest release of ARTFL’s text analysis software. The full version contains all 38,377 documents dating from the 16th century to the end of the 19th century. Our second build attempts to eliminate duplicate documents, is restricted to the period 1787-1799, and thus contains 26,445 documents.   Additional implementation information and full open access to both versions of the FRC collection are available online. The quality and coverage of the FRC texts makes it an ideal environment to test a variety of experiments and algorithms to enhance access and open new kinds of approaches using the 1787-99 sample data. At the bottom of the ARTFL FRC page, we have provided links to several different models for examining the collection which are based on extensions to the PhiloLogic4 package.

The simplest model is a document level search which returns matching documents by relevancy ranking based on Python Whoosh. This functions somewhat like a Google search on the collection, with links to the page images of the document or specific instances of the search words in context. For example, the results of a search for “conspirateurs aristocrates ennemis étrangères royalistes” can be seen here.

The second approach is the application of a Topic Model algorithm to the collection. Topic Models are a set of unsupervised learning algorithms that divide collections into a specified number of clusters based on vocabularies of each document which is widely used in digital humanities. The results of the Topic Model has been added to the metadata of the PhiloLogic4 build of the 1787-99 sample data. Each document is identified as having a first and second topic, denoted as A or B, with a number from 00-49 as listed in this TABLE. This first column is the topic number, the second is one or more english keywords which can also be searched. The third column is the top 3 weighted words (features) of that topic, and the 4th column is the rest of the top 10, all of which are shown in relative weight order. Thus, A29 will return the documents that have money assignats as the top weighted topic. Searching for “money” in topic models will get this as eight the first or second topic.   An alternative use of this data is to copy some or all of the terms in columns 3 and 4 into the Whoosh search form and get the documents in a ranked relevancy order.

Our first presentation of our work at the Digitizing Enlightenment III showed results from applying the latest version of our sequence aligner to detect text reuse – citations, borrowings, plagiarisms, and so on – from pre-Revolutionary documents during the Revolutionary period. Sequence alignment is a family of algorithms used in a surprising range of disciplines from genetics to text analysis to identify similar segments of arbitrary length. For this work, we aligned the FRC 1787-99 sample against ARTFL’s Frantext pre-1788 collection. The Frantext sample contains 1,263 documents and is particularly strong in 18th century holdings. We loaded the results of the alignment run in a dedicated database which can be queried in a variety of ways, such as source and/or target metadata as well as by words in matching passages.

The public database (June 22, 2018 build) found 8,937 aligned passages, or which around 1,000 were identified algorithmically as banalities. Filtering out shorter alignments, less than 10 words, results in just under 7,000 passages. It is important to note that these numbers are very relative, since they can vary significantly depending on the approach we use to identify and merge, where appropriate, longer passages. The general frequencies are not particularly surprising. The following is a table of the number of borrowed passages in the FRC by author.

Montesquieu – 1,315

Rousseau – 1,133

Voltaire – 979

Mably – 303

Aulony – 263

Racine – 168

Helvétius – 167

D’Holbach* – 146

 

Saint-Simon – 135

Bossuet – 110

La Fontaine – 94

Diderot – 85

Corneille – 72

Mirabeau – 71

Boileau – 69

Bernardin – 67

Montaigne – 65

*D’Holbach appears as two entries due to slight metadata differences.

The yearly distribution of borrowings from the top three Enlightenment authors again follows a reasonable pattern.

The annual distribution in the FRC of the 536 passages derived from Rousseau’s Contrat Social, seems reasonable and would match expectations based on other things we know.

While the global numbers are interesting, if not very surprising, there are number of specific texts and authors which would warrant further investigation. There are numerous chapbooks, such as the Calendrier moral, 1794, which are interesting because of their selection of inspiring passages from various authors. Jean-Jacques Barthélemy’s L’Accord de la religion et de la liberté (1791) features some 25 long extracts from d’Holbach’s Système social.

The alignment database is available to the public. The database has a variety of useful features. This link will push a search for all of the aligned passages in the FRC from Rousseau’s Contrat Social greater than 10 words. The report is laid out chronologically (in this case by FRC year). Each instance shows the matching passages with available metadata, links to the context of each passage, and a button to highlight the differences in each matching pair. The facets on the right will allow you to get frequencies by author, title, year and so on. Clicking on those will return the corresponding text pairs.

We anticipate further experimental work on the FRC, most notably in using the excellent subject information as ways to assess the accuracy of Topic Modelling and to consider supervised learning algorithms to further classify the collection by subject.

It is our pleasure to acknowledge that the Newberry Library has released this extraordinary resource under the Open Data Commons Attribution License, ODC-BY 1.0.   We believe that this splendid collection and the Newberry’s release of all of the data will facilitate a generation of ground-breaking work in Revolutionary studies. If you find the collection useful, please do contact the Newberry Library to congratulate them on this wonderful initiative and how their efforts contribute to your research.

We would love to hear from you. Please send comments, suggestions and problem reports to artfl@artfl.uchicago.edu.

– Clovis Gladstone and Mark Olsen

 

Advertisements

Voltaire Foundation appoints Digital Research Fellow

I am delighted to announce my appointment as Digital Research Fellow at the Voltaire Foundation for the academic year 2017-2018. This is the first Digital Humanities appointment in French at Oxford, and is made possible by the generosity of M. Julien Sevaux and the John Fell Fund. As Digital Research Fellow, I will oversee the creation of a pilot Digital Voltaire project, establishing a dataset that for the first time contains all of Voltaire’s works, including his correspondence, as well as undertake a series of computational experiments around the theme of ‘Visualising Voltaire’.

Voltaire, by Maurice Quentin de La Tour, 1735.

Voltaire, by Maurice Quentin de La Tour, 1735.

As the monumental print edition of the Complete Works of Voltaire nears completion, the Voltaire Foundation is currently preparing the ground for Digital Voltaire, an interactive and innovative digital edition of Voltaire’s Œuvres complètes. The pilot project we are embarking upon will thus bring together two key existing datasets: TOUT Voltaire, developed in collaboration with the ARTFL Project at the University of Chicago; and Voltaire’s letters, drawn from Electronic Enlightenment. The combined dataset will include more than 20,000 individual documents and over 11 million words, making this one of, if not the largest single-author databases available for digital humanities research. This resource, together with a focused research project to scope and understand its potential uses and applications, will enable the Voltaire Foundation to begin to create a conceptual and infrastructural framework for a broader, transformational Digital Voltaire, for which fundraising efforts have already begun.

The Visualising Voltaire project will become part of the soon-to-be-created ‘Voltaire Lab’ – a virtual space for new research experimentation and dissemination centred on Voltaire’s textual output and its relationship to the broader field of eighteenth-century studies. By interrogating the ‘big data’ of Voltaire’s texts at both a macro- and microscopic level, we hope to shed new light on Voltaire’s use of intertextuality, his most commonly used themes and literary motifs, his intellectual networks, and his development as a thinker. This research project will further benefit from close existing ties with the ARTFL Project and the newly-established Textual Optics Lab at the University of Chicago, and with the Labex OBVIL (‘Observatoire de la vie littéraire’) based at the Sorbonne; centres for digital humanities research and development in French studies where much of this type of analysis has been pioneered.

Visualising Voltaire will include a number of literary experiments to test the scholarly and critical value of a combined digital archive of Voltaire’s texts. Following on from the work of Franco Moretti and the Stanford Literary Lab, the project will investigate how we can apply distant reading approaches to this large corpus in order to discover new connections and patterns at scale, and, at the same time, how these new approaches can interact and intervene with our traditional close reading modes of analysis. To this end, we have identified two areas of research that we will pursue in 2017-2018, and that we hope will lead to further projects in the future.

Sequence alignment.

Sequence alignment in the intertextual edition of Raynal’s Histoire des deux Indes, Centre for Digital Humanities Research, Australian National University.

In the first instance, we will focus on Voltaire’s ‘intertextuality’ and how computational techniques such as sequence alignment – borrowed from the field of bio-informatics – can help us better understand the rich complexity of Voltaire’s writing practices. Indeed, one of the major research questions that has arisen from the preparation of the Complete Works of Voltaire concerns Voltaire’s unacknowledged use and reuse of other texts. This takes two forms: the widespread reuse (borrowing/theft/imitation) of works by other writers, and the equally widespread reuse of his own work. This is a huge subject that has never been satisfactorily studied until now.

In a second instance, the completion of the Complete Works of Voltaire on paper has also created the opportunity to provide an index to the whole of his writings, notably using automatic indexing and classification techniques developed in the fields of artificial intelligence and machine learning. In addition to our ‘traditional’ indexes of the paper editions, which can be digitised and leveraged for computational analysis, we will also aim to generate ‘thematic maps’ of Voltaire’s works and correspondence using both supervised and unsupervised machine learning algorithms such as vector space analysis and topic modelling. These sorts of approaches will, we hope, open up Voltaire’s writings in wholly new and exciting ways, creating opportunities for high-profile public engagement activities such as hackathons, and generating new areas of investigation for potential doctoral research students.

Choix de Chansons.

From Jean-Benjamin de Laborde’s Choix de Chansons, 1774 – subject of the ARC Discovery grant ‘Performing Transdisciplinarity’.

And finally, beyond these specific research projects, my role as Digital Research Fellow will entail making and maintaining connections with digital humanities teams both locally and internationally, building on past and current relationships to generate new research initiatives moving forward. We are interested, for example, in establishing a better understanding of the importance of Voltaire’s Enlightenment network and its participation in the larger eighteenth-century Republic of Letters, questions that can be addressed in collaboration with the Center for Spatial and Network Analysis at Stanford, and the Cultures of Knowledge project based in Oxford. The Voltaire Lab can thus become a venue for engaging with other complementary Oxford digital projects, such as the Newton Project, which will allow for broader access as well as further fundamental research. Newton is often seen as the key thinker who sets the agenda for Enlightenment scientific thinking – through his emphasis on empiricism and the experimental method – while Voltaire, the dominant intellectual figure of the Enlightenment, helps to popularise Newton’s scientific method across Europe. Voltaire’s role as a key critic and disseminator of ideas and texts is also an area of research to which digital approaches can bring much to bear, in particular by linking his correspondence to projects such as Western Sydney University’s French Book Trade in Enlightenment Europe and Mapping Print, Charting Enlightenment.

We are equally keen to investigate the deeply interdisciplinary nature of Voltaire’s work beyond the purely literary or even textual, and, more generally, of his role in the often-overlooked interplay of music, images, and text in eighteenth-century print culture. This is in fact the subject of our recently awarded Australian Research Council Discovery Grant, ‘Performing Transdisciplinarity’, which brings together a team of interdisciplinary researchers from the Australian National University, the Universities of Melbourne and Sydney, and Oxford.

The above are just a few of the countless avenues of research opened up by digital approaches to Voltaire’s work and legacy, and to which many more will be added as the larger Digital Voltaire project takes shape over the next few years. As the newly appointed Digital Research Fellow at the VF, I very much look forward to keeping you all informed on the results of these experiments and of the project’s evolution in due course.

– Glenn Roe