Not Just “Can We?” but “Should We?” and “Why?” Understanding Digital Manuscripts as (Big) Data

Last Spring, Bridget Whearty and I published our article “Not Just “Can We?” but “Should We?” and “Why?” Understanding Digital Manuscripts as (Big) Data” in Digital Philology. Here is a taste of that, and a link to the open access copy if you would like to read more.

Abstract: It has been over a decade since the runaway popularity of “big data” hit industry and academia in the early 2010s. Today, big data and the artificial intelligence (AI) programs that feed on them are no longer being embraced with quite as much enthusiasm. With this growing awareness of the darker sides of (big) data and AI as its foundation, this article explores the challenges of working with datasets of digital manuscripts, particularly for those who seek to treat digital manuscripts as big data. We examine big data approaches that focus on digital manuscript images and those working with structured manuscript metadata, outlining what end-users must be aware of before plunging into the world of digital manuscripts as big data. Throughout, we emphasize the ethical and environmental costs in using machine learning and other AI approaches to analyze massive datasets of medieval manuscripts. The article concludes with a series of considerations for researchers who want to engage with digital manuscripts as (big) data.

Read more here