Jun 4, 2026 Matt Basso, EGD and History, 2026 DM Faculty Grant Recipient
“Using Artificial Intelligence to Examine the Oral Histories of 10,000 Veterans.”
Briefly describe your project and the challenges, lessons learned, and obstacles overcome in the execution of it. What were the professional, academic, and personal motivations underlying your project?
The goal of this project was to build and, with the assistance of artificial intelligence technology, analyze a database of approximately 10,000 oral history interviews of military veterans who fought in US wars between 1945 in 2023. While we were confident we could create a database of 10,000 oral histories with transcripts, after initial success we set ourselves a bigger challenge: find as many veterans oral histories as we could in archives in every state in the union. To date we have found more than 43,000 interviews and met our geographic objective. However, two unexpected challenges emerged with the nation’s largest collection at the Library of Congress. First, batch downloading interviews proved more difficult than expected. With the assistance of undergraduate Data Sciences major Aiden De Boer, we have just solved this problem. In the process we discovered that a large percentage of interviews – 2/3 of our first batch download – that the LOC listed as having transcriptions did not. Instead, they were only available as sound or video files. To analyze these at scale with an LLM will necessitate creating a rapid transcription process, which we believe we can do but it is outside of our scope for this grant. Given the number of oral histories we have found, we also won’t need to transcribe these to reach our 10,000-interview threshold. The question of using AI to work with audio raised a second issue – and a remarkable new possibility. In collaboration with the U’s RAI/SCI team, we submitted a grant to build an AI platform able to interpret the prosody – how something is said, not just what is said – of interviews. Doing such analysis, though somewhat rare even on a small number of interviews, is considered a cornerstone of cutting-edge oral history interpretation.
My motivation for undertaking this project is based on my long history of working with oral history. I have used oral histories in my scholarship for 30 years, have led oral history projects for 20 years, and have taught oral history classes for more than 10 years. In the first and last of those contexts I have only ever imagined using a small number of oral histories. That is because while they are rich sources, they are complex and time consuming to analyze. Assessing one 90-minute interview can take a day; if you’re really pushing it, you might be able to get through three in that time.
Last year I began wondering whether large language models (LLMs) could facilitate the analysis of a previously impossible numbers of oral histories. Two things motivated me in this direction. First, I wondered what new questions I and others could ask and what patterns we might see if we could work with interviews at a much larger scale. Second, I thought that perhaps we could get both scholars and the public to listen to many more oral histories if we could provide them a tool that opened new possibilities for learning from these compelling sources. And because people like me, in both the academy and in the community, love taking oral histories – there are so many of them! Our estimates show that there are at least 1,000,000 oral histories in large and small archives around the world. There have been more than 400,000 interviews taken by StoryCorps volunteers alone. One of my long-term goals prompted by this DM grant is to help build an oral history and AI
platform that makes accessing oral histories much easier and allows users, including the public, to ask fascinating new questions based not just on scope and pattern, but also on comparative analysis.
How did the Digital Matters internship dovetail with your academic pursuits? What interested you in applying for this grant?
I think of Digital Matters as an enabler of new approaches to core research questions. That is what makes the DM grant and internship program such an incredible opportunity. My sense of what DM could offer wasn’t theoretical. I had the good fortune of having a DM grant several years ago that allowed me to think through how to more dynamically present World War II home front history, one of my specialty areas, to a public audience. I see this oral history and AI project in the same light – though I’m now sure it has the chance to offer far more innovative findings.
What insights have you gained in regard to your specific field as a result of your project and grant experience?
My early use of AI to assist with identifying patterns across many oral histories has underscored both the ways individuals share some common aspects of a national story and how intriguing when differences emerge. Even more fundamentally, I’m now convinced there are fascinating conclusions that can come from looking at hundreds if not thousands of oral histories to answer research questions. I believe that, done well, the magic of individual stories, which is what attracts so many people to oral history, can be retained at the same time as scholars use this new tool to develop new arguments about the experience of veterans – or any other group or topic that has been the focus of oral history projects.
What would you tell potential intern applicants to help them shape their own digital scholarship project?
For faculty, like me, who have limited digital skills – don’t be scared to apply. Think of DM as a space where you can work collaboratively bringing your expertise and joining it with that of potential research partners. For me that has been graduate students, like History PhD student Cathy Gilmore, who has stronger digital skills than I do and an interest in oral history, and undergraduates like Aiden De Boer. DM grant funds are a great way of supporting our excellent students!