Written by Priyamvada Nambrath (PhD Candidate, South Asia Studies) and Eleanor Webb (PhD Candidate, History)

During our time as Graduate Fellows at the Schoenberg Institute of Manuscript Studies (SIMS) (2025-2026), we explored how Handwritten Text Recognition (HTR) could be used to transcribe premodern manuscripts from Penn’s collections. HTR uses machine learning techniques to build models that can transcribe handwritten manuscripts. We came to the project with different disciplinary backgrounds and geographical specializations. Priya worked on UPenn Ms. Coll. 390 Item 1914, Item 2478, and Item 1167, all Indic manuscripts written in Devanāgarī script from the eighteenth and nineteenth centuries on omenology, devotional poetry and philosophy respectively, while Ellie worked on transcribing UPenn Oversize Ms. Codex 1663, a manual of mathematics, astrology, and chiromancy that was produced in Italy in the late seventeenth century. We nevertheless shared an interest as researchers in HTR and in exploring how using it might change the process of researching premodern manuscripts.

We worked alongside one another, and under the guidance of Lynn Ransom (Curator of SIMS Programs and Schoenberg Database Manager) and Jessie Dummer (Digitization Project Coordinator), with technical support from Doug Emery (Special Collections Digital Content Programmer). Dot Porter (SIMS Curator of Digital Humanities), and Matt Hunter (Head of Digital Scholarship, Research Data and Digital Scholarship) stepped in as needed with timely technical guidance and support.

We used the platform eScriptorium, an open source software that serves as an interface to Kraken, an automatic text recognition (ATR) system that is described on its website as “a universal text recognizer for the humanities.” eScriptorium allows users to segment manuscript pages into their individual lines and regions (title text, main text, marginalia, and diagrams) and to transcribe them. This data is used to create eScriptorium’s ground truth, or manually verified data. Kraken uses the ground truth to learn which characters on the written page correspond to which characters in the alphabet. In eScriptorium, segmentation divides the page into text regions and individual lines, allowing Kraken to learn from line-level images while the separation of different regions provides a basic analysis of page layout. The ground truth, which includes accurate page segmentation and transcriptions, can then be used to create a new transcription model from scratch, or to help refine an existing one. This makes it especially valuable for materials poorly served by conventional optical character recognition (OCR), such as handwritten texts. The advantage of eScriptorium’s open source status is that it is free for users and fosters a culture of open data between researchers and institutions. The processing capacity required for training models, however, required that the software be hosted on the university’s servers.

Could HTR alleviate the labor-intensive work of transcribing manuscripts? If so, what might be the downsides of taking this “shortcut”? Could HTR facilitate different kinds of engagement with the premodern manuscripts in Penn collections? Might the availability of HTR technology prompt or enable researchers to ask different questions about these manuscripts? As we began transcribing our manuscripts and testing our respective HTR models, these questions motivated many rich discussions in our weekly meetings. Conducting our transcription projects alongside each other also highlighted HTR’s different capabilities when applied to digitally well-represented scripts (like seventeenth-century Italian cursive) and to digitally under-represented scripts (like Devanāgarī).

We spent the first few weeks getting familiar with eScriptorium, learning a lot of new technical language, and beginning to build “ground truth” for our respective manuscripts (basically, lots of hours of old fashioned transcribing!) We first imported IIIF manifests for our respective manuscripts into eScriptorium (IIIF, the International Image Interoperability Framework, is a well-established standard for image exchange among institutions). We then asked the platform to “segment” each page. This is a process by which the software recognizes the lines, diagrams, and different regions of the page. We imported an ontology that organized the regions into recognizable categories, such as the main text, marginal text, diagrams, and others. When projects are exported from eScriptorium in XML format, the XML preserves the ontology and records the information about where on the page the handwriting occurs, in which region it occurs, and on which lines – so the output is interoperable with other systems in which researchers may want to display their work.

Ellie’s manuscript, Ms. Codex 1663, has a lot of mathematical diagrams and, in the section on chiromancy, several annotated diagrams of hands (Figure 1). For this manuscript, eScriptorium did a fairly good job of identifying the different regions of the page, although it had some difficulty in accurately numbering the lines. Ellie spent some time working through the manuscript and making minor corrections to this segmentation.

Priya began her work on Bulbul (बुलबुल), UPenn Ms. Coll. 390 Item 1914. This manuscript contains a text that is, as far as we know, not present elsewhere. It is an omenological work of performative sortilege, or the practice of fortune-telling by the random drawing of a card from within a collection. Both sides of the eight folios are identical in layout: on the left side is a painting of a bird with its name given below, and on the right side are ten numbered lines of text in the Hindustani language. The first page has a diagram of a bird titled ‘Bulbul,’ and this has been used as the title of the entire manuscript in the catalog record. Figure 2 shows the segmented version of a sample folio side of this manuscript, with the lines and regions indicated.

The next step was to begin transcribing the text to create the ground truth. This was the first step at which our projects diverged. Creating ground truth can proceed from scratch – by producing your own transcription – or by fine tuning a model built by other researchers (and made available on platforms like HuggingFace) by running the model over the manuscript and then correcting the transcription that model produces. The suggested amount of ground truth to build a reliable HTR model from scratch is considered by most estimates to be upwards of 1000 lines, while less is needed for fine-tuning a model. While many models are available for medieval and early modern Latin and other European languages (including Italian), we could find no models for Devanāgarī that worked in eScriptorium. Ellie was therefore able to apply a model to Ms. Codex 1663 and then refine it, while Priya was required to build ground truth from scratch for the Bulbul (बुलबुल) manuscript. She transcribed fifteen pages of the manuscript manually, and then used her transcription to read the sixteenth and final page, which was then used as the validation set for the model.

The process of creating ground truth required us to make many decisions that will be familiar to scholars who have produced a transcription or critical edition from a premodern manuscript. Ellie’s manuscript contained several common abbreviations in Italian (see Figure 3). What was distinct about producing a transcription using eScriptorium was the obligation to consider how the model would interpret the abbreviation. Unlike the normal transcription process, which is concerned with rendering the text to be comprehensible to a human reader, we needed to ensure our transcription would enable the computer to learn to recognize the characters depicted on the manuscript and connect them to known letters. This meant taking a different approach to abbreviations, errors, and scribal corrections. For instance, while one might usually transcribe the abbreviation for the word “per” in Italian diplomatically (p[er]), or silently (per), doing either in eScriptorium risked confusing the model by including letters that were not in the manuscript. It was therefore necessary to reproduce the text as accurately as possible – by, for instance, simply transcribing “p” (for per) and “dl” (for d[e]l). In the case of errors and scribal corrections, we tried to reproduce the visual appearance of the manuscript as best as possible in the transcription so that the model could learn to recognize the characters depicted. Ellie used unicode symbols to represent the zodiac and other astrological signs in her manuscript.Where errors or areas of damage rendered the text illegible, we excluded those regions from the computer’s scope of vision to avoid confusion.

Priya’s initial work on the Bulbul (बुलबुल) manuscript was an experimental effort, as this manuscript was smaller and more manageable than the others she was working with. The eight folios of this document together comprise around 160 lines, which seemed like an achievable amount of input data that needed to be manually input into the model. On transcribing the text over the next couple of weeks, however, she realized that all sixteen sides of the folios essentially had the same textual content even though their individual bird images were different: the same ten fortune outcomes were listed on each face in varying sequences. For purposes of creating the ground truth, this meant that the seven folios had essentially collapsed to just one folio, and the test results reflected this, returning an accuracy rate of just over 5%. As seen in the transcription output in the right of Figure 4 below, the model was unsuccessful in deciphering all the lines on the page, returning the same single, unreadable character for all ten lines. Given the highly restricted character set of the input data supplied to the training model, this quality of output was not really surprising. So this manuscript turned out to be a useful learning exercise in how to use eScriptorium, even though it was unsuccessful in generating a model.

In Ellie’s case, she applied a model that she found on Zenodo called catmus-medieval-1.6.0 to her manuscript. This is a Kraken HTR model that was trained on Old and Middle French, Latin, Spanish, and Italian. Interestingly, while this model returned a 94.9% accuracy rate on eScriptorium, a cursory examination of the transcription it produced for Ms. Codex 1663 indicated a far lower accuracy. Ellie transcribed over 50 folios (around 2800 lines) of the manuscript taken from different sections of the volume in order to produce a representative sample (while the scribe is consistent across the manuscript, the handwriting is rather messier in later sections, which also include more complex diagrams). She then trained the catmus model using the transcription. The first subsequent test of the trained model returned an accuracy rate of 93.3% – which was less than before any training! While this prompted discussion of how precisely eScriptorium produces its accuracy ratings, it also sent Ellie back on the hunt for models that might be better suited to her seventeenth-century manuscript. The task of searching for appropriate models is a learning experience in itself. While many models are available online, for the uninitiated researcher, it is not always immediately obvious where to find them or which one might be best suited for a particular manuscript. There are also many large and small gaps in the digital representation of premodern scripts, which is reflected in the available models. After her failed experiment with catmus, Ellie found another model – called Tridis_v2_Medieval_EarlyModern – which had been trained on early modern manuscripts. This model was markedly more successful at reading Ms. Codex 1663, and after several rounds of training produced an eScriptorium accuracy rating of 96.8% (and an Ellie accuracy rating of around 80%).

For Priya’s second attempt, she decided to select a more well-known Sanskrit manuscript, as this would help to improve the accuracy of her manual transcription. She chose a manuscript copy of the Saundaryalaharī Stotra (सौन्दर्यलहरी स्तोत्र), UPenn Ms. Coll. 390 Item 2478, a very popular tāntric work of contemplative poetry composed by the 8th-century philosopher-saint Śaṅkarācārya, transcribed in 1716, and comprising 37 folios (Figure 6). In terms of its physical characteristics, the manuscript is relatively small, measuring 10 × 16 cm, and is laid out with six lines per side. Consisting of 103 verses of four lines each, the manual transcription of ground truth took much longer than the earlier manuscript. She provided the textual input for 34 folios, and ran the resulting model on folios 35-36.

The resulting accuracy percentage from the training model generated using this document was 78.2%. Figure 7 shows a sample output page that was read and transcribed by the model. While there are several errors in transcription, this was not unexpected, given that the input data set, at around 200 lines, was still much smaller than the recommended size. Still, the significant improvement in output accuracy with the second manuscript, going from 5.3% to 78.2%, was quite encouraging. This motivated Priya to try to raise the accuracy by continuing to train the model with one more manuscript.

Priya’s third attempt focused on a well-known Upaniṣadic text, the Kaṭhopaniṣad (कठोपनिषद्), UPenn Ms. Coll. 390, Item 1167, which was accompanied by a vedāntic commentary, the Kaṭhopaniṣadvyākhyā, by Dāmodara, and transcribed in 1860. Consisting of 17 leaves, it is nevertheless a much larger manuscript than the Saundaryalaharī, with the leaves measuring 13 x 35 cm, and consisting of 9-11 lines per leaf. It is, therefore, twice the physical length of the Saundaryalaharī, with nearly twice the number of lines on each side. Two sample folio faces are shown in Figure 8. This manuscript offered additional interest as it also featured occasional vertical lines of text, and numerous corrections, editorial emendations and marginal annotations. A marked-up version of this page with regions, masks and lines indicated, is also presented in Figure 9. This model also showed remarkable improvement from the previous version, returning an accuracy rate of over 97%. The most persistent problem at this point was achieving a precise delineation of masks to completely include the mātras, or diacritical and combining marks that appear above and below the base letters in the script. Too often, the automatic masking process would assign these strokes to the sentence above or below the intended sentence.

Our experience with producing ground truth and testing our models put to rest any assumption we may have had coming into this project that HTR could substantially reduce the time spent transcribing manuscripts. In fact, as we moved through the process, we realized that there is still a long path ahead until HTR-produced transcriptions are widely available for premodern manuscript collections.

Towards the end of the Fall semester, we began to discuss the transcription capabilities of now well-known Large Language Models (LLMs), including ChatGPT 5.3 and Gemini 3. Having heard that other scholars had experienced some success using Gemini 3 to transcribe manuscripts, we began to experiment with that platform. We began by simply uploading images of individual folios from our manuscripts, and prompting the LLM to “transcribe this page exactly.” This was another point at which our experiences were markedly different. In the case of Ms. Codex 1663 – a fairly legible seventeenth-century cursive in Italian – Gemini was remarkably successful in accurately transcribing the folio provided. The LLM also (unprompted) identified the text as part of an early modern manual on mathematics.

In Priya’s case, the LLMs were also successful in identifying the text of Ms. Coll. 390, Item 1167 as belonging to the well-known Kaṭhopaniṣad, and providing relevant background information on the work. But these models failed to return an acceptable transcribed output, and consistently hallucinated in places where the manuscript text differed significantly from standard printed versions available online. This was most noticeable in the commentarial text of Dāmodara, which is not so easily available in print or online versions. This would imply that these AI models are not yet entirely reliable sources for the transcription of handwritten Sanskrit text, and human monitoring and verification of their output is still essential. In the case of Sanskrit at least, the LLM’s ability to read handwritten text very likely appears to be based on reading enough text to establish a correspondence and identification with a known source if possible, and once that identification has been made, in then extracting the output from a printed or online version, rather than the input manuscript. This would help explain why variations in the manuscript text from widely available print and online versions were largely resolved in favor of externally sourced standard versions over the input manuscript. Fine-tuning the prompt to disregard external sources and only read the input document was not successful, as the models continued to hallucinate. The LLMs’ tendency to hallucinate is of course not isolated to non-European languages and scripts. In the course of her experimentation with Gemini, Ellie also found that without very clear and bounded instructions, the LLM would “create” parts of the text, or simply repeat transcriptions for parts of the text it had already produced.

Two sample LLM outputs can be seen in Figures 10-12 below, which show, respectively, a sample source text, and the corresponding outputs using ChatGPT 5.3 and Gemini 3 AI transcription of Ms. Coll. 390, Item 1167. For both models, the instructions were coded to return the ‘@’ character for unreadable characters in the manuscript. While this character is seen to be present in both LLM outputs, it is interesting that there are far more errors of the model misreading the manuscript and returning an incorrect result that it believes to be correct. In fact, the greater part of the manuscript folio has been misread by both ChatGPT and Gemini, reinforcing the need for careful human involvement and mediation of the output produced by LLMs on handwritten manuscripts. Within this project, LLMs were useful in providing a tentative identification of the manuscript with much useful background information, but not for reliably accurate transcription of handwritten text.

We were assisted in our efforts to get the LLMs to obey our requests by Davide Pafumi, a doctoral student at the University of Lethbridge and visiting scholar at SIMS, who led us through the process of prompt engineering to improve output quality. We also experimented with asking the LLMs to produce transcriptions of multiple pages of Ms. Codex 1663 at a time using a CSV table. The CSV separated each folio image (using links from the IIIF manifest) into rows, with a blank column for the LLM to input its transcription. This was a largely successful experiment: the LLM produced a highly accurate transcription, and accurately described the layout of the text on each folio. While this process (and the more primitive one of manually inputting images into the LLM chat) are both rather cumbersome at present, Ellie found that the transcriptions produced by Gemini 3.0 were on the whole more accurate than those produced by the eScriptorium models she tested. The AI’s tendency to hallucinate meant that its outputs still required careful checking.

What ultimately proved difficult to resolve was that the LLM model largely functions as a black box, and the user has limited control over how the model decides to interpret the user’s instructions. Over the course of this project, eScriptorium therefore emerged as a more reliable alternative, although the manual transcription process involved in creating a sufficient quantity of ground truth can be painstakingly laborious and time-consuming. Regardless, it offers a much greater degree of control and determinability in training the system to respond to carefully defined and restricted inputs.

At the same time, the question of how LLMs might be responsibly and productively used by researchers remains an open one. With appropriate development and guardrails LLMs may offer opportunities for the large-scale transcription of premodern manuscript collections. Our experimentation with and continued discussion about the potential opportunities afforded by AI – as well as its issues and potential risks – prompted broader reflection on how scholars and librarians can engage with AI productively and responsibly. We discussed how AI tools could be thought of in terms of the medieval categories of “mechanical” and “thinking” arts. The former denotes manual, material labor, and the latter denotes contemplative, interpretive and theoretical work. HTR is a good example of an area in which AI might, if properly controlled, do the “mechanical” work of transcribing a medieval manuscript. In their current form LLMs are most successful and most palatable to researchers when confined to (in tightly controlled parameters) the performance of such mechanical tasks, rather than being let loose on the interpretive work, say, of analyzing the meaning of a manuscript and writing an article about it. Interestingly, one of the most fruitful results of the year-long project was that it brought together humanists and software specialists to work in close and animated collaboration to accomplish results that neither might have accomplished alone. It is clear to us that scholars, archivists, and librarians have a lot to offer ongoing discussions and debates about what AI and machine learning can and cannot contribute to scholarship, education, and culture.