Transcription Guide
We are in the process of transcribing source materials into Wiki pages for each treaty. To the best of our knowledge, these source documents have not been transcribed in this way before, and we aim to present them in a manner that is both accessible and supportive of research and analysis. The materials available for each treaty include images of the original treaty document and transcriptions of the text from the images in .txt format.
The .txt files are created using the Tesseract OCR software which analyzes the image and converts the document into a text file. The software can sometimes misinterpret the images, have trouble with spatial formatting, or create digital particles that show up incorrectly in the text. As a result, the text in the .txt files need to be edited to match the text in the images available in each treaty page, and then added to that page.
For example, on treaty 24's (old version) page, the file 24.txt can be found at the bottom of the page:
Opening the file reveals Tesseract-OCR’s attempt at converting the image to text:

Some letters and characters have been misanalysed by the software, carriage returns are at the end of every line, other particles and blemished created by the printing, scanning/digitization processes have resulted in other transcription errors that need to be corrected in our editing process to get the document ready to be published on the treaty page. Since treaties in the document often start or end part way down the page, the text files will also include text from preceding and following treaties.
Import the contents from the text file into a word processor to make the corrections before publishing the transcription onto the treaty’s wiki page.
The section of each treaty containing signatures in two columns creates additional Tesseract OCR transcription errors which need to be edited as follows:

See the page for treaty 24 to see the final transcription.
Some treaties contain tables of items traded as part of the treaty agreement which also pose difficulties for the Tesseract-OCR transcription and which need to be manually edited in the final transcription. See treaty 5 for an example of how this table is incorporated into the transcribed document.
Using AI to help with transcription tasks.
Many transcription tasks are time consuming, AI platforms and other editing software may help speed up these tasks. See the following gallery chat with an AI bot which helped process some of the editing for Treaty 30's transcription.