Lead
Deep within the restricted archives of the Vatican Library, a 408-page handwritten manuscript scrawled with 34 obscure symbols, Roman letters, and an Arabic title page remained undeciphered for more than four centuries. Known to scholars as the Borg cipher, the text's cryptic contents have finally been unraveled by an international research team utilizing machine learning algorithms. According to a report published by BBC Future, the computational achievement represents a major milestone in digital humanities, demonstrating how artificial intelligence can streamline the painstaking process of historical codebreaking.
Nut Graph: Why Computational Decipherment Matters
Archival estimates indicate that approximately 1% of all historical manuscripts preserved in global libraries, religious institutions, and state records remain fully or partially encrypted. Historically, deciphering even a short coded letter required months or years of manual statistical analysis and trial-and-error hypothesis testing.
By training machine learning models on linguistic patterns, character substitution tables, and frequency distribution profiles, computational linguists can now process encrypted archives at scale. This algorithmic shift is beginning to expose critical historical information that was previously missing from academic literature, spanning secret diplomatic intelligence, medical formulas, political conspiracies, and the private correspondence of historical figures.
Deep Dive: Unraveling the Vatican's Borg Cipher
The Borg cipher—cataloged under shelfmark Borg.lat.898 at the Vatican Library—was long suspected of concealing secret medical remedies. During the early modern period, unorthodox healing practices were often kept heavily encrypted to avoid religious suspicion or accusations of witchcraft.
According to Beáta Megyesi, a professor in computational linguistics at Stockholm University who participated in the decoding project, analyzing the manuscript required overcoming significant structural and physical obstacles:
Document Scale and Composition: The manuscript spans 408 aged pages composed primarily of 34 distinct cipher symbols, interspersed with occasional Roman letters and an opening page inscribed in Arabic script.
Physical Degradation: Decades of ink fading, paper deterioration, and physical damage to page margins created substantial visual noise, making initial transcription difficult.
Cryptographic Structure: The research team determined that the manuscript utilized a monoalphabetic substitution cipher, in which each unique cipher symbol corresponds directly to a specific Roman letter.
Decrypted Medical Remedies: Once the cipher key was established, the decoded text exposed thousands of unusual medical recipes and treatments. Examples included fermenting nutmeg inside dough to treat dysentery and prescribing precise measures of high-quality red wine for various bodily ailments.
Megyesi characterized the computational process as algorithmic detective work, noting that even with machine learning assistance, identifying partial character patterns and building a functional cipher key required extensive manual verification.
Cryptographic Complexity: Why Historical Codebreaking Is Difficult
While substitution ciphers like the Borg manuscript are conceptually straightforward once character frequencies are matched, many historical ciphers were intentionally engineered to defeat frequency analysis.
As detailed in the BBC Future report, historical cryptographers employed various noise techniques to prevent unauthorized reading:
1. Nomenclators and Homophonic Substitution
Rather than mapping one symbol to one letter, advanced historical ciphers assigned multiple distinct symbols to represent a single high-frequency letter (such as 'E' or 'T'). This flattened the statistical frequency curve, rendering traditional letter-counting techniques ineffective.
2. Null Characters and Decoys
Scribes frequently inserted "nulls"—meaningless extra symbols—throughout the text to disrupt word boundary recognition and throw off prospective snoops.
3. Code Words and Shorthand
Important entities, such as monarch names, military locations, or sensitive topics, were often replaced with arbitrary code words or shorthand symbols that bore no relationship to the surrounding language's phonetic rules.
4. Unknown Source Languages
Codebreakers frequently face situations where the underlying language or regional dialect of the original plain text is completely unknown, requiring machine learning systems to test multiple linguistic dictionaries simultaneously.
5. Manual Transcription Bottlenecks
Before codebreaking software can analyze a document, handwritten manuscripts must be converted into digital, machine-readable formats. Faded inks, erratic historical handwriting, and non-standard spelling conventions make this initial digital transcription phase labor-intensive.
Cecile Pierrot, a cryptologist at the French National Institute for Computer Science Research (INRIA), noted that transcribing just two pages of unfamiliar handwritten cipher symbols typically requires a full day of meticulous manual work. Pierrot and her team spent six months unravelling a three-page letter from 1547 written by Charles V, the Holy Roman Emperor and King of Spain. Encrypted using 120 distinct symbols, the decrypted text revealed that the powerful monarch was paralyzed by fear over an alleged assassination plot led by an Italian mercenary serving French King Francis I.
Historical Case Studies: Rewriting Known Narratives
The application of digital decryption techniques has already produced major historical discoveries that alter current academic consensus.
A notable example cited by BBC Future involved the recent discovery and decipherment of an extensive collection of coded letters written by Mary, Queen of Scots, during her 19-year imprisonment in England. Once decrypted, the primary documents revealed intimate details regarding her active participation in political plots to overthrow Queen Elizabeth I and regain her throne, as well as her complex, tense diplomatic negotiations with her son, King James VI of Scotland.
Technical Breakdown: The Role of AI vs. Human Researchers
While media headlines often suggest that artificial intelligence breaks historical codes independently, computational linguists emphasize that machine learning functions as an accelerating tool within a hybrid workflow:
| Stage of Decipherment | Role of Machine Learning / AI | Role of Human Researcher & Historian |
| Image Preprocessing | Cleans visual noise, enhances faded ink contrast, and isolates individual glyphs. | Verifies character segmentation accuracy on damaged manuscript physical edges. |
| Digitization & Transcription | Optical Character Recognition (OCR) clusters visually similar symbols into digital categories. | Manually transcribes highly irregular or faded handwriting where OCR confidence is low. |
| Pattern Analysis | Evaluates character frequency distributions and identifies recurring symbol sequences across large datasets. | Identifies historical context, jargon, and likely regional dialects used by the scribe. |
| Key Generation | Iteratively tests thousands of potential substitution keys against historical linguistic dictionaries. | Evaluates whether resulting plaintext outputs form grammatically and historically coherent sentences. |
What Has Not Been Confirmed
Fully Automated Decryption: Current machine learning models cannot decipher complex historical manuscripts autonomously without human expert intervention in transcription, context validation, and language selection.
Universal Application: AI codebreaking techniques optimized for simple substitution ciphers cannot automatically solve highly complex polyalphabetic ciphers or visual steganography without custom model adaptation.
What Happens Next in Digital Humanities
Research groups led by Beáta Megyesi and international collaborators are currently building standardized digital pipelines to transcribe and index historical ciphers systematically. By combining automated character recognition with statistical language modeling, computational linguists aim to create open-access databases that will allow archives worldwide to digitize, decode, and catalog their encrypted collections, making centuries of hidden primary sources accessible to global historians.
Sources
Primary Reporting
(Author: Sandrine Ceurstemont)BBC Future - Plots, love letters and remedies: The medieval secrets being revealed by AI Photo Credit:
Photo: 2026 Biblioteca Apostolica Vaticana / Beáta Megyesi
.png)
0 Comments