Computer Science
History
History Department
Department of Computer Science and Software Engineering
The Historia Augusta is an anonymous collection of biographies of 30 Roman emperors who ruled from 117 to 284 CE and serves as the principal Latin source for this period of Roman history. Although the work is attributed to six named authors, these identities are widely considered pseudonymous, leaving its true authorship unresolved. This project applies computational stylometry to evaluate whether contemporary writers, including Aurelius Victor, Symmachus, Eutropius, Festus, and Julius Obsequens, exhibit stylistic patterns consistent with the Historia Augusta. While no single author demonstrates overwhelming similarity, the results capture meaningful stylistic signals and align with prior observations of overlap with Eutropius. These findings suggest that, despite limitations in available data, computational methods can reflect established patterns and provide a scalable framework for future authorship analysis.
Modern scholarly consensus is that all of the biographies were written by a single author in the late 4th or early 5th century who, for whatever reason, wished to remain anonymous. This project proposes using machine learning to attempt to solve the mystery of the authorship of this important ancient work. Machine learning techniques have been used successfully in several studies of ancient literature, particularly in stylometry studies analyzing word selection and grammatical structure to identify authors. Since 2005, over 225 articles have been published on machine learning and ancient languages. A recent example was an analysis of the works of the Roman historian Tacitus, comparing his writings to the earlier historian Livy, searching for passages in Tacitus quoting Livy without attribution, which was successful. The first step has been to identify contemporary authors who plausibly might have written the Historia Augusta as targets for testing. The selected targets were Aurelius Victor, Symmachus, Eutropius, Festus, and Julius Obsequens. In addition to being from the correct time, these authors have left substantial bodies of work which - like the Historia Augusta - are available in digital formats, giving us enough ready material for an accurate analysis.
Data
This project uses texts from the Historia Augusta alongside works from known Latin authors, including Aurelius Victor, Eutropius, Festus, Julius Obsequens, and Quintus Aurelius Symmachus. All texts were obtained from public domain sources and combined into a unified corpus for analysis.
Preprocessing
All text was normalized by converting characters to lowercase and removing non-content metadata such as section headers. Orthographic variation was reduced by standardizing Latin character usage (e.g., treating j as i).
Chunking Strategy
Texts were divided into fixed-length chunks to enable consistent comparison across authors. Remainder chunks smaller than 80% of the target size were discarded to avoid unstable samples. A chunk size of 500 words was selected to maintain stylistic signal while maximizing the number of usable samples under limited Latin data availability.
Feature Extraction
Character-level 4-grams were extracted from each text chunk. A global frequency analysis was performed across the corpus, and the top 400 most frequent n-grams were selected as features for comparison. Each chunk was represented as a normalized frequency vector over these features.
Similarity Measure
Burrows' Delta was used to measure the stylistic similarity between text chunks. Feature vectors were standardized using z-scores, and distances were computing using Manhattan (L1) distance.
Visualization
Principal Component Analysis (PCA) was applied to the standardized feature space to visualize clustering patterns between chunks for assessment of stylistic similarity between authors.
PCA visualization shows partial clustering by author, with some overlap between texts. Notably, chunks from the Historia Augusta exhibit proximity to those of Eutropius, suggesting stylistic similarity. However, no author forms a clearly dominant cluster, indicating that attribution cannot be definitively resolved using the available data.
This study found no single author with overwhelming stylistic similarity to the Historia Augusta, suggesting that authorship cannot be definitively attributed using the available data. However, the results are consistent with prior scholarship indicating a stylistic overlap between the Historia Augusta and Eutropius, demonstrating that the methodology is capable of capturing meaningful signals.
These findings also highlight key limitations, particularly the scarcity and uneven distribution of available source texts, as well as variation in writing domains, both of which introduce instability into the analysis.
Future work will focus on expanding the corpus as additional texts become digitized, allowing for a broader set of candidate authors and more robust comparisons. The methodology presented here provides a scalable framework that can be applied to larger datasets and refined with additional features to improve the reliability of computational authorship studies in classical literature.
The following is an image of poster presented at the 2026 Undergraduate Research Forum [remember to include alt text]
Martins, Armando, Clara Grácio, Cláudia Teixeira, Irene Pimenta Rodrigues, Juan Luiz Garcia Zapata, and Lúgia Ferreira. “Historia Augusta authorship: an approach based on measurements of complex networks.” Applied Network Science 6 (2021): 50. https://doi.org/10.1007/s41109-021-00390-7.
This project developed several NACE Career Readiness Competencies, including Technology, Career & Self-Development, Teamwork, and Communication.
Technology skills were strengthened through the implementation of a computational stylometry pipeline using Python, including data preprocessing, feature extraction, and statistical analysis. The use of version-controlled code and digital collaboration tools reflected industry-standard practices.
As an interdisciplinary collaboration between History and Computer Science, the project emphasized teamwork and communication. This required translating domain-specific concepts across fields and ensuring that both technical methods and historical context were clearly understood by all contributors.
Career & Self-Development was demonstrated through independent learning of stylometric techniques and machine learning applications in classical studies, as well as adapting to challenges such as limited and uneven data.