Bolívar, Márquez, and the Machine


Simón Bolívar lies in bed, surrounded by military officers and a priest.

Advances in digital data and computing power have made machine learning a juggernaut in modern life. From ChatGPT to music recommendations to Amazon pricing, it often feels like life is increasingly measured and evaluated in ways that once could not be predicted. Fueled by machine learning, Natural Language Processing (NLP) has proven itself one of the most pervasive technologies in our increasingly digital lives. As a tool, it has long fueled search engines and automated grammar checking on the internet, and with the advent of Large Language Models (LLMs), it has only grown more relevant. Via NLP, computational methods inevitably found their way into the world of literature. While literary academia has long relied on close reading of texts, NLP frameworks allow for what scholar Franco Moretti calls “distant reading,” using computing techniques to provide a quantitative understanding of literature that traditional close reading cannot provide.1

Digital humanities methodologies are already disruptive in the field of literature, and they take on a new nature when used in the context of historical fiction. Historical fiction already exists in an odd paradigm. In novels and cinema, historical fiction can capture and shape the popular zeitgeist of history among the general population without necessarily resembling history as understood by scholars. After all, this is a genre often associated with “hacks and clichés” and the ripping of bodices.2 On the other hand, historical fiction can offer intelligent perspectives on history only possible through fictitious presentation. Narrative can fill gaps in historical knowledge, challenge understandings of history, and innovate scholarship from the outside.

An excellent example of an innovative perspective on history achieved through fiction comes through the Colombian literary luminary Gabriel García Márquez’s writing about the Liberator himself, Simón Bolívar. While United States-based audiences may be unaware, the Liberator is a juggernaut in the South American zeitgeist as the man who freed six nations from the Spanish. His portrayal in works of historical fiction often reflects this, such as the heroic fictional Bolívar portrayed by Édgar Ramírez in the 2013 Spanish-Venezuelan film Libertador.3 In his 1989 book The General in His Labyrinth, Gabriel García Márquez paints a far more transgressive picture: Bolívar traversing through Colombia at the end of his life in a state of sickness and disillusionment.4 This work radically differed from the Bolívar central to the founding national myth of many countries and was controversial among literary critics after its release.5

Applying NLP to a work as influential as Bolívar’s is an avenue for a new understanding of the novel and its commentary on Latin American history. This project aims to understand The General in His Labyrinth through “distant reading,” as Moretti would describe it. Methodologically, it will begin with basic counting-based analysis made simple through computing and work up to more computationally intensive NLP techniques.

Why Do This?

This project exists within the discipline of the digital humanities, which has faced intense criticism. University of Minnesota professor Timothy Brennan argues that digital humanities methods confuse “more information with more knowledge,” that computers can only answer the questions suited to their limitations, and that the field ultimately functions as “a wedge separating the humanities from its reason to exist.”6 These critiques have merit — quantifying language requires embedding subjective assumptions into algorithms, and the questions computational methods answer are not always the ones literary scholars have long wrestled with. Nevertheless, NLP methods reveal patterns in texts that no close reader could reliably detect, and when paired with historical literacy they can illuminate how literary works reflect and create cultural phenomena. Furthermore, as Franco Moretti has shown, distant reading can produce findings (like the century-long shift in British novels from moral to descriptive language) that are simply beyond the reach of traditional scholarship.7 Applied to a work of historical fiction as influential as The General in His Labyrinth, these methods carry the additional virtue of accessibility: numbers and visualizations can engage a general audience that literary criticism typically cannot reach.8

Márquez and Digital Humanities

Widely considered the most famous Latin American author, Gabriel García Márquez’s novels have been subject to intense scholarly review. Texts like One Hundred Years of Solitude have long captured both the scholarly and popular imagination. Despite this, existing digital humanities scholarship on his works is sparse. The two closest examples of digital humanities work on Márquez come from a hobbyist project and a scholarly book involving archival and network analysis, both focused on One Hundred Years of Solitude.

In an informal but interesting analysis, Francesco Cauteruccio subjected the seminal Márquez novel to a computational analysis, focusing on word frequency, lexical richness, collocation patterns, and degree of centrality for all characters in the novel.9 The analysis found the novel contains 144,739 words with 11,027 distinct words, and used network graph methods to rank characters by importance, with Úrsula Iguarán ranking highest by degree centrality at 33 edges, followed by Amaranta and Remedios.10

Notably, the analysis used the novel’s English translation and ran the text through NLP frameworks using Python.

In contrast, sociologist Álvaro Santana-Acuña’s book Ascent to Glory: How One Hundred Years of Solitude Was Written and Became a Global Classic traces the novel’s development across its manuscripts through network analysis.11 The author used new documents from Márquez’s archive in addition to manuscripts, intermediary versions, and early drafts of the novel to trace textual evolution and map the social networks that made the book’s publication possible.12 While not as clear an NLP-based analysis as that of Cauteruccio, Santana-Acuña’s book, in principle, resembles techniques used in the digital humanities such as the analysis of networks within literature. His analysis importantly worked with the original Spanish.

These two examples represent the vast majority of digital humanities scholarship done on Márquez. Few peer-reviewed studies use NLP on Márquez novels, and none work with The General in His Labyrinth. This project therefore occupies a unique niche. Unfortunately, because research on the novel is limited, there is little prior scholarship to base this analysis on. Conversely, this means the project makes a novel contribution to the analysis of Latin American historical fiction. This work draws on traditional literary criticism. Literary scholarship on The General in His Labyrinth has emphasized the novel’s power to humanize a deified historical figure, and these analyses shape the flow of analysis in this project.13

It is important to note that this project works using the English translation of the novel, specifically the 1990 Alfred A. Knopf edition of the Edith Grossman translation. While previous scholarship has worked in both translation and the original language, it would be ideal to work in Spanish, since translation changes the novel’s word choice. However, many of the tools in the NLP toolkit exist primarily in English. The conclusions of this work should be held in the context of Edith Grossman’s translation, a potential drawback of this analysis. Working with the original Spanish would be a worthy extension of this project, but it would require creating NLP data sets beyond this project’s scope.

The Book Itself

Chapter lengths. Names are assigned from each chapter’s opening phrase.

The General in His Labyrinth is an odd book. Consisting of 8 unnumbered and untitled chapters, the book traces Bolívar’s journey up the Magdalena river and ends with him dying in poverty, a shadow of the man he used to be.14 For the sake of clarity, this analysis assigns names to each chapter based on their first phrases. The book typically refers to Bolívar as simply “The General.” In a typical NLP analysis, the entire text is flattened to lowercase. This would be inappropriate in the context of the novel, as it would eliminate the distinction between the word “general” and its use as Bolívar’s title. Thus, this analysis treats the word’s use as a title as “General_TITLE” and considers it a distinct word, with the rest of the text flattened. In structure, the book is quite repetitive. When reading the novel, this comes across as symbolic of Bolívar’s plodding journey to his death. This is revealed in the chapter structure as well, each being roughly the same length despite their unclear division.

Word-Frequency Analysis

Before running any complex NLP algorithm, it helps to understand word frequency and choice at a high level with a simple word frequency analysis.

The most common words across the whole book include “the,” “of,” and “and.” In NLP, words like these (words that connect ideas rather than express them) are called “stop words” and are often stripped from the text to give a more meaningful understanding of the book.


With “stop words” removed, word frequency becomes more meaningful. The most common word by far is “General,” as in referring to Bolívar. José Palacios, Bolívar’s mayordomo, comes next. In the novel, Bolívar’s state of sickness leaves him highly dependent on those around him, and this often falls to Palacios, perhaps the most instrumental character in the novel. The word “time” is also very frequent. Characters in the novel are highly aware of time — both its passing and its dictates over their journey.

Sentiment & Emotion Analysis

The next direction in NLP is analyzing sentiment across the novel. As sentiment, the emotion behind the words of the novel, is subjective, NLP analyses typically rely on sentiment dictionaries, which assign sentiments to words. This analysis uses three dictionaries: AFINN, Bing, and NRC. The Bing dictionary assigns a binary positive/negative value to each word in its set, while AFINN numerically rates words from most to least positive. The NRC dictionary is the most complex, as it assigns words to 8 different emotions (Anger, Fear, Anticipation, Trust, Surprise, Sadness, Joy, Disgust) and two sentiments (positive or negative). As stop words occasionally impact sentiment, we will compare the three sentiment dictionaries across the novel with and without them.

Sentiment across the novel, with and without stop words.

With and without stop words, we see negative sentiment increase as the novel progresses, quantitatively showing the reader’s sense that the characters are slowly circling the drain as the plot moves forward.

The NRC dictionary’s emotional profile is also useful. Tracking the eight NRC emotions across chapters reveals how the emotional texture of García Márquez’s prose shifts as Bolívar’s journey progresses.

NRC emotion counts by chapter.

The heatmap above shows which emotions dominate each chapter. Fear and sadness tend to darken as Bolívar nears death; anticipation and trust signal chapters where he still holds political authority or meaningful relationships. Interestingly, words coded as signaling anticipation are common in the novel and increase toward the end. This is darker than it sounds — the prose of the novel reveals anticipation of Bolívar’s death increasingly reaching its zenith.

Distinct Words by Chapter

Term Frequency – Inverse Document Frequency (tf-idf) is a framework for evaluating what a document is about by examining which words in it are the most distinctive.15 Term frequency, as examined earlier, can only tell us so much. Inverse document frequency decreases the weight for commonly used words and increases the weight for words that aren’t used very much within the document. Multiplying this by term frequency creates the TF-IDF score, which adjusts a term’s frequency by how rarely it is used. This can identify which terms are most distinctive in one document relative to others. In this case, we can use TF-IDF scores by chapter to show which words are the most distinctive for each chapter of the novel.


TF-IDF shows mainly which characters and objects are the most important for each chapter. “Reverend” popping up in the final chapter refers to Bishop Estevez’s appearance toward the end of the novel. He is the clergyman summoned to hear Simón Bolívar’s confession on December 10, 1830, as the General’s health rapidly declines. The private meeting lasted fourteen minutes, left the bishop visibly unsettled, and he later refused to officiate at the funeral.16 The priest serves as a counterpoint to the General’s final journey, emphasizing the spiritual decay and isolation accompanying his physical decline.

Where the Machine Fails

Latent Dirichlet Allocation (LDA) is an unsupervised machine learning algorithm that reads a collection of documents and tries to discover hidden “topics” running through them.17 In the case of the novel, LDA attempts to find four topics, then describes them with representative words.

The four latent topics revealed by the algorithm were as follows:

  1. Bolívar, Palacios, and war
  2. Bolívar and Palacios as related to their bodies
  3. Bolívar and Palacios as related to time and date
  4. Bolívar and Palacios as related to location

The clearest next step from this is to evaluate where these topics are most common in the novel. The simplest way to do this is to track topics by chapter, as shown here with a heatmap.

LDA topic mixture by chapter, using four topics.

The above heatmap shows that topics in each chapter often don’t intersect. Anyone who has read The General in His Labyrinth knows this isn’t true in the novel. In this case, this likely indicates that the LDA topics aren’t very distinct—they mainly relate to Bolívar and Palacios in four different ways —so the algorithm splits them arbitrarily. Either LDA isn’t a good method for understanding this book, or getting more insight requires a more bespoke implementation. A better way to understand connections between characters and topics in the novel, therefore, is through network analysis.

Networks in the Labyrinth

The final method in this project evaluates co-occurrence of thematic words within the novel.

Character co-occurrence correlation network. All figures by Nathan Hertzberg.

Characters with high correlations frequently share scenes; low or absent edges suggest parallel but separate narrative threads. This reveals interesting relations. Fernando, Urdaneta, and Iturbide are highly related. All three are figures from the independence era who represent failed or compromised republican projects. Urdaneta was Bolívar’s most loyal general and briefly seized power in New Granada in 1830 in a last-ditch attempt to keep Bolívar’s project alive; Iturbide was the Mexican independence leader who declared himself Emperor Agustín I and was later executed. Fernando VII of Spain is the monarchical shadow hanging over the whole independence generation. Their co-occurrence could suggest García Márquez is using Bolívar’s dying consciousness to stage a kind of tribunal of failed liberators: men whose republican ideals curdled into autocracy, exile, or irrelevance. This would reinforce the novel’s central irony: Bolívar, architect of the independence of 6 modern nations, dying in poverty and political disgrace.


This project has read The General in His Labyrinth from a distance — and what that distance reveals is, paradoxically, a novel profoundly preoccupied with closeness. Close quarters on a river boat, close dependence on a mayordomo, the closing in of death. The quantitative methods applied here do not replace the rich understanding a close reader brings to García Márquez’s prose, but they make visible patterns that no close reader could reliably detect across hundreds of pages.

The findings converge on a coherent portrait of the novel. Word frequency established the primacy of José Palacios and his gradual recession as the narrative progresses, suggesting a structural shift from Bolívar’s physical dependency to his psychological deterioration. Sentiment analysis confirmed what readers feel but rarely quantify: the novel grows measurably darker as it advances, with negative sentiment intensifying in the final chapters as Bolívar approaches death. The NRC emotion heatmap added nuance, revealing that anticipation (dark, death-ward anticipation) is among the novel’s most persistent emotional registers.

Topic modeling via LDA proved less illuminating than the other methods, a finding meaningful in itself. The algorithm’s inability to produce clearly distinct topics reflects the novel’s deliberate structural monotony. Its chapters are roughly equal in length, similarly populated, and unified by the same relentless trajectory toward death. Where LDA fell short, network analysis succeeded. The co-occurrence network revealed that Bolívar is narratively bound to the themes of time, night, and dependence on Palacios, while the character correlation network surfaced the unexpected cluster of Fernando, Urdaneta, and Iturbide: three figures of failed or compromised republicanism whose narrative proximity suggests García Márquez is using Bolívar’s dying consciousness to conduct a continent-wide autopsy of the independence generation.

The most significant limitations of this work are its reliance on the English translation and the small corpus size. Eight chapters is a small corpus for methods like LDA that benefit from scale. Working with Edith Grossman’s translation rather than García Márquez’s Spanish means the findings are technically findings about the translated text, a point that warrants serious acknowledgment in any future extension of this work. Applying these methods to the original Spanish, and to García Márquez’s broader corpus, would constitute a meaningful contribution to a field that has so far left one of the world’s most important authors largely untouched.

by Nathan Hertzberg

Notes

  1. Jennifer Schuessler, “Reading by the Numbers: When Big Data Meets Literature,” The New York Times, October 30, 2017. The New York Times.
  2. Steven Mintz, “Why Historical Fiction Matters,” Inside Higher Ed, March 5, 2023. Inside Higher Ed.
  3. Alberto Arvelo, Libertador, Cohen Media Group, 2013.
  4. Gabriel García Márquez, The General in His Labyrinth, trans. Edith Grossman (New York: Alfred A. Knopf, 1989).
  5. Rodica Grigore, “Gabriel García Márquez, History and the Labyrinth of Literature,” Theory in Action 13, no. 4 (October 2020): 107–30. Theory in Action.
  6. Timothy Brennan, “The Digital-Humanities Bust,” The Chronicle of Higher Education, October 15, 2017. The Chronicle of Higher Education.
  7. Schuessler, “Reading by the Numbers.” The New York Times
  8. Emma Belnap, “In Defense of the (Digital) Humanities,” BYU Humanities Center, October 29, 2023. BYU Humanities Center.
  9. Francesco Cauteruccio, “One Hundred Years of Solitude: How I Analyzed My Favorite Book,” 2018. Francesco Cauteruccio on Medium.
  10. Cauteruccio, “One Hundred Years of Solitude: How I Analyzed My Favorite Book.” Francesco Cauteruccio on Medium.
  11. Álvaro Santana-Acuña, Ascent to Glory: How One Hundred Years of Solitude Was Written and Became a Global Classic (New York: Columbia University Press, 2020).
  12. Santana-Acuña, Ascent to Glory: How One Hundred Years of Solitude Was Written and Became a Global Classic.
  13. Grigore, “Gabriel García Márquez, History and the Labyrinth of Literature.” Theory in Action
  14. García Márquez, The General in His Labyrinth.
  15. Julia Silge and David Robinson, “Analyzing Word and Document Frequency: Tf-Idf,” in Text Mining with R: A Tidy Approach (2017). Text Mining with R.
  16. García Márquez, The General in His Labyrinth.
  17. Julia Silge and David Robinson, “Topic Modeling,” in Text Mining with R: A Tidy Approach (2017). Text Mining with R.

Author

  • Nathan is a junior from Hamilton, VA, majoring in statistical science and history. He is one of three Politics Editors at The Lemur.


Discover more from The Lemur: Duke's Big Ideas Magazine

Subscribe to get the latest posts sent to your email.

Recent


Discover more from The Lemur: Duke's Big Ideas Magazine

Subscribe now to keep reading and get access to the full archive.

Continue reading