Repetition, with and without variation, pervades language use. We see it, for example, in spontaneous conversation, in memes on social media, in traditional rhetoric and, of course, in poetry and oral tradition. This paper presents recurrence plotting as a visualisation tool for the analysis of exact and inexact repetition in language. It is a particularly powerful tool for the exploration of un(der)studied data, such as those arising through language documentation work, as its reliance on the visual sense offers instant, holistic insights into the structures of large amounts of data. Recurrence plotting can be used with any type of connected source material, including audio (e.g., full-spectra, f0 contours, etc.) and written data (e.g., transcripts, glosses and translations). We illustrate in this paper the diverse potential of recurrence plotting with data from two oral traditions: the Old Indo-Aryan Rigveda, and the oral tradition Igu of the eastern Himalayan society of the Kera’a.
Repetition, with and without variation, pervades language use. We see it, for example, in spontaneous conversation, in memes on social media, in traditional rhetoric and, of course, in poetry and oral tradition. This paper presents recurrence plotting as a visualisation tool for the analysis of exact and inexact repetition in language. It is a particularly powerful tool for the exploration of un(der)studied data, such as those arising through language documentation work, as its reliance on the visual sense offers instant, holistic insights into the structures of large amounts of data. Recurrence plotting can be used with any type of connected source material, including audio (e.g., full-spectra, f0 contours, etc.) and written data (e.g., transcripts, glosses and translations). We illustrate in this paper the diverse potential of recurrence plotting with data from two oral traditions: the Old Indo-Aryan Rigveda, and the oral tradition Igu of the eastern Himalayan society of the Kera’a.
This paper presents recurrence plotting as a visualisation tool for the analysis of both exact and inexact repetition in language. In contrast to other visualization tools used to highlight repetition (for example, frequency and network analysis tools in stylometrics, see e.g., Päpcke et al., 2023), recurrence plots visualize language use over time. Recurrence plotting has been employed for diverse purposes in the natural sciences, but with limited applications to the study of language - primarily focussing on conversation analysis and speech recognition. The tool has not so far become widely used, or even well-known, in linguistics, despite its remarkable versatility. In this paper we argue that it has, moreover, special potential in revealing linguistic structures in little or un-annotated data.
Repetition, with and without variation, pervades language use (Aitchison, 1994; Brown, 1999; von Contzen et al., 2024). We see it, for example, in spontaneous conversation (Tannen, 1987; Couper-Kuhlen, 2020; Auer & Pfänder, 2007), in memes on social media (Milner, 2016), in traditional rhetoric (Nash, 1989) and, of course, in poetry and oral tradition (Jakobson, 1966; Fox, ed., 1988; Gaenszle, 2018). Recurrence plots visualise repetition patterns in any data type and in accordance with any given unit of comparison: a fixed time-based unit (e.g., a millisecond), a prosodic or written word, a line in poetically structured language, a sentence, a paragraph etc. In standard recurrence plotting, each chunked unit is compared with each other unit in the same data set. These comparisons reveal repetitions within a text or recording. In cross-recurrence plotting, units from different data sets are compared to reveal similarities and differences between, e.g., different texts or audio recordings. In either kind of recurrence plotting, a selected colour palette encodes levels of similarity. In this paper, we demonstrate the use of two-colour as well as multi-colour palettes, reflecting binary or gradient levels of similarity.
While repetition is ubiquitous in language, it is within the ritualised structures of oral art traditions or traditional poetry where it becomes a defining principle (e.g. Jakobson, 1966). Therefore, after introductory examples from English, we choose two unrelated oral art traditions to illustrate the diverse potentials of recurrence plotting. Note that these only serve as illustrations, as the main thrust of this paper is the demonstration of the versatility of recurrence plotting for linguistic exploration and analysis. The first oral tradition that we choose for illustration is the Rigveda (ca. 1300 BC), one of the best-known as well as better-studied oral artworks in the world. As such, it can serve as evidence for the utility of recurrence plots. The other oral tradition is Igu, the shamanic language of the Kera’a, a Tibeto-Burman speaking society of the eastern Himalayas (Reinöhl, 2022). While Igu rituals are similar in breadth and richness to those of the Vedic tradition, research on their unique language (Reinöhl et al., 2024; Reinöhl, 2026) and their cultural role (Dele/DelleyFootnote 1, 2018; 2021; 2023) is only beginning. In un(der)studied language data like that collected for Igu, recurrence plotting can serve as an important exploratory tool. It can reveal repetition patterns even before a detailed linguistic analysis.
Section 2 provides a literature survey of recurrence plotting applications. Section 3 gives some more background regarding the role of repetition in language use and on the two oral traditions, the Rigveda and Igu, chosen here for illustration. The subsequent methods section introduces the notion of time series as the foundational concept underlying recurrence plotting as well as more technical background. This section also offers guidance in how to read recurrence plots. Section 5 describes diverse applications of recurrence plotting, highlighting repetition in full-spectrum acoustic signals, orthographic representations, interlinearizations, formulas (with multi-word formulas being ubiquitous in many oral art forms), and metrical structure. This section ends with an example of cross-recurrence plotting, highlighting repetition between rather than within data-sets. Section 6 concludes.
The term recurrence plot was introduced by Eckmann et al. (1987). They describe a method for the visualisation of patterns of repetition in dynamical systems. Zbilut et al. (1990) and Webber and Zbilut (1994) applied the methodology to the exploration of dynamic systems in physiology, such as heartbeat and breathing respectively. Applications to the social sciences also start early. Scheinkman & LeBaron (1989) use a number of recurrence plots to look at patterns in US GNP time-series data. The most cited article in the field is a technical discussion of recurrence plotting in general by Marwan et al. (2007).
Helfman (1994) offers an early example of recurrence plotting (albeit not under that name) of language data. His work includes a study on the complete works of Shakespeare, looking at word repetitions. The study also examines recurrence patterns in programming language code and natural language software manuals translated into multiple languages.
Another early application of recurrence plots to language data is that by Orsucci et al. (1997), who apply them to the comparison of poems in original language versions (in English and Swedish orthographies) and their ‘translations’ (Italian orthography and audio, English audio)Footnote 2. Dale and Spivey (2006) explored recurrence relationships in a corpus of child language data, highlighting how children pick up expressions used by adults. Angus et al. (2012) step from processing language at the purely form-based level to looking at conceptual repetition. They use a cosine metric between cooccurrence vectors as their distance measure to explore how the topics discussed in a conversation evolve over time. Other aspects of language have also been addressed with this methodology. Schultz (2006), for example, reports on a meeting of scholars discussing the use of recurrence plots to analyse speech. Angus (2019) offers a survey of applications of recurrence plots to communication studies. Overall, however, there has so far been more interest in recurrence plots in the fields of communication research than in exploring linguistic structures themselves.
It is worth noting that the same technique has appeared within digital humanities under another name. Foote (1999) (see also Foote & Cooper, 2001, Cooper & Foote 2003) looked at “self-similarity matrices” in music. Their methods and resulting visualisations are formally identical with those in the “recurrence plot” tradition.
In contrast to the above previous work, in this paper we focus on how recurrence plots can assist a linguist in both their first foray in the structure of relatively unstudied works, as well as revealing hitherto unrecognized structures in better-studied data.
Humans employ repetition at scales ranging from the smallest units of language to the largest, e.g., from phonemes or syllables, to schematic discourse structures such as those found in storytelling (see Contzen et al., 2024 for a recent overview). It occurs with a wide variety of functions, e.g., for emphasis, correction, humour, or sarcasm. Its use can build community through shared linguistic practices. Repetition also features in diverse grammatical domains, such as in morphological reduplication for the marking of a plural (as in Indonesian) or intensive meaning (as in Sanskrit). Aitchison (1994) offers the following enumeration of repetition phenomena: “Alliteration, anadiplosis, antimetabole, assonance, battology, chiming, cohesion, copying, doubling, echolalia, epizeuxis, gemination, imitation, iteration, parallelism, parroting, perseveration, ploce, polyptoton, reduplication, reinforcement, reiteration, rhyme, ritual, shadowing, stammering, stuttering.” In this paper we focus on repetition, or more broadly similarity relations, in connected text or speech. The interplay between degrees of similarity and dissimilarity is fundamental to the organisation of any natural language use. This is true whether or not repetition is frequent, preferred or accepted, or whether high-similarity repetitions are avoided. Many genres can be characterised by their preferred types and densities of repetition - however, there is one genre where repetition has been taken as defining: oral art. Oral art designates orally transmitted, poetically structured discourses of often extensive length, which are ubiquitous in societies that do not practice writing or only relatively recently introduced writing practices. This also includes types of language that are orally rooted, even if they are eventually written down and passed on through writing (examples include the Hebrew Bible and Homer’s epics).
After illustrating recurrence mapping with simply structured examples from English (Sect. 4.2), we will illustrate its potentials with data from Vedic Sanskrit and Igu, which were chosen based on the first author’s direct research experience with them. The Vedic Sanskrit data source is the Rigveda, the oldest attested Indo-Aryan hymn collection and the foundational text of Hinduism. The Rigveda has been studied intensively with regard to its written form, even though it was created orally before the advent of writing in South Asia and continues to be orally transmitted to the present day. We will illustrate recurrence plotting with both written and spoken data. The audio data comes from recordings of recitations performed by D. P. Kinjawadekar in 1983, which have recently become available through the collection of Guni Hesting Kirchheiner. The recitations are a unique resource, giving insight into one of the longest-living, though endangered, oral traditions: the Rigveda is estimated to have been assembled in its current form by the 13th century BC. All Rigveda data are made available through “VedaWeb 2.0” (Kiss et al., 2019; Casaretto et al., 2023; Kölligan et al., n.d.).
The second data source comes from Igu, the shamanic language of the Kera’a people, a society of the eastern Himalayas whose languages belong to the Trans-Himalayan family (a.k.a. Sino-Tibetan or Tibeto-Burman; Reinöhl et al., 2025; Reinöhl, 2024; Reinöhl, 2026). The Kera’a term “Igu” designates both the shamanic beliefs and practices of the Kera’a, and the shaman him- or herself. It is also used as the term for the language of the shamanic rituals. The language Igu shares about half of its vocabulary with Kera’a (i.e., the language of everyday use among the Kera’a people), and the two varieties also show substantial morpho-syntactic similarities. Igu is, however, sufficiently different in both lexicon and grammar to support classification as a distinct language from Kera’a, matching how it is seen within Kera’a society. Igu is endangered, as few young Kera’a are becoming shamans today. However, both the language and the shamanic tradition remain comparatively vibrant. Compared to several other Eastern-Himalayan societies, the Kera’a have been less influenced by Christianity and/or Hinduism, and shamanic rituals are still regularly performed. No linguistic or cultural influences from these other now-dominant belief systems in the region on Igu can be discerned at this point (Reinöhl et al., 2025). The data for this paper come from three performances of two different healing rituals. The shorter ritual, the Kãliwu, lasts between 20 min and an hour. Both performances used here were conducted by the same shaman, Igu Pachu Pulu. The longer ritual, the Ayi, lasts approximately 7 h, and was performed by Igu Mola Mili. All three performances were recorded in Lower Dibang Valley, India, in August 2019. The recordings, with interlinear glosses and translations (complete for the two Kãliwu rituals, and available for the first 30 min of the Ayi ritual), are archived with Language Archive Cologne (Reinöhl, 2024). All Igu performance recordings are also made available to the Kera’a society in a linguistically-enriched yet accessible form, as video recordings with subtitled transcripts and free translations.
In this section, we introduce the notion of time series in linguistics and how these can be used to create recurrence plots. We will start with a pair of examples of English data, before moving on to a more theoretical account of the relationship between recurrence plotting and time series.
4.1 Time series and recurrence plotsRecurrence plotting is based on the concept of time series. A time series is any sequence of data items where each data item is associated with a time. The time units can be extrinsic, measured for example in milliseconds or seconds. Alternatively, they can be intrinsic, measured in ordinal numbers just showing the temporal ordering of the items in the sequence (e.g., letters, intonation units or lines).
Many kinds of linguistic information can be thought of as forming a time series. An audio recording, for example, is a sequence of recorded amplitudes, occurring at evenly spaced intervals (the sample rate). Video recordings encode frames (pictures) at regular time intervals – also constituting a time series. More symbolic forms of language rarely have extrinsic markings reflecting objective time (exceptions are diaries with dated entries), but rather the passage in time is defined internally by relative temporal encoding. This may be no more than a relative ordering of the expressed items, with no marking of temporal quantity. For example, the characters that make up this paper form a time-series, as do its words, sentences or paragraphs, with the first preceding the second, and the second preceding the third, and so on.
In mathematical terms, we can describe a time-series as a series of values (xt)t∈T. This expression says that there is a set of time values T (which may be absolute timestamps, or relative temporal positions), and associated with each of these is a value xt.
A recurrence plot takes two time series and compares them. It does this by defining a comparison function which produces a colour to indicate how similar two items, one from each time series, are. If the time series are laid out along the x and y axes of the plot, then we can relate any point on the plot to a pair of items (xi,yj), where xi is the ith time-point in the time series Tx along the x axis (a whole vertical line of points will have the same x value), and yj is the jth item in the time series Ty along the y axis (likewise, horizontal lines will all share the same y value). The defining property of a recurrence plot is that we paint the point corresponding to (xi,yj) with a colour that reflects how similar xi is to yj.
The kind of similarity measure we use depends on the nature of the items in the time series. For example, if we are looking at a text letter-by-letter, a good comparison to use is identity: letters are either identical (similarity = 1) or different (similarity = 0). We might mark identity with a dark colour, and difference with a light one (as seen in Sect. 4.2 and 5). If our items are words, we might use a scaled Levenshtein distance (Levenshtein, 1966) to again achieve a similarity measure between 0 and 1, but capable of many values between those extremes. This similarity measure reflects how many characters occur in both words being compared, and in the same order.
For acoustic data, we might measure the similarity of two F0 (fundamental frequency) values as some non-linear (e.g. logarithmic) scaling of the difference between them. Spectral slices can be compared using cosine similarity measures. In all cases, the colour at each point in the plot reflects the similarity of the corresponding two points in the input time-series.
The software for creating these plots was created by the second author in the PYTHON programming language (version 3.12, Python Software Foundation, https://www.python.org/), using the jupyter software development environment (https://docs.jupyter.org/en/latest/). It includes a functional programming library tmefunc5, a library for manipulating time-series tmeseries, and a library tmerecmap for creating recurrence plots. The software is available in the OSF repository. The workflow for creating the plots consisted of the following steps: (a) organising the data usually in a spreadsheet, saved as a csv file, (b) reading the data into the script, (c) performing any needed preprocessing of the data, (d) breaking the data into components of the desired granularity in a time-series, (e) specifying the distance measures and colorisation, and (f) calling the library function to create the recurrence plot. Steps (b)-(f) occur within scripts created to process each specific data-set.
4.2 Recurrence plots: how to read themRecurrence plotting allows the self- or other-comparison (“cross-recurrence plotting”) of any kind of sequenced linguistic data. It doesn’t matter if it is written data, oral or visual, as in sign language or gesture analysis. And it can be used at any linguistic level (e.g., for phonetics, morpho-syntax or discourse structure). A recurrence plot is defined by the data type being plotted (e.g., the audio spectrum), the unit of comparison (e.g. 10ms samples) and the type of comparison (e.g., cosine similarity between vectors). We can define two primary classes of comparison: item identity and item similarity. We begin our examples with a simply structured recurrence plot, doing identity comparison of the lines of a poem.
This first example is a recurrence plot of the Litany of the Blessed Virgin Mary. This Christian prayer shows highly repetitive line structure. Let us, to start with, focus on the first four lines only. These are:
(1)
Lord, have mercy on us.
Christ, have mercy on us.
Lord, have mercy on us.
Christ, hear us.
In terms of line identity, we here see an a-b-a-c pattern, i.e. identity of the first and third line (Lord, have mercy on us.) and different lines in the second and fourth line. Now consider the recurrence plot (plot 1), for which we have chosen a two-tone colour scheme, with line identity in dark red and non-identity in cream. In the lower left corner, we see a small chessboard pattern of first a red dot, then a cream dot, then red, then cream. The very first dot compares the first line with itself, as the whole prayer is plotted, identically, onto the x- and y-axes. Now let’s consider the next dot next to the first one - in cream (Given that the prayer is plotted onto both axes, it does not matter whether we shift right on the x-axis or up on the y-axis.). This dot indicates that the first and second line (plotted onto x = 1 and y = 2, or, alternatively, x = 2 and y = 1) are non-identical. Moving onto the third dot (again, shift, e.g., to the right on the lowest line), we see another red dot. This red dot indicates that the first line is identical to the third line (either, x = 1, y = 3 or x = 3, y = 1), as we saw in ex. (1) above. Repetition structures abound throughout the prayer, with the line pray for us alternating with other invocations. The recurrence plot allows us to easily see the structure of the whole prayer at a glance, including the large uniform chessboard-like section, and the distinct patterns occurring initially and finally. When we see square structures of the kind in plot 1, we know that some item is being repeated frequently in this section of the time-series. The edges of the square show the boundary times of this interval of local repetition. In this case, we see four repetition regions around the main diagonal: a minimal repetition at the start of the prayer, then a short region of repetition, then a large region where pray for us is repeated regularly, and finally another region where a line is repeated three times. This structure is intuitively accessible in the plot.
Recurrence plot showing a line-identity comparison of the Litany of the Blessed Virgin Mary
Our second example explores William Blake’s well-known poem The Tyger. We could have taken the original text as input, but given the complex grapheme-to-phoneme correspondence and the importance of rhyme and other sound play, we decided to work with a phonological transcription instead. For this, we used ChatGPT to produce a phonological transcription appropriate for the historical era of Blake’s writing, in the late 18th century. (For the full prompts and responses, see the Appendix.) On the basis of the transcription returned by the AI (after minor manual correction), we encode phoneme identity and non-identity. We once again work with binary colour coding, using the same colour scheme as in the previous plot (red for identity, cream for non-identity). Among the many repetitions, what stands out are the two lines parallel to the diagonal in the upper left and lower right corner, marked with blue boxes. This pattern arises from the near-identity of the poem’s first and last verse: Tyger Tyger burning bright, In the forests of the night: What immortal hand or eye, Could/Dare frame thy fearful symmetry? In other words, this verse starts at x = 1, y = 1 extending to x = 85 and y = 85, and it also starts at x = 450, y = 450, extending to x = 534, y = 534. This full-verse repetition is visually revealed by the off-set shorter diagonals, starting at x = 0, y = 450 and x = 450, y = 0, respectively. The granularity of this plot is much finer than that of the previous plot, as we are comparing each phoneme with each other phoneme, rather than entire line units with each other as in plot 1. Nonetheless, full-line near-identity emerges from the identity of almost all phonemes in the two sections (verses) at the beginning and end. Also note the square pattern around the diagonal in the middle of the poem (marked with a red box). This pattern reflects the consistent repetition of the word “what” along with shorter repetitions of other words (plot 2).
Recurrence plot of a phonemic transcription of William Blake’s The Tyger. The units are phonemes, and identity is shown in dark red, non-identity in cream. Given the rich and detailed structure, readers of electronic copies are encouraged to zoom in on this picture
We have selected an array of diverse data and recurrence plotting applications in order to illustrate the flexibility and breadth of this tool for linguistic analysis. We illustrate applications to the auditory signal (Sect. 5.1), the transcript (5.2) and the glossing line (5.3). We show both high-granularity applications, zooming in to the realization of single words and formulaic couplets (5.4) as well as of entire ritual performances (5.5). This selection illustrates uses for phonetic or phonological analysis, synonymy and formula research, and morpho-syntactic corpus analysis. While the examples chosen are necessarily selective, we hope that they create in the reader an understanding of the essentially unlimited applicability of recurrence maps to connected linguistic data.
5.1 Sound parallelism through a full-spectrum analysisPlot 3 shows a single spoken Igu couplet compared with itself. The similarity measure compares full spectra with each other, using rainbow colours to show the degree of similarity ranging from red (most similar) to violet (least similar). This plot highlights the sound-based parallelism within the couplet provided, in transcription, in (2)Footnote 3. The sound file can be accessed in this project’s OSF archive (see also Reinöhl, 2026, for poetic structures and strategies in Igu):
(2) | aloju | atogi | ji-mi-ma | |
|---|---|---|---|---|
holy_turmeric | tong | sit-neg-aff | ||
‘(Igu says: I have) Turmeric! Tongs! Stay away! | ||||
abrapo | alone | ji-mi-ma | ||
surname_of_turmeric | short_tong | sit-neg-aff | ||
(Igu says: I have ) Turmeric! Short tongs! Stay away!’ | [ayi_a_8:51] | |||
Recurrence plot of a full-spectrum analysis of an Igu couplet. Units of comparison are frequency energy spectra of short fixed intervals. These spectra, expressed as vectors, are compared with a cosine similarity measure. The colours follow a rainbow spectrum from red (high similarity) through green to blue and violet (low similarity)
What is important here is that we processed frequency energy spectra of the raw acoustic information, rather than the transcript. A transcript is, by its nature, greatly reduced and standardised, relative to the much richer information available in an audio record of a performance. Nonetheless, the three-word line structure, (near-)identity of the line-final verbs, and phonological similarities (same syllable count, alliteration, near-identical phonotactic structure) of the synonymous nouns emerge from the recurrence plot. We superimposed two black boxes in plot 3 to frame the couplet’s two lines, while the white box delimits the pause between the lines.
This application thus illustrates how we can detect couplet structure, using analyses available even prior to the creation of a transcript. In other words, recurrence plotting can help as a short-cut to explore the internal organisation of un(der)studied material even before any additional work has been carried out beyond the recording of the data.
5.2 Sound parallelism through letter analysisWhen the next step of analysis has taken place - a transcript has been produced - a plot can be generated parallel to the one visualised in the previous section. The new plot shows the binary distinction of letter identity/non-identity (again, in dark red and cream). Since the data sources, and the mode of comparison are different, the graphs do not look the same, but similar structures clearly exist, and are apparent in the two graphs. The graphs are reproduced side-by-side in plot 4. There is no analogue to the silence in the audio recording between chanted lines, and so the squares for the two lines touch in the plot on the left.
Recurrence plot of letter comparison in an Igu couplet. For comparison, plot 3 is reproduced on the right
An important feature of parallelism is the repetition of parallel - i.e. similar, but not identical - terms, in analogous lines (e.g., in the two lines of a couplet). Parallel terms often stand in a semantic relationship of synonymy or antonymy, with different similarity degrees (Jakobson, 1966).
In order to capture semantic similarity one could use a qualitative ontology for the domain being considered, for instance based on the Gold Upper Ontology (Farrar, 2003), or a quantitative semantic distance measure based on Large Language Models (dos Santos & Leal, 2024). In the former case, the distance measure might count the number of stepping stone concepts needed to move from meaning to another within the ontology, e.g. cat - pet - animal - reptile - lizard might give a distance of 4 between cat and lizard.
For languages without the LLM option, or their own qualitative ontology, we can approximate a basic semantic similarity measure by letter-based comparison of glossed meanings. This will capture terms with identical glosses, or glosses with shared morphology. This approach to parallel term analysis is illustrated with the same Igu couplet used in the two previous sections, which we repeat here for convenience:
(3) | aloju | atogi | ji-mi-ma | ||||
|---|---|---|---|---|---|---|---|
holy_turmeric | tong | sit-neg-pol | |||||
‘(Igu says: I have) Turmeric! Tongs! Stay away! | |||||||
abrapo | alone | ji-mi-ma | |||||
surname_of_turmeric | short_tong | sit-neg-aff | |||||
(Igu says: I have ) Turmeric! Short tongs! Stay away!’ | [ayi_a_8:51] | ||||||
This couplet contains two synonymous parallel term pairs: aloju/abrapo for ‘turmeric’ and atogi/alone for ‘(short) tong(s)’. Their glosses are partially identical, i.e. ‘holy_turmeric’/’surname_of_turmeric’Footnote 4 and ‘tong’/’short_tong’. Accordingly, a letter-based comparison results in partial identity, as indicated by the four red-framed boxes in Plot 5. This letter-based comparison roughly mimics semantic mark-up, where e.g. ‘tongs’ and ‘short tong’ are related to each other in something like a generic-specific relation.
Recurrence plot of ex. (3). Note the superimposed red-boxed areas containing short lines parallel to the diagonal. These reflect the semantic parallelism between the strings "_turmeric" and "tong sit-neg-aff" respectively
Our approach to synonym detection here resembles the formula detection for the Rigveda described in Sect. 5.4. However, while plot 5 is based on per-letter comparison, plot 6 shows the identity or non-identity of full word forms.
5.4 Formula detectionFormulas as “repeated, semantically unified word-group[s]” (Dunkel, 2021: 13) are particularly prominent types of repetition structures in many oral art traditions (Parry, 1928), including in the Rigveda (Bloomfield, 1916; Dunkel, 2021). In formulas, the first word “created a strong presumption that the other would follow” (Hainsworth, 1962: 57-68). Thus they act as entrenched combinations that can be considered holistic gestalt-like units.
In the Rigveda, inflection-based variation of formulas is common (Dunkel, 2021: 15–16). We thus choose to disregard inflections and use only the citation forms of words while preserving the order in which the lexemes occur in the Rigveda. The lemmas have been derived from the VedaWeb database (Casaretto et al., 2023).Footnote 5
The recurrence patterns characteristic of short formulas can be seen in plot 6 (framed by a blue square). The visual structure to look out for are short lines parallel to the main diagonal (similar to our previous The Tyger example). These smaller, offset diagonals indicate the use of the same multi-word group at different points in the text. For the plot, we selected 500 lemmas between RV 1.14 and 1.20. Dark red indicates the identity of two word lemmas, while cream indicates non-identity. In this plot, even the difference in one letter results in the classification (and visualization) as a different lemma - and so we mark these in the plot as distinct (cream). You can see that diagonals abound especially in the upper right corner (see superimposed blue-framed area). This region is a part of RV 1.19 where the last pāda (metrical ‘foot’) in each stanza contains the following identical lemmas in this sequence: marút- (‘Marut’, ‘storm god’), agní- (‘Agni’, ‘fire’, ‘god of fire’), ā́ (‘to, near, towards’), √gam- (‘to go’). The first two stanzas (given as (4) and (5) respectively below, with lemmatised glosses and translationsFootnote 6) serve to illustrate (in bold print) the pattern of stanza-final formula repetition in this hymn.
Word-based comparison in the Rigveda. The grid pattern highlighted by the red box shows the frequent regular repetition of the lexeme Indra. The blue-highlighted region shows the frequent repetition of the short formula With the Maruts, o Agni, come hither
(4) | práti | tyám | cā́rum | adhvarám |
práti | syá- ~ tyá-.acc.sg.m | cā́ru-.acc.sg.m | adhvará-.acc.sg.m |
gopīthā́ya | prá | hūyase |
|---|---|---|
gopīthá-.dat.sg.m | prá | √hū-2sg.prs.ind.pass |
marúdbhiḥ | agne | ā́ | gahi | |
|---|---|---|---|---|
marút-.ins.pl.m | agní-.voc.sg.m | ā́ | √gam-.2SG.AOR.IMP.ACT | |
‘Toward this pleasing ceremony you are called, for its protection. – With the Maruts, o Agni, come hither.’ | (RV 1.19.1) | |||
(5) | nahí | deváḥ | ná | mártyaḥ |
nahí | devá-.nom.sg.m | ná | mártya-.nom.sg.m |
maháḥ | táva | krátum | parás |
|---|---|---|---|
máh-.gen.sg.m | tvám.gen.sg | krátu-.acc.sg.m | parás |
marúdbhiḥ | agne | ā́ | gahi | |
|---|---|---|---|---|
marút-.ins.pl.m | agní-.voc.sg.m | ā́ | √gam-.2sg.aor.imp.act | |
‘For no god nor mortal is beyond the will of you who are great. – With the Maruts, o Agni, come hither.’ | (RV 1.19.2) | |||
Another structural type stands out in plot (6)Footnote 7. Matrices of dots or short line-segments indicate identical, contiguous (if the red dots are connected) or non-contiguous (if separated by white) sequences repeating the same lemma or short lemma sequence. An example comes from RV 1.16.2-4, where the god Indra is named repeatedly with only a few words separating the repetitions. The lemma-based recurrence plot highlights these mentions regardless of inflectional differences (see the superimposed red-framed area in plot 6). Here is the text that gives rise to the framed pattern.
(6) | imā́ḥ | dhānā́ḥ | ghr̥tasnúvaḥ |
ayám.acc.pl.f | dhānā́-.acc.pl.f | ghr̥tasnú-.acc.pl.f |
hárī | ihá | úpa | vakṣataḥ |
|---|---|---|---|
hári-.nom.du.m | ihá | úpa | √vah-.3du.aor.sbjv.act |
índram | sukhátame | ráthe |
|---|---|---|
índra-.acc.sg.m | sukhátama-.loc.sg.m | rátha-.loc.sg.m |
índram | prātár | havāmahe |
|---|---|---|
índra-.acc.sg.m | prātár | √hū-.1pl.prs.ind.mid |
índram | prayatí | adhvaré |
|---|---|---|
índra-.acc.sg.m | √i-loc.sg.m.prs.act | adhvará-.loc.sg.m |
índram | sómasya | pītáye |
|---|---|---|
índra-.acc.sg.m | sóma-.gen.sg.m | pītí-.dat.sg.f |
úpa | naḥ | sutám | ā́ | gahi |
|---|---|---|---|---|
úpa | ahám.acc/dat/gen.pl | √su-.acc.sg.m | ā́ | √gam-.2sg.aor.imp.act |
háribhiḥ | indra | keśíbhiḥ |
|---|---|---|
hári-.ins.pl.m | índra-.voc.sg.m | keśín-.ins.pl.m |
suté | hí | tvā | hávāmahe | ||
|---|---|---|---|---|---|
√su-.loc.sg.m | hí | tvám.acc.sg | √hū-.1pl.prs.ind.mid | ||
‘[16.2] Here are the roasted grains, bathing in ghee; the fallow bay pair will convey Indra here right to them in the best-naved chariot. [16.3] Indra we invoke early in the morning, Indra as the ceremony advances, Indra to drink of the soma. [16.4] Come up here to our pressed soma, Indra, with your shaggy fallow bays, for when it is pressed we invoke you.’ | [RV 1.16.2-4; highlighting added] | ||||
We now turn to how recurrence plots can help us visualise metrical structure as another repetition-based feature characteristic of many oral art forms. The Rigveda is organised into books, which are divided into hymns, which in turn consist of a number of stanzas each. Generally, the metrical pattern of a hymn is maintained in all of its stanzas. In order to bring out such metrical differences, we selected two consecutive hymns with distinct metres. Hymn RV 1.9 is composed in the Gāyatrī metre, which consists of three parts (pādas), each of which contains eight syllables. The next hymn, RV 1.10, uses the Aṇuṣṭubh metre, wherein stanzas have four parts, each of eight syllables.
The difference in stanza lengths between the two metrical types is reflected in distinct breath unit patterns. The pattern emerging from the visualization is that the recitation involves a pause at the end of two pādas since the last pause, or at the end of the stanza, whichever comes first. Plot 7 shows part of RV 1.9, highlighting differences in acoustic intensity. For easier recognition, we framed each stanza in black. Inside each stanza, the reverse-colour lines (yellow when crossing a red area or red when crossing another yellow line) reflect pauses, because silence is very different in intensity to speech. Within each stanza we see a strong break somewhere around two thirds of the way through. We also see that one breath pause is longer, and the audio signal suggests that this is for recovery purposes, as the recitation is delivered at a certain speed.
A recurrence plot of part of hymn RV 1.9, based on acoustic data. Red indicates similarity in intensity, while yellow indicates difference in intensity. The black boxes frame individual stanzas. Note the vertical (or horizontal) yellow lines approximately ⅔ of the way from the left (or bottom) of each box to the right (or top)
A different picture is visible in plot 8, where we see not only verses in the Gāyatrī metre (lime-boxed stanzas), but some in the Aṇuṣṭubh metre (blue-boxed stanzas) as well. In the latter, we see the line marking a pause as a strong bisector of the region, occurring very close to the middle of the stanza, in contrast to the two-thirds position in stanzas of the Gāyatrī metre. This is because the breath-delimited units are of relatively equal lengths in the Aṇuṣṭubh stanzas, consisting of two pādas each.
Recurrence plot of intensity across hymns RV 1.9 and 1.10. Notice the shorter stanzas of the Gāyatrī 2+1 pāda pattern in RV 1.9 (yellow boxes) in comparison to 2+2 pāda pattern of stanzas (blue boxes) in RV 1.10
The audio quality of the Rigvedic recitational data is good considering that the chants were recorded in the 1980s. With a transcript at hand, the recitation can be followed along, especially if one is somewhat acquainted with the Rigveda and Vedic Sankrit. (The speed of delivery, however, will require appropriate focus.) However, there is background noise throughout, and the audio quality is far from current acoustic standards in phonetic or phonological fieldwork for language description. The interested reader can listen to the audio files in our OSF archive. The clarity with which the metrical differences emerge, notwithstanding the less-than-optimal audio quality, is noteworthy. It shows the potential of recurrence plotting for exploring audio data even when state-of-the-art recordings are not available.
5.6 Cross-recurrence plotsSo far we have only used recurrence plotting for self-comparison. However, it is possible to use it also for other-comparison - “cross-recurrence plotting” - where two different data strings are compared. For example, we can compare one performance of a ritual with another performance of the same ritual. We illustrate this use by comparing the full transcripts of two Kãliwu performances by Igu Pachu Pulu, both recorded on the same day in August 2019, but for different clients.Footnote 8
The visual patterns show the structure of the performances in two ways. Firstly, every (dark) dot indicates a word shared between the two performances. We see that the two performances share quite a few words, shown by the many dots. At the same time, there is also a lot of difference, indicated by the cream-coloured areas (Plot 9).
Plot 9 shows the comparison of two data sets: Both performances took place on the 4th of August 2019. The shaman performing the chant is the same in both, but the clients, and their ailments, differ. This difference explains the slight difference in length, as the recitation is adapted to each client’s clan lineage and their ailment. The performance on the x-axis is slightly longer with a total of 795 words, in contrast to the 550 words of the other performance
The dark dots form patterns of vertical lines, and to a lesser extent, horizontal ones. This results, among other things, from the fact that 17 word tokens in the longer performance are instances of the word b(w)ea, ‘earlier, in mythological time’, and this word occurs 23 times in the shorter performance, resulting in the associated dark dots being closer together vertically. The higher density of dark dots along the vertical lines corresponding to this word makes them more visually salient than the horizontal sequence of dots for this word. Secondly, we also see several small diagonals. These indicate not just the overlap in individual words, but in entire phrases or formulas (similar to the previous Rigveda example). The most frequent recurrence is the repetition of the three-word, line-initial formula bea echa go ‘long ago, once upon a time’, situating us in mythological, primordial times. We see a number of repetitions of this formula as micro-diagonals spaced out at the bottom of the plot.Footnote 9
The differences in relative internal repetitiveness can, of course, also be brought out by comparing self-comparison plots of the two performances. These self-recurrence plots would capture repetitions even of words not shared between the two performances.
An important insight gained from plot 9 is the high number of words that are not, in fact, shared between the two performances. As mentioned, we see this in the large number of rows and columns which contain no dark dots at all. This means that even the same Kãliwu ritual, performed by the same shaman, on the very same day, can be realised by two performances that are quite distinct. This finding is significant in light of the fact that much research into ritual language relies only on one documented instance of a ritual. Of course, reasons why many scholars have focussed on only single versions of rituals include the lack of other recorded performances and the great effort needed to compare different versions should they exist. Regarding the latter, recurrence plots can decrease workload dramatically by exploiting our visual pattern-recognition capacity to survey substantial amounts of linguistic data at a glance.
This study has illustrated the range of potential applications of recurrence plotting in exploring the structure of oral and written language. While exemplified with data from different verbal art traditions - Vedic Sanskrit and Igu - recurrence plotting has potential uses for any type of connected linguistic data. In order to showcase the breadth of applicational options, we illustrated recurrence plotting with acoustic and written data in the object language, as well as translational information in glosses. We showed uses both for longer stretches of language as well as for local, fine-grained analysis. We also illustrated use-cases at different levels of linguistic structure addressing phonological, metrical, semantic and lexical research questions. We highlighted that this tool has great potential for un(der)explored data, where it can serve as a short-cut to visualise structure prior to the creation of transcriptions or translations, or where data have not been extensively studied. For existing analyses, too, recurrence plots offer visualizations that facilitate data recognition and representation.
The data files and code for this paper are made available in this OSF repository: https://osf.io/w4ek6/.
Spellings of this author’s surname vary across publications.
Unfortunately, we have not been able to find the plots of the recurrence maps associated with the arxiv paper, as these appear to have been hosted externally on sites that are no longer accessible.
abbreviations used in this paper are: ACC=accusative case, ACT=active voice, AOR=aorist, DAT=dative case, DU=dual, F=feminine, GEN=genitive case, IMP=imperative, IND=indicative, INS=instrumental case, LOC=locative case, M=masculine, MID=middle voice, NEG=negation, NOM=nominative case, PASS=passive voice, PL=plural, POL=politeness marker, PRS=present tense, SBJV=subjunctive mood, SG=singular, VOC=vocative case, 1 = 1st person, 2 = 2nd person, 3 = 3rd person.
The use of ‘surname’ here reflects the Igu Pachu Pulu’s own assessment of parallel terms, which he explains as relating to each other like first name and surname. It is an apt terminological choice, and so U.R. has decided to retain the label of ‘surname’ for the non-basic members of parallel term pairs or sets (see also Reinöhl,).
Extraction date, 20.09.2024.
Translations of the Rigvedic verses quoted in this paper are from Jamison and Brereton (2014).
The local particle prá, here prefixed to the verb root √i, is not glossed separately in nominalized forms in the VedaWeb lemmatization, given its bound status.
The Kãliwu ritual contains non-verbal sections of ritualised purification by blowing, waving and sucking actions. The transcripts compared for plot 9 only involve the verbal parts of the two performances.
Our other-comparison in this section illustrates recurrence plotting potentials on transcripts that have not been strongly cleaned. We did not overly standardize the transcripts on purpose (e.g., allowing for stylized variation in pronunciation) in order to show what recurrence plotting can do with even rough inputs. In more refined analyses, one might choose to compare not just words, but morphemes (which would increase repetitiveness in our Kãliwu transcripts, with compounding and derivation being ubiquitous). Another refinement would involve a differentiation between content and function words, where one might only wish to target content words in the recurrence plotting visualization.
Aitchison, J. (1994). Say, say it again, Sam: The treatment of repetition in linguistics. Swiss Papers in English Language and Literature, 7, 15–34.
Angus, D. (2019). Recurrence Methods for Communication Data, Reflecting on 20 Years of Progress. Front Appl Math Stat, 5, 54. https://doi.org/10.3389/fams.2019.00054
Angus, D., Smith, A., & Wiles, J. (2012). Conceptual recurrence plots: revealing patterns in human discourse. Ieee Transactions On Visualization And Computer Graphics, 18, 988–997. https://doi.org/10.1109/TVCG.2011.100
Auer, P., & Pfänder, S. (2007). Multiple retractions in spoken French and spoken German: A contrastive study in oral performance styles. Cahiers de Praxématique, 48, 57–84. https://doi.org/10.4000/praxematique.758
Bloomfield, M. (1916). Rig-Veda Repetitions: The repeated verses and distichs and stanzas of the Rig-Veda in systematic presentation and with critical discussion (2 vols.). Harvard University Press.
Brown, P. (1999). Repetition. Journal of Linguistic Anthropology, 9(1/2), 223–226. https://doi.org/10.1525/jlin.1999.9.1-2.223
Casaretto, A., Halfmann, J., Korobzow, N., Kölligan, D., & Reinöhl, U. (2023). The morphologically glossed Rigveda — The Zurich annotation corpus revised and extended. VedaWeb – Online Research Platform for Old Indic Texts University of Cologne. https://doi.org/10.5281/zenodo.8410655
Cooper, M., & Foote, J. (2003). Summarizing popular music via structural similarity analysis. 2003 IEEE workshop on applications of signal processing to audio and acoustics (IEEE Cat. No. 03TH8684), 127–130. IEEE
Couper-Kuhlen, E. (2020). The prosody of other-repetition in British and North American English. Language in Society, 49(4), 521–552. https://doi.org/10.1017/S004740452000024X
Dale, R., & Spivey, M. J. (2006). Unraveling the dyad: Using recurrence analysis to explore patterns of syntactic coordination between children and caregivers in conversation. Language Learning, 56(3), 391–430.
Dele [Delley], R. (2018). Idu Mishmi shamanic funeral ritual Ya: A research publication on shamanic oral chants and rituals related to mortuary behaviour among the Idu Mishmis of Arunachal Pradesh. Bookwell.
Delley [Dele], R. (2021). Igu: Study on four Idu Mishmi shamanic rituals: Kanliwu, Machiwu, Anongko, A–Taye (birth ritual). Bookwell.
Delley [Dele], R. (2023). Reh: Ethnography of the Idu Mishmi shamanic ritualistic festival of Arunachal Pradesh. Bookwell.
dos Santos, A. F., & Leal, J. P. (2024). Early findings in using LLMs to assess semantic relations strength (Short Paper). In 13th Symposium on Languages, Applications and Technologies (SLATE 2024) (p. 4:1–4:9). Schloss Dagstuhl – Leibniz-Zentrum für Informatik. https://doi.org/10.4230/OASIcs.SLATE.2024.4
Dunkel, G. E. (2021). The oral style of the R̥gveda. Oral Tradition, 35(1), 3–36.
Eckmann, J. P., Kamphorst, S. O., & Ruelle, D. (1987). Recurrence plots of dynamical systems. Europhysics Letters, 4(9), 973–977. https://doi.org/10.1209/0295-5075/4/9/004
Farrar, S. (2003). An ontological account of linguistics: Extending SUMO with GOLD. In Proceedings of the International Conference on Natural Language Processing and Knowledge Engineering (IEEE, 2003, 797–806.
Foote, J. (1999). Visualizing music and audio using self-similarity. In Proceedings of the seventh ACM international conference on Multimedia (Part 1) 77–80.
Foote, J., & Cooper, M. (2001). Visualizing Musical Structure and Rhythm via Self-Similarity. In ICMC 1, 423–430.
Fox, J. (1988). To speak in pairs: Essays on the ritual languages of Eastern Indonesia. Cambridge University Press.
Gaenszle, M., & Gaenszle, M. (Eds.). (2018). Ritual speech in the Himalayas: Oral texts and their contexts (3–16). Harvard University Press.
Hainsworth, J. B. (1962). The Homeric formula and the problem of its transmission. Bulletin of the Institute of Classical Studies, 9, 57–68. https://doi.org/10.1111/j.2041-5370.1962.tb00696.x
Helfman, J. I. (1994, October). Similarity patterns in language. In Proceedings of 1994 IEEE Symposium on Visual Languages (pp. 173–175). IEEE.
Jakobson, R. (1966). Grammatical parallelism and its Russian facet. Language, 42(2), 398–429. https://doi.org/10.2307/411699
Jamison, S. W., & Brereton, J. P. (2014). The Rigveda: The earliest religious poetry of India. Oxford University Press.
Kinjawadekar, D. P. (1983). Rgveda Samhita. Royal Danish Library. https://loar.kb.dk/items/520b8f85-a8ef-4b50-8c94-d72114b6fb1e
Kiss, B., Kölligan, D., Mondaca, F., Neuefeind, C., Reinöhl, U., & Sahle, P. (2019). It takes a village: Co–developing VedaWeb, a digital research platform for Old Indo–Aryan texts. In S. Krauwer & D. Fišer (Eds.), TwinTalks: Understanding Collaboration in Digital Humanities. Fourth Digital Humanities Conference in the Nordic Countries 2019 (DHN2019) (pp. 35–44). CEUR–WS.
Kölligan, D., Neuefeind, C., Reinöhl, U., Sahle, P., Casaretto, A., Fischer, A., Kiss, B., Korobzow, N., Rolshoven, J., Halfmann, J., & Mondaca, F. (n.d.). VedaWeb – Online research platform for Old Indic texts. University of Cologne. https://vedaweb.uni-koeln.de/
Levenshtein, V. I. (1966). Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10(8), 707–710.
Marwan, N., Romano, M. C., Thiel, M., & Kurths, J. (2007). Recurrence plots for the analysis of complex systems. Physics Reports, 438(5–6), 237–329. https://doi.org/10.1016/j.physrep.2006.11.001
Milner, R. M. (2016). The world made meme: Public conversations and participatory media. MIT Press. https://doi.org/10.7551/mitpress/9780262034999.001.0001
Nash, W. (1989). Rhetoric: The wit of persuasion. Blackwell.
Orsucci, F., Walter, K., Giuliani, A., Webber Jr, C. L., & Zbilut, J. P. (1997). Orthographic structuring of human speech and texts: linguistic application of recurrence quantification analysis. arXiv preprint cmp-lg/9712010. https://doi.org/10.48550/arXiv.cmp-lg/9712010
Päpcke, S., Weitin, T., Herget, K., Glawion, A., & Brandes, U. (2023). Stylometric similarity in literary corpora: Non-authorship clustering and Deutscher Novellenschatz. Digital Scholarship in the Humanities, 38(1), 277–295. https://doi.org/10.1093/llc/fqac039
Parry, M. (1928). L’Épithète traditionnelle dans Homère: Essai sur un problème de style homérique. [The traditional epithet in Homer: Essay on a problem of Homeric style]. Société d’éditions Les belles lettres.
Reinöhl, U. (2022). Locating Kera’a (Idu Mishmi) in its linguistic neighbourhood: Evidence from dialectology. In M. W. Post, S. Morey, & T. Huber (Eds.), Ethnolinguistic Prehistory of the Eastern Himalaya 232–263. Brill.
Reinöhl, U. (2024). Documentation of Igu, Data Center for the Humanities. [Online]. Available: https://doi.org/10.18716/dch/b.00000017
Reinöhl, U. (2026). Poetic structures and strategies in Igu, an Eastern Himalayan shamanic language. Oral tradition.
Reinöhl, U., Pulu, P., & Wallner, U. (2025). A sketch grammar of Igu, the shamanic language of the Kera’a. Himalayan Linguistics, 24(1), 58–87.
Scheinkman, J. A., & LeBaron, B. (1989). Nonlinear dynamics and GNP data. In Economic complexity: chaos, sunspots, bubbles, and nonlinearity, Proceedings of the Fourth International Symposium in Economic Theory and Econometrics 213–227. Cambridge University Press, Cambridge.
Schultz, A. (2006). Recurrence quantification analysis of speech. The Journal of the Acoustical Society of America, 119(5), 3339. https://doi.org/10.1121/1.4786429
Tannen, D. (1987). Repetition in Conversation: Toward a Poetics of Talk. Language, 63(3), 574–605. https://doi.org/10.2307/415006
von Contzen, E., Pfänder, S., Reinöhl, U., & Sulimma, M. (2024). Repetition, again: Cross-disciplinary approaches to practices and forms of repeating. DIEGESIS Interdisciplinary E-Journal for Narrative Research, 13(2), 138–159. https://doi.org/10.25926/dr7c-jt58
WebberJr, C. L., & Zbilut, J. P. (1994). Dynamical assessment of physiological systems and states using recurrence plot strategies. Journal of applied physiology, 76(2), 965–973. https://doi.org/10.1152/jappl.1994.76.2.965
Zbilut, J. P., Koebbe, M., Loeb, H., & Mayer-Kress, G. (1990). Use of recurrence plots in the analysis of heart beat intervals. In [1990] Proceedings Computers in Cardiology (pp. 263–266). IEEE. https://doi.org/10.1109/CIC.1990.144211
The recordings, transcriptions and translations of the Igu data were produced and carried out in collaboration with Igu Pachu Pulu and Usha Wallner. U.R. wants to express special thanks to the Kera’a people for inviting her into their community and sharing their linguistic and cultural knowledge with her. The idea to generate recurrence maps for the Litany of the Blessed Virgin Mary came from Anna Nowicka, a student in U.R.’s class on repetition in language in the summer term 2023. We thank an anonymous reviewer for pointing us at Foote’s (and others’) work on self-similarity matrices.
Open Access funding enabled and organized by Projekt DEAL. Partial funding was received from the SFB 1252 Prominence in Language (German Research Foundation), individual project C09, for TME.
University of Freiburg, Freiburg, Germany
Uta Reinöhl
Universität zu Köln, Köln, Germany
T. Mark Ellison
Authors
U.R. provided the linguistic data analysis of Vedic Sanskrit and Igu, and is the primary author of the linguistic sections of the main text. T.M.E. developed the recurrence plotting software and produced the recurrence plots presented in this paper. T.M.E. is the primary author of the technical sections of the main text. Both authors take equal responsibility for the conceptual content.
Corresponding authorCorrespondence to Uta Reinöhl.
The authors declare no competing interests.
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
ChatGPT 4.0 Prompt: Please provide phonetic transcription in the International Phonetic Alphabet of Blake's poem "The Tyger" using the best reconstruction of his intended pronunciation. Please think about your answer before responding.
The programme returned a transcription and commentary of the first verse. So we added a second prompt: Please provide similar transcriptions and descriptions of verses 2 to 5 of the same poem.
Result:
To transcribe William Blake's *The Tyger* in the International Phonetic Alphabet (IPA) based on a reconstruction of late 18th-century London English, we must consider the phonological and phonetic changes of that period. Blake's dialect likely reflected the Southern English pronunciation of his time, influenced by Early Modern English (the 17th century) and leading into Modern English.
### Considerations for the Transcription:
**Non-rhoticity**: Blake's English was likely non-rhotic, meaning the /r/ at the end of syllables or before consonants would not be pronounced (e.g., "Tyger" would lack the final /r/).
**Tense vowels**: Some vowel sounds would differ from today's modern English, with long vowels or diphthongs in words like "eye" and "Tyger."
**Consonantal differences**: Words like "fear" might have a more centralized /iə/ sound instead of a diphthong /ɪə/ in contemporary speech.
**Voicing and aspiration**: Voiceless stops /p/, /t/, and /k/ might have less aspiration than in present-day English.
### IPA Transcription (in reconstructed 18th-century London English):
#### First stanza of *The Tyger*:
**The original text**:
```
Tyger Tyger, burning bright,
In the forests of the night;
What immortal hand or eye,
Could frame thy fearful symmetry?
```
**IPA transcription**:
/ˈtaɪɡə ˈtaɪɡə, ˈbɜːnɪŋ braɪt,
ɪn ðə ˈfɒrɪsts əv ðə naɪt;
wɒt ɪˈmɔːtl̩ hænd ɔːr aɪ,
kʊd freɪm ðaɪ ˈfɪəfʊl ˈsɪmɪtriː?/
### Notes on Transcription:
**/taɪɡə/**: "Tyger" is non-rhotic, without the final /r/ sound.
**/ˈbɜːnɪŋ/**: "Burning" has a more open /ɜː/ sound, typical of the period.
**/ðə ˈfɒrɪsts/**: "Forests" retains a more fronted /ɒ/ vowel, closer to that of Early Modern English.
**/wɒt ɪˈmɔːtl̩/**: "What immortal" has a broad "o" in "what" and a syllabic /l/ in "immortal."
**/aɪ/**: "Eye" is pronounced with a diphthong typical of Southern English dialects of the time.
**/kʊd freɪm/**: "Could" is a short /ʊ/, and "frame" retains the contemporary pronunciation of /freɪm/.
**/ˈfɪəfʊl ˈsɪmɪtriː/**: "Fearful" features a more open /ɪə/, and "symmetry" ends with /triː/, reflecting the earlier pronunciation of the final vowel.
This phonetic transcription attempts to balance historical reconstruction with comprehensibility, reflecting Blake's pronunciation as it may have been spoken in late 18th-century London.
### Verses 2-5 of *The Tyger* by William Blake
We will follow the same phonetic principles for reconstructing Blake’s late 18th-century London pronunciation.
### Verse 2
**The original text:**
```
In what distant deeps or skies
Burnt the fire of thine eyes?
On what wings dare he aspire?
What the hand, dare seize the fire?
```
**IPA transcription:**
/ɪn wɒt ˈdɪstənt diːps ɔː skaɪz
bɜːnt ðə ˈfaɪə əv ðaɪn aɪz?
ɒn wɒt wɪŋz dɛə hi əˈspaɪə?
wɒt ðə hænd, dɛə siːz ðə ˈfaɪə?/
---
**Notes on transcription:**
**/wɒt ˈdɪstənt diːps/**: "What distant" uses a rounded /ɒ/ in "what" and a lengthened /iː/ in "deeps".
**/bɜːnt ðə ˈfaɪə/**: "Burnt" uses a long /ɜː/, and "fire" maintains the diphthong /aɪə/, which may have been realized with a more central vowel glide compared to modern English.
**/aɪz/**: "Eyes" retains the long diphthong /aɪ/, typical of the Southern English accent of the time.
**/dɛə hi əˈspaɪə/**: "Dare he aspire" has a more open /ɛə/ in "dare", and "aspire" preserves a broad /aɪə/ diphthong.
---
### Verse 3
**The original text:**
```
And what shoulder, & what art,
Could twist the sinews of thy heart?
And when thy heart began to beat,
What dread hand? & what dread feet?
**IPA transcription:**
/ənd wɒt ˈʃoʊldə, ænd wɒt ɑːt,
kʊd twɪst ðə ˈsɪnjuːz əv ðaɪ hɑːt?
ənd wɛn ðaɪ hɑːt bɪˈɡæn tə biːt,
wɒt drɛd hænd? ænd wɒt drɛd fiːt?/
---
**Notes on transcription:**
**/wɒt ˈʃoʊldə/**: "Shoulder" contains an earlier /oʊ/ diphthong.
**/ɑːt/**: "Art" has a long /ɑː/ reflecting a non-rhotic realization.
**/ˈsɪnjuːz/**: "Sinews" has the long /uː/ vowel typical of older English pronunciation.
**/hɑːt/**: "Heart" would be pronounced with a long open back vowel /ɑː/, without the final /r/.
**/wɛn/**: "When" would be pronounced closer to /wɛn/ rather than contemporary /wɪn/.
**/drɛd fiːt/**: "Dread feet" features long /ɛ/ and /iː/ vowels.
---
### Verse 4
**The original text:**
```
What the hammer? what the chain?
In what furnace was thy brain?
What the anvil? what dread grasp
Dare its deadly terrors clasp?
```
**IPA transcription:**
/wɒt ðə ˈhæmə? wɒt ðə tʃeɪn?
ɪn wɒt ˈfɜːnɪs wɒz ðaɪ breɪn?
wɒt ðə ˈænvɪl? wɒt drɛd ɡrɑːsp
dɛə ɪts ˈdɛdli ˈtɛrəz klɑːsp?/
---
**Notes on transcription:**
**/ˈhæmə/**: "Hammer" uses an open /æ/ and drops the final /r/.
**/ˈfɜːnɪs/**: "Furnace" retains a long /ɜː/ vowel, typical of earlier forms of English.
**/wɒz ðaɪ breɪn/**: "Was thy brain" uses a more rounded vowel in "was" /wɒz/ and retains the broad diphthong /aɪ/ in "thy".
**/ɡrɑːsp/**: "Grasp" retains an /ɑː/ vowel in the first syllable and the final /sp/ cluster.
---
### Verse 5
**The original text:**
```
When the stars threw down their spears,
And water’d heaven with their tears:
Did he smile his work to see?
Did he who made the Lamb make thee?
**IPA transcription:**
/wɛn ðə stɑːrz θruː daʊn ðɛə spɪəz,
ənd ˈwɔːtəd ˈhɛvən wɪð ðɛə tɪəz:
dɪd hi smaɪl ɪz wɜːk tə siː?
dɪd hi huː meɪd ðə læm meɪk ðiː?/
---
**Notes on transcription:**
**/stɑːrz/**: "Stars" retains a long /ɑː/ vowel and non-rhoticity.
**/θruː daʊn/**: "Threw down" has the diphthong /aʊ/ in "down," while "threw" maintains a high rounded vowel /uː/.
**/ˈwɔːtəd/**: "Water'd" features the historical broad /ɔː/ in "water."
**/wɜːk/**: "Work" uses a centralized vowel /ɜː/ in a non-rhotic pronunciation.
**/dɪd hi huː/**: "Did he who" uses a rounded /uː/ in "who," which is a more archaic feature.
**/læm/**: "Lamb" retains the short /æ/ vowel.
---
### Conclusion
This IPA transcription reflects phonological features typical of late 18th-century London English, including non-rhoticity, broader diphthongs, long vowels, and some archaic vowel distinctions. Although Blake's exact pronunciation cannot be fully reconstructed, this transcription provides a detailed approximation based on historical phonology.
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
Reinöhl, U., Ellison, T.M. Recurrence plots – a tool for exploring repetition in un(der)studied language data. Int J Digit Humanities (2026). https://doi.org/10.1007/s42803-026-00130-2
Received: 10 June 2025
Revised: 11 March 2026
Accepted: 07 July 2026
Published: 23 July 2026
Version of record: 23 July 2026
DOI: https://doi.org/10.1007/s42803-026-00130-2
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Modelling non-stationary extremal dependence through a geometric approach | 0 | 5.83 | 24-07-2026 |
| 2 | RUMBA: русскоязычный бенчмарк для оценки долгосрочной памяти | 0 | 7.52 | 24-07-2026 |
| 3 | Group LASSO for multiple change-point detection in a generalized integer-valued autoregressive model | 0 | 9.18 | 24-07-2026 |
| 4 | Quantile adaptive feature screening for ultra-high dimensional longitudinal heterogeneous data | 0 | 8.98 | 24-07-2026 |
| 5 | Folklore Fusion | 0 | 14.4 | 14-07-2026 |
| 6 | Renewable Diesel Boom and Market Transition Towards Resilience in Soybeans | 0 | 7 | 15-07-2026 |
| 7 | Great ape laughter reveals a hidden origin of human speech | 0 | 6.72 | 02-07-2026 |
| 8 | Help with the protocol for accurately recording an orbit for the IAU | 0 | 5 | 12-07-2026 |
| 9 | Systems Forecasting By-Parts: A Study on Fed Beef Production | 0 | 7 | 15-07-2026 |