Section 6 of 9
The language–machine analogy
Matteo Pasquinelli · about 5 minutes
As much as the organism–machine analogy is key for understanding the modern episteme until cybernetics, the language–machine analogy plays a similar role for the twentieth century and the age of AI. The analogy is about the possibility that a machine can emulate and automate the generative structure of language—the interplay of ‘information, mechanism, and meaning’ (MacKay 1969). This issue can be historically illustrated by looking at the telegraph as an epistemic mediator between the domains of machine and language in late modernity (Wise 1988). Telegraphy performed a material and heuristic formalisation of language into bits of information that influenced, with the help of mathematics, the rise of both information theory and structural linguistics. In the history of science, the telegraph represented a watershed innovation: it de facto implemented a formalisation of language into a mathematical code (i.e. Morse code), often under-recognised as a forerunner of the binary code of the digital age. In different intellectual moments, the telegraph played the role of model machine for sensation, cognition and computation. In nineteenth century Berlin, scientists Emil du Bois-Reymond and Hermann von Helmholtz (Hoffmann 2003) took the telegraph as a model for the physiology of the nervous system (Otis 2001). British mathematician Alan Turing (1936, 1950) adopted the telegraph as a model of the Turing Machine, and implicitly also in the ‘imitation game’ known as the Turing Test. US mathematician Claude Shannon (1938, 1948) inaugurated information theory as a quantification of language in telegraphy to solve problems of compression, encryption and transmission along communication channels. After Shannon, cyberneticians Warren McCulloch and Walter Pitts (1943, 1947) adopted the telegraph as the model to formalise biological neurons into artificial neurons. At the 1948 Hixon Symposium at the California Institute of Technology, McCulloch insisted on conceiving ‘neurons as telegraphic relays’ (Jeffress 1951:45).
Looking attentively at the history of cybernetics and information theory, it must be noted that two paradigms of information co-existed through two models of machine, which engendered further hybrids, such as the current form of AI. The coupling of the paradigms of machine and organism was the concern of Wiener (1948). The coupling of the paradigms of machine and language was the concern of Shannon (1948). The former adopted the model of information from organicist biology (Uexküll 1920b), the latter from military telegraphy and signal intelligence. Wiener maintained a model of information that was substantially analogue, concerned with control and adaptation, whilst Shannon advanced a discrete one, concerned with encoding, bandwidth and transmission. It is the latter paradigm of linear information (i.e. information codified in a linear medium such as written language) that became hegemonic and grounded the computer age. However, the current form of AI emerged from the former paradigm of self-organising information. More precisely, one should see the invention of artificial neural networks such as the Perceptron (Rosenblatt 1958, 1962) as the confluence of the two paradigms of organism–machine and language–machine. The confluence of these two paradigms can be easily detected also in McCulloch and Pitts (1943, 1947), when they introduced the idea of artificial neuron. The artificial neuron (McCulloch and Pitts 1943) was envisioned to encode Boolean logic with the later ambition to automate the recognition of visual patterns (McCulloch and Pitts 1947). Today, deep neural networks, such as LLMs, are computing networks made of trillions of parameters (also known as parametrised machines) that adjust and self-organise their weights in order to compute a statistical model of large data repositories.
Looking at the trajectory of computation over the last century, the formalisation of language is found both at its inception—in telegraphy and information theory—and at its culmination in the planetary hegemony of LLMs. Shannon (1948) founded information theory to quantify the ‘intelligence’ or interpretability of a signal in telegraphy, drawing on earlier work from Andrey Markov, who analysed the first 20,000 letters of Pushkin’s Eugene Onegin. Shannon defined ‘information’ as a statistical measure of token frequency in printed English. The intuition of digital information, in fact, is not simply about the discretisation of language into a binary code, as in telegraphy: it is an index of the predictability of its letters into sequences and compounds. It is often said that information theory is not concerned with content (semantics), but with reducing language to its grammatical structure (syntax). Like Markov, Shannon sought to demonstrate that a natural language is statistically dependent: therefore, any text can be measured, computed and transmitted as a stable message across a noisy channel. In time, computation grounded the large statistical models of AI, which came to analyse and automate increasingly complex linguistic artefacts, inspiring a new phase in linguistics, that is the automation of linguistics (Léon 2015/2021).
Thus, LLMs revive the old structuralist hypothesis that language is a vast system of dichotomic differences which now can be measured accurately as complex positions in a vector space. Old structuralism turned into neo-structuralism (Guariento 2026): whereas the former encoded meaning as a binary opposition or dichotomy in a low-dimensional space, the latter measures meaning as cosine similarity in a multi-dimensional vector space of greater flexibility. The performance of deep learning neural networks (such as Transformer models, e.g. ChatGPT) has given new impetus to the ergodic theory in linguistics according to which AI models would demonstrate the computability of human language as such. Already Shannon (1948) had treated language as a ‘stochastic process’ and noted that an English corpus behaves like a ‘stationary ergodic source’ for the purposes of prediction. By training on massive textual corpora (treated as samples from an underlying stationary ergodic linguistic source), contemporary LLMs behave as if the long-run frequencies of words and constructions can approximate the full generative rules of language. The Italian semiotician Umberto Eco once warned that this process is algorithmically impossible, because language is a creative and living process, not just the imitation of the past. One of the entities that LLMs struggle to decode and encode is a new poetic metaphor: ‘No algorithm exists for the metaphor, nor can a metaphor be produced by means of a computer’s precise instructions, no matter what the volume of organized information to be fed in’ (Eco 1986: 127).