Individual Submission Summary
Share...

Direct link:

The Architecture of "Attention": Uncrumpling NLP’s Linguistic Ideologies

Thu, September 5, 2:45 to 4:15pm, Sheraton New Orleans Hotel, Floor: Five, Grand Ballroom B

Abstract

In 2017, researchers at Google published a paper entitled “Attention is All You Need” describing a neural network model they called the ‘Transformer’, which achieved state-of-the-art performance on English-French/French-German machine translation (and other tasks of Western European denotational transduction), dramatically reducing past feats of bespoke rule-based engineering to a few hundred lines of code. This was achieved through a holistic elevation of a technique known in the field of natural language processing (NLP) as ‘attention’ which, as I will describe here, is a material and pragmatic attempt to cope with the intrinsic temporality of speech/text within infrastructural systems like Google’s, which instead prefer to harness and manipulate human language as a synchronic, unchanging, and wholly vectorized structure. What, then, does this ‘attention’ of artificial neural networks have to do with attention as historical and cultural phenomena? In this paper, I will suggest that just as the attention of 18th-century naturalists to their environments was bound up with notions of utilitarian ‘fitness’ (Daston 2004), the ‘attention’ of deep learning (DL) makes this relationship profoundly literal: the trained parameters of such models are always optimized through reference to some ‘utility function’, which expresses the success or failure of a translation as a single commensurable value (Espeland and Stevens 1998). ‘Attending’ to this correspondence can help unpack, or in the high-dimensional DL parlance, ‘uncrumple’ the ideological underpinnings of a conception of language which the large-scale deployment of these models currently reifies.

Author