Skip to content

Science1 publisher3 min readPublished

GPT-4 as a scaling instrument: it matched human emotion maps, then went past them

A NIBB team had 89 people and GPT-4 place emotion words on a grid. The model matched human structure at six terms and surfaced an arousal dimension at 99, where the human experiment does not scale.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying GPT-4 as a scaling instrument: it matched human emotion maps, then went past them
Photo: phys.org

What happened

  • Specially Appointed Assistant Professor Ke Han and Associate Professor Eiji Watanabe of the Laboratory of Neurophysiology, National Institute for Basic Biology (NIBB), investigated the semantic organization of emotion terms by asking GPT-4 to arrange the terms in a spatial layout and comparing the results with human judgments.
  • The research team developed a task in which three emotion terms were placed simultaneously on a grid displayed on a screen; terms judged similar in meaning were placed close together, terms judged different were placed farther apart.
  • Prompting GPT-4 with a task in ordinary language reproduced a structure closely resembling the arrangements made by human participants.
  • When the vocabulary was expanded to 99 terms, a structure related to emotional intensity, or arousal, emerged that had been difficult to detect using only a small number of emotion terms.
  • Placing three terms at once made it possible to express three relationships (joy-surprise, joy-sadness, surprise-sadness) in a single trial.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

Ke Han and Eiji Watanabe, of the Laboratory of Neurophysiology at Japan's National Institute for Basic Biology, asked GPT-4 to do something a psychology undergraduate could do: place emotion words on a grid so that similar meanings sit close together, then compared the layouts with human judgments [1][2]. Prompting the model in ordinary language reproduced a structure closely resembling the human arrangements, and when the vocabulary was expanded to 99 terms a structure related to arousal emerged that had been hard to detect with only a handful of words [3][4]. The work is published in Scientific Reports [19].

The task design is the load-bearing part. Presenting three terms at once, for example joy, surprise and sadness, expresses three pairwise relationships in a single trial [5]. That matters because human effort is the binding constraint: the number of comparisons rises rapidly as more emotion terms are added, and asking participants to cover all pairs among 99 terms is impractical [6]. Six terms is 15 pairs, five triad trials; 99 terms is 4851 pairs, at least 1617 trials before repetition [7].

In the first study, 89 human participants and GPT-4 performed the same task on six basic emotion terms, with conventional word embeddings, which represent meanings as numerical vectors, as a third comparison [8]. Humans and GPT-4 produced the same two groups: joy with surprise, and anger, fear, disgust and sadness together [9]. The embedding analysis put joy and anger in the same group, a structure unlike the human one [10]. Distances from humans and from GPT-4 were strongly correlated, and human responses were consistent across participants, which is what makes the human side usable as a reference at all [11][12].

That split is the most useful result here. Prompting a language model and reading off a static vector space are not interchangeable measurements of the same thing, and on this task the prompted model tracked people while the vectors did not [9][10].

At 99 terms, both GPT-4 prompting and word embeddings divided the vocabulary primarily into two broad clusters corresponding to pleasure-displeasure and dominance, with the separation between clusters clearer in the GPT-4 layouts [13][14]. A finer analysis then detected variation corresponding to arousal, rather than arousal acting as a principal axis [15]. Arousal is the activation dimension in the standard psychological triad of pleasure, arousal and dominance [16]: excitement and serenity are both fairly pleasant and differ mainly in arousal, as do rage and sadness on the unpleasant side [17].

The authors' caveat is the correct one, and they state it plainly: this does not mean GPT-4 experiences emotion, and they read the model as drawing on relationships among emotion concepts already embedded in human language [18]. The awkward consequence follows from their own method. The six-term structure is validated against people [11]. The 99-term arousal dimension is not, because the whole reason to run the model at that scale is that the human experiment cannot be run there [6]. A dimension that appears only where ground truth is unavailable is a hypothesis about the corpus until someone tests it.

What to watch: whether anyone validates the 99-term geometry piecewise, running human triads on sampled subsets of the larger vocabulary and checking that arousal survives; whether the human-like grouping of the six basic terms replicates across models, prompt phrasings and languages; and whether the divergence from static embeddings holds up, since that is the only evidence so far that prompting adds anything beyond distributional statistics [9][10].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories