Science1 distinct publisher3 min readUpdated
A NIBB team had 89 people and GPT-4 place emotion words on a grid. The model matched human structure at six terms and surfaced an arousal dimension at 99, where the human experiment does not scale.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Ke Han and Eiji Watanabe, of the Laboratory of Neurophysiology at Japan's National Institute for Basic Biology, asked GPT-4 to do something a psychology undergraduate could do: place emotion words on a grid so that similar meanings sit close together, then compared the layouts with human judgments [1][2]. Prompting the model in ordinary language reproduced a structure closely resembling the human arrangements, and when the vocabulary was expanded to 99 terms a structure related to arousal emerged that had been hard to detect with only a handful of words [3][4]. The work is published in Scientific Reports [19].
The task design is the load-bearing part. Presenting three terms at once, for example joy, surprise and sadness, expresses three pairwise relationships in a single trial [5]. That matters because human effort is the binding constraint: the number of comparisons rises rapidly as more emotion terms are added, and asking participants to cover all pairs among 99 terms is impractical [6]. Six terms is 15 pairs, five triad trials; 99 terms is 4851 pairs, at least 1617 trials before repetition [7].
In the first study, 89 human participants and GPT-4 performed the same task on six basic emotion terms, with conventional word embeddings, which represent meanings as numerical vectors, as a third comparison [8]. Humans and GPT-4 produced the same two groups: joy with surprise, and anger, fear, disgust and sadness together [9]. The embedding analysis put joy and anger in the same group, a structure unlike the human one [10]. Distances from humans and from GPT-4 were strongly correlated, and human responses were consistent across participants, which is what makes the human side usable as a reference at all [11][12].
That split is the most useful result here. Prompting a language model and reading off a static vector space are not interchangeable measurements of the same thing, and on this task the prompted model tracked people while the vectors did not [9][10].
At 99 terms, both GPT-4 prompting and word embeddings divided the vocabulary primarily into two broad clusters corresponding to pleasure-displeasure and dominance, with the separation between clusters clearer in the GPT-4 layouts [13][14]. A finer analysis then detected variation corresponding to arousal, rather than arousal acting as a principal axis [15]. Arousal is the activation dimension in the standard psychological triad of pleasure, arousal and dominance [16]: excitement and serenity are both fairly pleasant and differ mainly in arousal, as do rage and sadness on the unpleasant side [17].
The authors' caveat is the correct one, and they state it plainly: this does not mean GPT-4 experiences emotion, and they read the model as drawing on relationships among emotion concepts already embedded in human language [18]. The awkward consequence follows from their own method. The six-term structure is validated against people [11]. The 99-term arousal dimension is not, because the whole reason to run the model at that scale is that the human experiment cannot be run there [6]. A dimension that appears only where ground truth is unavailable is a hypothesis about the corpus until someone tests it.
What to watch: whether anyone validates the 99-term geometry piecewise, running human triads on sampled subsets of the larger vocabulary and checking that arousal survives; whether the human-like grouping of the six basic terms replicates across models, prompt phrasings and languages; and whether the divergence from static embeddings holds up, since that is the only evidence so far that prompting adds anything beyond distributional statistics [9][10].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Specially Appointed Assistant Professor Ke Han and Associate Professor Eiji Watanabe of the Laboratory of Neurophysiology, National Institute for Basic Biology (NIBB), investigated the semantic organization of emotion terms by asking GPT-4 to arrange the terms in a spatial layout and comparing the results with human judgments.
The research team developed a task in which three emotion terms were placed simultaneously on a grid displayed on a screen; terms judged similar in meaning were placed close together, terms judged different were placed farther apart.
Prompting GPT-4 with a task in ordinary language reproduced a structure closely resembling the arrangements made by human participants.
When the vocabulary was expanded to 99 terms, a structure related to emotional intensity, or arousal, emerged that had been difficult to detect using only a small number of emotion terms.
Placing three terms at once made it possible to express three relationships (joy-surprise, joy-sadness, surprise-sadness) in a single trial.
Investigating the entire structure using human participants alone is not easy because the number of comparisons rises rapidly as more emotion terms are included; repeatedly asking human participants to compare all pairwise relationships among 99 terms would be impractical, whereas GPT-4 can perform the same task across a large number of combinations.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed single study with a baseline comparison, but no reported effect sizes or replication
The core result rests on a published Scientific Reports study with a human control group of 89 participants, a preregistered-style identical task for model and humans, and a competing method (word embeddings) as a contrast, which is stronger than a bare demo. It is weakened by everything not reported in the supplied material: no correlation coefficients, no model snapshot or sampling procedure, no named embedding baseline, and no independent replication. All of it reaches us through one article.
One disclosed research use by the originating lab
The only observed use is the authors' own study. There is no evidence in the supplied material of other labs, tools, or products adopting the prompt-as-instrument method, no released code or dataset, and no downstream citation or replication signal. Scoring low reflects that single disclosed use rather than an inference that uptake is happening elsewhere.
Mildly overstated framing over a carefully caveated result
The write-up's own caveats keep it close to aligned: it states the model does not experience emotions and that arousal was a within-cluster dimension rather than the principal axis. The residual gap comes from framing a single-study method as 'a new tool' when there is no adoption beyond the originating lab, no released prompts or code, and no numeric agreement statistics behind 'strongly correlated' or 'clearer separation'.
Originating-lab announcement channel with an interest in method uptake
The sole source reads as an institutional research announcement relayed by a science-news aggregator: the researchers are the narrators, the closing section is an author quote advocating their method's accessibility over word-embedding pipelines, and no independent voice or critic appears. That is a normal academic promotion incentive rather than a commercial one, and it is partly offset by the authors' own limiting caveats. No funding, sponsorship or vendor relationship is disclosed in the supplied material, so nothing beyond this is assessed.
Credible peer-reviewed core, single unreplicated channel
Confidence is limited chiefly by structure rather than substance: one publisher, one apparently author-sourced article, no numbers behind the qualitative agreement claims, and a body text that is truncated mid-sentence. The peer-reviewed venue, the 89-participant human arm and the internal caveats keep it from being low.
science
Two Neanderthal pelvises suggest the strange hip belongs to the modern human male1 distinct publisher
science
What you expect from your own old age shows up a decade later in who you still see1 distinct publisher
science
Eastern US extreme rain is pooling into fewer, wider storms, and station records hide it1 distinct publisher
science
A named ship and a published timetable turn the Arctic into a bookable lane1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026