Skip to content

Science1 publisher3 min readPublished

MIT researchers score suicide risk in 16,000 crisis chats with a 49-factor word list

MIT researchers report that a lexicon of 49 suicide risk factors accurately predicted risk levels in about 16,000 Crisis Text Line chats. Each match traces to a named factor, so the scores can be audited, but they were checked against the crisis line's own ratings and still need clinical validation.

The Scientist · Science desk

Illustration accompanying MIT researchers score suicide risk in 16,000 crisis chats with a 49-factor word list

What happened

  • Daniel Low, working in Satra Ghosh's group at MIT's McGovern Institute, built a text tool that searches for words and phrases tied to 49 suicide risk factors.
  • The team analyzed de-identified texts from about 16,000 conversations between people in distress and Crisis Text Line's volunteer counselors.
  • In the Journal of Psychopathology and Clinical Science, Ghosh, Low and colleagues report that the tool accurately predicts suicide risk from crisis-line texts.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • constraint Reported accuracy measures agreement with Crisis Text Line's own ratings, so a service cannot yet tell whether the tool would catch texters whom counselors themselves misjudge.
  • capability A clinician reviewing a flagged conversation could see which named risk factors drove the score and check them against the transcript, an audit an opaque score does not allow directly.
  • decision A crisis service considering the tool has to wait for, or run, clinical validation first, since the researchers make further testing the condition for clinical and crisis-support use.

"Accurately" needs a reference point, and in this study the reference point is a human judgement. Crisis Text Line's own assessments sorted the roughly 16,000 conversations into three levels: non-suicidal, suicidal ideation without imminent risk, and imminent risk [4][5]. The model learned to reproduce those ratings. The top tier covers texters with a plan for suicide or an intent to die within the next 48 hours [6]. It is a sharp category, but it is assigned during the conversation. A rating made mid-chat is a different thing from a record of who later made an attempt.

The MIT account notes that even trained clinicians struggle to identify who will make an attempt among people with suicidal ideation [16]. A model trained on counselors' ratings learns their misses along with their judgement.

I find the design the most persuasive part of the work. The team had AI draft a list of words and phrases tied to established risk factors. They then reviewed it by hand, keeping about 60 entries for each of 49 factors, each confirmed as relevant by expert clinicians [7]. That comes to roughly 2,940 entries [14]. The model predicts risk from matches against that list. Because each entry belongs to a named factor, the team could work out which factors sit closest to imminent risk [8]. According to the MIT account, the pattern agreed with earlier research [9].

Low's case for this dataset is about timing. Crisis Text Line gave the team specialized training and controlled access to the restricted, de-identified data [15]. Earlier work on the question typically relied on epidemiological surveys that ask people to recall symptoms, often after the crisis has passed, he said [13]. "Crisis Text Line gives us an opportunity to assess many different symptoms and potential risk factors as people are having the crises," Low said [12].

The factor ranking is still a ranking of association. A factor whose terms turn up more often in imminent-risk chats tracks the counselors' rating. Whether it drives risk is a separate question, and Low describes the factors as entangled. "You see all these 50 risk factors, and they're all interacting in ways we don't really understand," he said [11].

The MIT account does not report the accuracy metric, how many of the 16,000 conversations fell in the imminent-risk tier, or any comparison with an opaque model trained on the same chats. That leaves the case for explainable triage over black-box scoring resting on the design alone. If imminent-risk chats are a small share of the sample, an overall accuracy score says little about how many of those texters the tool would flag.

I think the transparent approach is the right starting point for crisis triage, on two conditions. The paper's catch rate for the imminent-risk tier has to hold up, and a prospective test has to tie scores to what happened to texters afterward. The researchers attach a condition of their own. With more validation, the MIT account says, the tool could help with risk assessment in clinical and crisis-support settings [10].

What to watch

  • The paper's own accuracy figures, especially how many imminent-risk conversations the lexicon model correctly flags.
  • A head-to-head test against an opaque model trained on the same Crisis Text Line conversations.
  • A prospective validation that links lexicon scores to what later happened to texters, beyond counselor ratings.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories