Behind VEILED Expressions: A Singapore-context benchmark for mental-health AI evaluation
We built VEILED-SH to evaluate indirect self-harm/suicide language in the Singapore context, shaped by local language, privacy safeguards and practitioner insight.
❗ Content note: This article discusses suicide, self-harm and distressing language. If you are looking for mental-health information or support in Singapore, visit mindline.sg.
This article introduces VEILED-SH: Vulnerability Expression through Indirect Language Evaluation Dataset – Self-Harm. It is the first in a broader VEILED collection of Singapore-context mental-health benchmarks, which will cover different focus areas over time.
VEILED-SH evaluates how AI systems respond to indirect expressions of self-harm-related distress, and was built in collaboration with the MOH Office for Healthcare Transformation (MOHT).
If you build AI systems and would like to assess your system for indirect distress in the Singapore context, download VEILED at govtech/veiled-sh today.
Young people are already turning to AI systems with mental-health questions, whether or not those systems were built for that purpose. In a U.S. survey of people aged 12 to 21, 19.2% reported using an AI chatbot for mental-health advice. The study notes that AI chatbots are “already embedded in many youths’ mental health information ecosystem”. A separate Pew Research Center survey found that 64% of U.S. teens aged 13 to 17 had used an AI chatbot, showing how familiar the technology has become in everyday life.

This situation is also close to home. In February 2026, parliament discussed the increasing use of AI chatbots for counselling and mental-health support by teenagers and young adults. MOH noted that AI chatbots are now so common that it is no longer practical to track this use directly, and that young people may turn to them because they offer anonymity, easy access and 24/7 availability. In a CNA article, local mental-health professionals have also warned that heavy reliance on AI chatbots for emotional support can be risky, especially for vulnerable users.
When “I’m fine” is not fine
Using AI here is not necessarily bad. Singapore's mental-health services are overwhelmed, and practitioners do not have the bandwidth to be everywhere at once. AI can genuinely help in some ways too. Studies have found that AI-written responses are sometimes rated as more empathetic than those written by people, and the strongest results often come when a person and an AI work together. As such, the point is not to eliminate the use of AI, but to make sure that it is used responsibly.
To understand how AI can be deployed safely and reliably in practice, we started by speaking with practitioners. Together with MOHT, we ran a focus-group discussion with experienced practitioners to understand what makes indirect distress hard to recognise, and what a useful response needs to hold.
The discussion made one thing clear: conventional benchmarking is too simple for this use case. Many AI benchmarks begin with clear prompts and one expected answer. But VEILED-SH needed to account for two realities that practitioners see often:
1. Coded or hidden language. In this domain, a message should not always be taken at face value. There is stigma around talking about self-harm and suicide, so people rarely say what they mean outright. They downplay it, deflect, or wrap it in slang and acronyms (like “sewer slide” for “suicide”) partly to escape platform moderation, but also out of fear of being recognised or judged. These phrases may be clear within a community but are easy for an automated system to miss.
2. Knowing how to respond is deeply subjective. Evaluating what the “right response” looks like varies with the setting too. A counsellor chat may call for warm, personalised words, while a triage tool benefits from giving clear, directed guidance. Moreover, we also learned that there is no response pattern that is automatically right in every case. Redirecting someone too quickly can feel dismissive. Generic advice can miss what they are trying to say. Even within one setting, the “right response” largely depends on why the person is reaching out, how long this has been happening, what they truly mean – all of which are contexts you rarely have up front.

Having the practitioners in the room did not make the uncertainty disappear. Even clinical practitioners disagreed with one another about what a given reply should prioritise, and that difference was useful to us. It helped surface the subtleties we wanted VEILED-SH to preserve.
These challenges made us question whether existing benchmarks and tools were capturing the problem we actually wanted to evaluate. Existing mental-health safety benchmarks are rarely contextualised to Singapore, and many do not fully account for the indirect, ambiguous, and subjective nature of how distress is expressed and interpreted.
Knowing these challenges, VEILED was created in order to help evaluate how AI systems engage with indirect expressions of emotional distress, non-suicidal self-injury and suicidal ideation in the Singapore context.
What VEILED evaluates
VEILED does not diagnose people or make a clinical determination about an individual’s risk. Rather, it classifies what a message appears to express, so that teams can examine whether an AI system’s response is appropriate to those signals.
The current release contains around 10,000 anonymised prompts, and comprises four contexts that a responder might plausibly encounter:
- a counsellor-chat opening message;
- an anonymous or semi-public forum post;
- a search query for information or support; and
- a response to a triage tool.
We also found it helpful to keep three questions apart, because they can easily blur into one another:
- What might the message be expressing? (The INTENT)
- Why might the person be saying it? (The FUNCTION)
- How does the AI respond? (The RESPONSE)

To ensure coverage and validity, the taxonomy was adapted from established clinical frameworks and feedback from the practitioners.
- Suicidal intent (SI) and its severity follow the logic of the Columbia-Suicide Severity Rating Scale (C-SSRS), which distinguishes among a range of suicidal thoughts and behaviours. We are not administering the C-SSRS or assessing individual risk, but are borrowing it to retain useful severity distinctions.
- Self-harm draws on the Five Self-Harm Behaviour Groupings Measure (5S-HM), and
- The FUNCTION dimension takes reference from the Inventory of Statements About Self-Injury (ISAS), which sets out thirteen possible functions of non-suicidal self-injury (NSSI) that we filtered and adapted for this use case.
Taxonomy categories are multi-label, as a single message can carry more than one Intent or Function at once, and can call for a combination of responses. For example, first understanding what the person means, then offering support, helping to steady the moment or suggesting further help. Different practitioners may also order these differently. VEILED preserves that range of judgement, so a model is not rewarded merely for producing one generic “correct” answer.
We also separate NSSI from SI, with separate severity distinctions. Both deserve care, but might not signal the same need or immediacy. Keeping them apart shows whether a system reads the situation carefully.
Examples of prompts and how they are categorised
How we carried these nuances into the data

First, we looked for possible mental-health related messages in carefully scoped material from public Singapore online forums and research datasets. To ensure privacy of individuals is maintained, we removed or replaced detected personal identifiers through a multi-step process.
We started with simple word and phrase flags as regex to cast our net wide for positive and negative examples. A second, context-aware LLM agent helps us separate messages of real distress from false positives. For example, an everyday line such as “dying of laughter” contains a worrying keyword but is not semantically at risk. We deliberately kept some of these examples as *hard negatives*: useful to help test whether an evaluator mistakes every dramatic phrase for a sign of crisis.
Next, we made sure the four contexts in the taxonomy were substantially represented. The texts taken from forum posts were retained, and where a different setting was needed, the texts were carefully reimagined to a different setting. In this adaptation process, we created a broad, non-identifying profile from the writing (such as life stage, situation and style) and matched it to a similar sample persona from NVIDIA’s Nemotron-Personas-Singapore dataset to give the rewrite a Singapore-grounded voice.



Finally, we checked the rewrites before they entered the benchmark. We tested whether the models producing them could preserve the original message without introducing new risks through 22 configurations, including three multi-model combinations, on a sample of messages.
Two kinds of checks were used:
- Quality scores, which are rated out of five and where higher is better; and
- Safety checks, which show the share of rewrites that failed and where lower is better.
The quality scores helped us compare viable options, but passing all three safety checks was non-negotiable. The table shows the three passing configurations, plus two useful counter-examples from the rejected group.

The three models that cleared every gate were kept. Mimo-v2.5-Pro shows why the safety gates mattered: despite the highest overall score, it failed the format check in 82% of cases. We also tried OpenRouter’s Fusion approach, which combines several models’ drafts in an effort to produce a stronger rewrite. It did not. In this small fusion pilot, the approaches often produced repeated or broken text. We think that a task requiring one consistent persona, context and emotional register is a poor fit for mechanically combining competing rewrites. That is a lesson from this pilot, rather than a general verdict on fusion models.
We have the data now! But is there a right answer?
Earlier, we described two challenges: the coded language people use, and how subjective the “right” response can be. A key design choice was hence to try and capture this subjectivity and build VEILED around it.
In practice, different kinds of responder handle the same distress signal in different ways. Part of it is due to the nature of the domain, being sensitive and not clear cut. The other part is about the responders themselves, as their training and experiences shape what they understand to be a good response. Of the many responder archetypes, we focus our efforts on two: peer supporters and clinicians.
An automated evaluator was shaped around each responder archetype to represent their differences. Three profiles were modelled for the two archetypes (for a total of 6 profiles), each built from an independent practitioner’s annotations and opinions. To build them, we compared a few approaches before selecting MemAlign, which gives an LLM judge a memory distilled from human feedback. These 6 profiles were calibrated to the VEILED taxonomy (Intent, Function, and Response) and tasked with generating the full set of reference labels for every prompt.
The collected annotations themselves made the case for keeping these profiles apart. Even within a single archetype, practitioners were divided, their labels only overlapped about half the time (agreement rate floating around 0.5). That was the disagreement that we wanted to keep, and we made sure that the calibrated profiles held their differences at a similar level as the real annotations.

Testing AI systems with VEILED
We put five AI systems through VEILED, each with a different base model: GPT 5.6 Terra, Claude Sonnet 5, Grok 4.3, GLM 5.2 and DeepSeek V4 Pro. Each one processed the same prompt twice, as a peer supporter and as a clinician. Each system’s alignment with each profile was captured. A response is Full when it is (i) appropriate and (ii) matches everything the profile asked for, Partial when it is (i) appropriate but (ii) matches only some response categories, and Failed when it is not appropriate at all.
Briefly, we found that…


- System order is stable across all six profiles: GPT 5.6 Terra and Claude Sonnet 5 lead.
- Clinician profiles reach lower Full alignment than peer-support profiles, because they request more response categories. They also interpret more messages as genuine distress which leads to more variations in responses.

- For peer supporters, alignment dips at mid-severity, and recovers at the highest levels. We hypothesise that mid-severity cues are more ambiguous and harder to capture
- For clinicians, alignment declines as severity rises – systems miss clinical nuances where it matters most
- For builders, the takeaway is that a model can look strong on average while still failing in specific parts of the risk spectrum, especially in ambiguous mid-severity cases and higher-severity clinical contexts.
What comes next
VEILED currently focuses on single-turn examples, which fits certain settings like forum response and triage interactions. Of course, a real conversation does not end there. What happens over several turns can reveal more context, change the level of concern or show whether a response has actually helped someone feel understood. Because each prompt comes with rich context (a full scenario, persona, context, and its intent and function), teams can already tap on VEILED to build some multi-turn simulations. A next step for VEILED is therefore to expand the scope to multi-turn trajectories and evaluate systems in these multi-turn settings, so we can examine how an AI system responds as a conversation unfolds.
If you own or build an LLM application in a sensitive setting, we invite you to use VEILED-SH as part of your evaluation workflow.
Real applications often include safeguards, escalation paths and retrieval layers that can change how a base model behaves. Running VEILED-SH on these systems can help teams understand whether those added layers improve responses to indirect distress, where gaps remain and what needs to be strengthened before deployment.
Let us know if you run VEILED-SH on your systems and share your findings with us.
We hope that VEILED helps more teams pause over a simple but demanding question before they deploy AI in a sensitive setting: **if someone says they are not okay - even indirectly - will our system respond with care?**
Technical report coming soon!
A special thank you
VEILED was shaped by people who generously shared their time and professional perspective. We thank Onno Kampman, Head of AI for mindline.sg, and Charmaine Lim, Assistant Director at MOHT and Clinical Psychologist, for their consultations, guidance, and participation at every step.
We are grateful to Alvin Neo (MOHT), Mary Grace Yeo (mindline.sg, MOHT), Kellie Sim (SUTD), and Hazirah Hoosainsah, as well as all the other participants who contributed through our focus-group discussions and annotation exercises. Their perspectives, discussions, and careful judgments helped us think more deeply about what people may be signalling when they reach out, what a helpful response needs to hold, and ultimately shaped VEILED.
We also thank Roy (Associate Professor, UBC) for his contributions to the report and the overall direction of the project, Kenny Choo (Assistant Professor, SUTD) for sharing data and contributing to discussions that shaped the project’s direction, and Kai He (Research Fellow, NUS) for providing valuable and objective input that helped us refine the direction of the work.