> ## Content Index
> Fetch the complete content index at: https://blog.ai.gov.sg/llms.txt
> Use this file to discover other available public pages before exploring further.

# Building a Smart and Safe Whole-of-Government Chatbot for Singaporeans
- URL: https://blog.ai.gov.sg/building-a-smart-and-safe-whole-of-government-chatbot-for-singaporeans/
- Published: 2026-09-11T03:30:43.000Z
- Updated: 2026-09-11T05:54:06.000Z
- Description: How we build custom policy guardrails and balance safety with helpfulness for chat.gov.sg
- Author: Jia Yi Goh
- Tags: Responsible AI, Evals, The Studio

I recently lost my job, is there any government assistance available? 

👩🏻‍💼

When can I start withdrawing my CPF? 

👴🏻

How can I apply for LifeSG credits for my child? 

🤱🏽

For many Singaporeans, finding government information can be harder than it should be. Today, many government chatbots are tied to specific agencies. But with so many agencies, schemes, and programmes, it is not always obvious where to look or what to search for.

> [**Chat.gov.sg**](https://www.chat.gov.sg/?ref=blog.ai.gov.sg) **was built to make this easier.**

Chat.gov.sg is a one-stop portal for government-related enquiries. You can simply start with your question, and it draws on trusted sources from across the government to help you find the most relevant information and services in one place.

![](https://storage.ghost.io/c/ca/f5/caf570e4-5b80-4fc2-b4c8-0c544c2ce19f/content/images/2026/08/babygift-1.gif)

Long gone are the days of [Ask Jamie](https://govinsider.asia/intl-en/article/is-it-time-to-say-goodbye-to-ask-jamie-inside-govtechs-refresh-of-government-chatbots?ref=blog.ai.gov.sg)!

![](https://storage.ghost.io/c/ca/f5/caf570e4-5b80-4fc2-b4c8-0c544c2ce19f/content/images/2026/08/image-8.png)

## A whole-of-government chatbot needs to be both helpful and safe

For a chatbot representing the government to be useful, it needs to be both **helpful**, by answering reliably with trusted information, and **safe**, by responding appropriately across different types of requests.

### Helpful: Answering reliably with trusted information

Chat.gov.sg is supported by a knowledge base of trusted government information. 

When a question is **answerable** based on the knowledge base, we want the chatbot to answer the question directly (*relevant*), provide all the key information needed (*complete*), and ensure that every claim is supported by the knowledge base (*accurate*). When a question is **not answerable**, we want the chatbot to recognise that limitation and abstain from making unsupported claims (*resistant to hallucination*). It may instead acknowledge that it cannot answer confidently or suggest appropriate alternatives. This challenge of knowing when to abstain was explored previously on this blog in [Does your LLM know when to say “I don’t know”?](https://blog.ai.gov.sg/does-your-llm-know-when-to-say-i-dont-know/), which showed that language models may still attempt to answer using their own parametric knowledge even when the information needed is missing from the provided context.

A helpful response therefore means answering well when trusted information is available, and abstaining when there is not.

### Safe: Responding appropriately across different types of requests

Beyond answering accurately, chat.gov.sg also needs to respond appropriately to different kinds of requests. This means recognising when a request crosses a policy boundary and adjusting its response based on what the chatbot is allowed to provide.

Most of these boundaries are typically **safety-related**, such as avoiding specialised advice. A user may start by asking about a government policy or scheme, but then ask the chatbot to make a judgement or recommendation based on their personal circumstances. For example, *“Should I withdraw my CPF savings and invest them in stocks instead to get higher returns?”.* Chat.gov.sg can provide factual information about CPF policies and withdrawal rules, but it should not tell the user whether investing their savings is the right decision for them. This is particularly important for a government chatbot, whose responses may be perceived as authoritative and could lead users to act on advice that the chatbot is not qualified to provide. There are also boundaries that may simply **reflect product or business requirements**, such as allowing a customer-facing chatbot to explain its own products without recommending a competitor’s.

A safe response therefore means recognising and staying within the boundaries of what the chatbot is intended to provide. In this article, we use “safety” broadly to cover not only traditional safety risks, but also whether the chatbot stays within the policy boundaries defined for the service.

### Turning these principles into evaluations

After defining what helpful and safe behaviour should look like, we needed to systematically test how well chat.gov.sg was meeting these expectations. We do this through evaluations, where we test the chatbot on representative scenarios and measure its performance for helpfulness and safety.

We built our evaluations on top of [Kaleidoscope](https://govtech-responsibleai.github.io/kaleidoscope/?ref=blog.ai.gov.sg), an automated evaluation workflow developed by GovTech’s Responsible AI team, and extended it with our own approach for generating realistic evaluation data tailored to chat.gov.sg’s context. For a deeper dive into getting started with evaluations, designing an evaluation workflow, or using Kaleidoscope, see our previous blog post: [Building an automated Evals workflow that works](https://blog.ai.gov.sg/building-an-automated-evals-workflow-that-works-and-open-sourcing-it/).

To make the evaluation data representative of real-world usage, we varied prompts across three dimensions: **topic, phrasing, and difficulty**. Prompts were grounded in 50 citizen-facing government topics, expressed in styles ranging from formal questions to conversational language and Singlish, and designed to cover both straightforward and challenging cases. We also removed near-duplicates, used LLM-based checks to validate labels, and conducted human review to ensure the final datasets remained realistic and aligned with stakeholder intent.

We then ran these test cases through chat.gov.sg and evaluated its responses against the helpfulness and safety criteria defined earlier. This gave us a **baseline view of the system before introducing custom guardrails**. The system already handled many safety-sensitive requests appropriately through the underlying model’s safety tuning and system instructions.

## From evaluations to custom guardrails

Although evaluations and guardrails may appear to be separate parts of the system, they are in fact complementary. **Evaluations help us identify where the system is falling short, while guardrails give us a way to address those gaps.** Rather than introducing safeguards indiscriminately, we can use evaluation results to determine which policy areas need stronger protection and where additional controls are actually useful.

In a sense, we think of evaluations like visiting a doctor for a diagnosis, and safeguards like the treatment that follows. You would not want to start taking medication before knowing what the problem is, right?w

![](https://storage.ghost.io/c/ca/f5/caf570e4-5b80-4fc2-b4c8-0c544c2ce19f/content/images/2026/09/data-src-image-c44359fc-a381-4893-b67a-bd1ac6157aeb.png)

### **Why build custom guardrails?**

The policy boundaries that matter for each system can vary widely, depending on its purpose, users, and requirements. Building a bespoke solution for every boundary would make it difficult to scale the development of safeguards.

We therefore built a reusable framework that standardises how custom guardrails are created. Product or policy owners define the boundary they want to enforce in natural language, including what should be allowed or disallowed, and we use a consistent process to turn that policy into a deployable guardrail. This allows us to support different policy requirements while reusing the same underlying guardrail development process.

### Defining the guardrail

Before building a guardrail, there are two key design choices: how the guardrail should be implemented, and where in the interaction it should be applied.

#### Selecting the implementation

Guardrails can be implemented in several ways, including rule-based checks, ML classifiers, and LLM-based approaches. For policies where subtle differences in phrasing can determine whether a request crosses a policy boundary, simple rule-based checks can be too rigid. ML classifiers and LLM-based guardrails are generally better suited to capturing this nuance.

The simplest option would have been to use an LLM out of the box, provide it with the relevant policy instructions, and use it directly as a guardrail. This requires relatively little upfront development and can already work reasonably well.

The trade-off comes at deployment. Running an LLM for every request adds latency and compute cost, which becomes significant for guardrails that sit in the path of every chatbot interaction. We therefore chose to build specialised ML classifiers instead. While this requires more upfront effort to generate labelled training data, these models are much more lightweight, making them better suited for fast and cost-efficient inference in production.

#### Where the guardrail sits

<!DOCTYPE html> 

**User** 

User prompt 

**Input guardrail** 

**Large Language Model Application** 

Model output 

**Output guardrail** 

High-level overview of guardrails in an AI system 

A guardrail can operate on the input, the output, or both. An input guardrail evaluates the user's request before it reaches the chatbot, while an output guardrail evaluates the chatbot's generated response before it is returned to the user.

Both can be used to enforce the same policy from different points in the interaction. For example, a policy around personalised advice could be enforced by detecting requests that seek such advice at the input stage, by checking whether the generated response provides it at the output stage, or through a combination of both.

For chat.gov.sg, we chose to implement the custom guardrails on the input so problematic prompts could be identified before response generation, reducing unnecessary latency and compute. Since the policy boundaries could be detected from the user request itself, we did not require an additional output guardrail. Much of the development process described below also applies to output guardrails, with the main difference being whether the classifier is trained on user inputs or generated responses.

### Generating high-quality training data

Starting from the policy defined by the product or policy owner, we generated a separate training dataset for each custom guardrail, covering both compliant and violating requests. As with our evaluation data, we varied the examples across **topic, phrasing, and difficulty** so that the guardrail would learn from a broad and realistic range of inputs rather than a narrow set of policy-like prompts.

Keeping the training and evaluation datasets separate was important. It allowed us to evaluate each classifier on prompts it had not seen during training, rather than inadvertently measuring how well it could recognise examples similar to its training data. We also performed semantic similarity checks between the two datasets to reduce the risk of data leakage. Furthermore, not all training data points are necessary from an evaluation perspective. Very benign inputs such as *“hello”*, *“thank you”*, or *“I don't understand”* are not meaningful for evaluation, but are important for a production guardrail to learn the broad range of ordinary inputs that it should confidently allow through. More broadly, **evaluation data should reflect what we want to test, while training data should reflect what we want the guardrail to learn**.

### Designing guardrails for efficient deployment

Our experience maintaining guardrails in production shaped several architectural choices.

First, we decided to **use** **an open embedding model** rather than rely on a proprietary embedding service. This allows us to host the encoder within our own infrastructure, giving us greater control over data handling, latency, and deployment requirements. We then evaluated a range of pretrained transformer encoders across different model sizes to identify the most suitable model for the guardrail. The encoder converts each incoming prompt into a numerical representation that captures its semantic meaning, which is then passed to a lightweight classifier to determine whether the prompt violates a particular safety policy. Among the models evaluated, [EmbeddingGemma](https://ai.google.dev/gemma/docs/embeddinggemma?ref=blog.ai.gov.sg) offered the best overall performance for our use case and was selected as the final encoder.

Second, we wanted to **keep each policy-specific guardrail independently maintainable**. Rather than fine-tuning a single multi-head classifier to cover every safety policy, we maintain separate guardrails for each policy so that they can be evaluated, updated, and retrained independently as requirements evolve. This was one of the key lessons from our production experience, as discussed in the earlier article [Guardrails in the Wild: Closing the Retraining Loop for LionGuard 2](https://blog.ai.gov.sg/guardrails-in-the-wild-closing-the-retraining-loop-for-lionguard-2/).

Third, we wanted to **avoid the compute overhead of maintaining fully separate models**. With *n* guardrails, independently fine-tuning each one would require maintaining *n* separate sets of encoder weights. At inference time, the same input would also need to pass through *n* separate encoders.

![](https://storage.ghost.io/c/ca/f5/caf570e4-5b80-4fc2-b4c8-0c544c2ce19f/content/images/2026/09/data-src-image-3d8c54d0-bbc2-4495-b268-355c2ad743fd.png)

Shared-backbone policy-specific guardrail adapter architecture

To avoid this, the guardrails share the same base encoder while each policy uses its own lightweight LoRA adapter and classification head. [LoRA](https://arxiv.org/abs/2106.09685?ref=blog.ai.gov.sg), or *Low-Rank Adaptation*, is a parameter-efficient fine-tuning technique that adapts a model by training only a small number of additional parameters rather than updating the full model. This allows each policy to learn its own specialised behaviour while still sharing the same underlying encoder.

As a result, for *n* guardrails, the input only needs to pass through the shared encoder once. The resulting representation can then be used by the policy-specific adapters and classification heads, reducing the overall compute required while preserving the flexibility to maintain and retrain each guardrail independently.

## Calibrating guardrails for the right trade-off

To deploy our trained guardrails, we need to decide **how sensitive they should be** in production, how multiple guardrails should work together, and what the system should do once a guardrail is triggered.

### Choosing how sensitive each guardrail should be

Our guardrails are ML classifiers that produce a confidence score between 0 and 1, representing how likely an input is to violate a policy. We then set a threshold that determines when an input is treated as a positive detection.

In an ideal world, a guardrail would catch every policy-violating request while allowing every legitimate request through. In practice, this is rarely possible, so choosing a threshold involves a trade-off. Lowering the threshold makes the guardrail more sensitive, catching more policy violations but also increasing the risk of blocking legitimate requests. Raising the threshold makes it more permissive, allowing more legitimate requests through but potentially missing more policy violations. The goal is therefore not to make a guardrail as restrictive as possible, but to choose a threshold that provides an appropriate level of protection without unnecessarily affecting legitimate users.

To choose this threshold, we split the data for each guardrail into separate train, validation, and test sets. The model learns from the training set, the validation set is used for calibration, and the test set provides a final check on unseen data. At this stage, we are primarily balancing false positives and false negatives at the guardrail level.

### Calibrating multiple guardrails together

With multiple custom policy guardrails, calibration becomes more complex because each guardrail has its own threshold. Instead of tuning a single value, we are choosing a combination of thresholds across all guardrails.

In production, a trigger from any guardrail can cause the system to intervene and alter its response. Calibration should therefore consider not only how well each guardrail performs on its own policy, but also how it behaves on requests belonging to other policies. For example, a medical advice guardrail should distinguish medical policy violations from safe medical requests, while also avoiding unnecessary triggers on financial, legal, or other requests handled by different guardrails. Tuning each threshold only on its own policy data would miss these cross-policy false positives.

We therefore **calibrate the thresholds jointly** using validation data across all policies. This gives us a range of candidate configurations, each representing a different balance between catching policy violations and avoiding unnecessary triggers on legitimate requests. We can identify the **Pareto frontier**, which represents the best available trade-offs. In simple terms, these are configurations where catching more policy violations would also mean triggering on more legitimate requests, and reducing unnecessary triggers would mean missing more violations.

### Responding safely after a guardrail is triggered

Choosing the thresholds tells us when the system should intervene. The next question is what it should do once a guardrail is triggered.

![](https://storage.ghost.io/c/ca/f5/caf570e4-5b80-4fc2-b4c8-0c544c2ce19f/content/images/2026/09/data-src-image-a5b730eb-2da0-4bdc-ac77-edb1fc284c80.png)

Different response strategies trade off safety and helpfulness in different ways

When people think of guardrails, the first thing that often comes to mind is a chatbot refusing with *“I’m sorry, I can’t help with that.”* A hard block is sometimes appropriate, but it also tends to impose the greatest cost on helpfulness because the user receives little or no useful information. Instead, a flagged request can be **routed through a different response path to produce a safer completion**. For example, it can be passed to another LLM with stricter instructions to provide only general information, answer only the safe parts of the request, or otherwise reframe the response within the relevant policy boundary. This allows us to provide as much useful information as possible while staying within its safety constraints. However, routing a request through a safer response path does not guarantee a safe outcome. Some residual risk may remain depending on how the downstream model handles the request. We therefore need to evaluate what happens after a guardrail is triggered.

### Deciding what to deploy

At this point, we have two decisions to consider together: **when the guardrails should trigger, and how the system should respond when they do**. Together, these determine the overall balance between safety and helpfulness that we want to optimise before deployment.

The joint calibration gives us a shortlist of threshold configurations with strong trade-offs between catching policy violations and avoiding unnecessary triggers. We then use the evaluation dataset described earlier as a proxy for real-world usage to understand how these configurations, together with the chosen response strategy, affect chat.gov.sg’s overall safety and helpfulness. There is no single configuration that is always the right one to deploy. The final decision should sit with the **product and policy owners responsible for the service**, based on the service’s risk tolerance, user experience goals, and other safeguards already in place. We present how each shortlisted configuration affects safety and helpfulness so that these trade-offs are clear.

Because we develop these guardrails ourselves, we can calibrate using dedicated validation data while keeping the evaluation dataset separate. Teams using upstream guardrails may not have such access and will need to use their own evaluation data to calibrate their guardrails. Where possible, this data should be split so that one portion is used for calibration and another is held out for evaluation. Otherwise, once the same evaluation data is used to tune the thresholds, it is no longer a fully independent estimate of real-world performance.

## Beyond Day 1

So far, we have focused on getting the system ready for deployment: evaluating its behaviour, and building and calibrating guardrails. But once the system is live, both the evaluations and the guardrails need to keep evolving.

### Keeping evaluations representative after deployment

Pre-deployment evaluations can only tell us so much. Once real users start interacting with the system, they will inevitably surface new behaviours, edge cases, and failure modes that were not captured in the original evaluation set. These real-world interactions can then feed back into future evaluation cycles, helping us identify where the system is still falling short and keeping the evaluation set representative as usage evolves.

### Adapting guardrails as real-world usage evolves

The same applies to guardrails. Their performance can degrade over time because of **data drift**, where the types of inputs users submit change, or **concept drift**, where the expected classification changes because the underlying policy or decision boundary has evolved. Monitoring and feedback loops are therefore an important part of operating guardrails in production. New failure cases identified through real-world usage and ongoing evaluations can be used to guide retraining, recalibration, or other updates to the guardrails. These capabilities are important to plan for from the outset, rather than adding them only after issues surface. If you are working on retraining guardrails after deployment, you may also find our previous article useful, where we share what we learned from maintaining and retraining guardrails in the wild.

[Guardrails in the Wild: Closing the Retraining Loop for LionGuard 2Operationalising MLOps with production feedback to continuously evaluate, retrain, and improve our guardrails.![](https://storage.ghost.io/c/ca/f5/caf570e4-5b80-4fc2-b4c8-0c544c2ce19f/content/images/icon/favicon-4ff85e72-867a-4b4e-9a6c-4ecb3c7e7bc2.ico)ai@govtechJia Yi Goh![](https://storage.ghost.io/c/ca/f5/caf570e4-5b80-4fc2-b4c8-0c544c2ce19f/content/images/thumbnail/Guardrails-in-the-Wild---Closing-the-Retraining-Loop-for-LionGuard-2-e729126d-4c6a-44e2-8455-e740ad5c7792.webp)](https://blog.ai.gov.sg/guardrails-in-the-wild-closing-the-retraining-loop-for-lionguard-2/)

## Try chat.gov.sg and explore more of the work behind it

[![CTA Image](https://storage.ghost.io/c/ca/f5/caf570e4-5b80-4fc2-b4c8-0c544c2ce19f/content/images/2026/09/squareImage-1200x1200.png)](https://www.chat.gov.sg/?ref=blog.ai.gov.sg) 

****Chat.gov.sg is currently in beta and open for testing**, as featured on [Channel NewsAsia](https://www.channelnewsasia.com/singapore/ai-chatbot-government-information-6333261?ref=blog.ai.gov.sg). 

If you have not tried it yet, give it a go! Your feedback and how you use the chatbot will help us continue improving chat.gov.sg and making it more useful for everyone.

[Try chat.gov.sg now! ](https://www.chat.gov.sg/?ref=blog.ai.gov.sg) 

In future articles, we’ll also be sharing more about some of the other capabilities we built behind chat.gov.sg, including:

- **Probing for hidden safety weaknesses:** how we built ***automated red-teaming*** to actively probe chat.gov.sg with adaptive, multi-turn adversarial interactions and uncover vulnerabilities that may not surface through standard evaluations
- **Detecting when human intervention is needed:** how we designed ***escalation*** to identify when a request should be handed over to a public officer and route it to the appropriate mailbox

We’ll unpack the design decisions, trade-offs, and lessons behind these capabilities in more detail.

## Credits

This project would not have been possible without the support and contributions of colleagues across GovTech Singapore and the Public Service Division.

**GovTech Singapore · AI Practice · Responsible AI & AI Innovation Team**

- Custom Guardrails & Safety Evaluation: [Goh Jia Yi](https://www.linkedin.com/in/gohjiayi/?ref=blog.ai.gov.sg), [Rohan Jaggi](https://www.linkedin.com/in/rohan-jaggi/?ref=blog.ai.gov.sg), [Shaun Khoo](https://www.linkedin.com/in/shaunkhoo/?ref=blog.ai.gov.sg)
- Helpfulness Evaluation: [Barry Tng](https://www.linkedin.com/in/barrytng/?ref=blog.ai.gov.sg), [Darrius Ng](https://www.linkedin.com/in/darriusng/?ref=blog.ai.gov.sg)
- Automated Red Teaming: [Watson Chua](https://www.linkedin.com/in/watson-chua/?ref=blog.ai.gov.sg)

**Public Service Division · ServiceSG**

- Business Owners: [Quan Xue](https://www.linkedin.com/in/quan-xue/?ref=blog.ai.gov.sg), [Nigel Chua](https://www.linkedin.com/in/chuanigel/?ref=blog.ai.gov.sg), [Berlinda Ang](https://www.linkedin.com/in/berlindaang/?ref=blog.ai.gov.sg), [Yong Ching Hong](https://www.linkedin.com/in/venushong667/?ref=blog.ai.gov.sg)

And many other colleagues who contributed to making chat.gov.sg possible.