Agent Wars, Episode 2: The Brain Wars
On shared memory systems for organisations and the complexities around them
Find all our past blog articles here. Filter by series, author, publication window, or keywords.
Latest articles
No articles match those filters.
On shared memory systems for organisations and the complexities around them
By Ng Shangru
Government communications can misfire in ways their authors never intended. We built an app that lets officers stress-test drafts against AI personas grounded in real Singapore voices, so they can catch what they missed in minutes rather than weeks.
By Shaun Ang
Tell us about yourself- what did you expect from this internship? I’m Rohan, a penultimate year Business Analytics student at the National University of Singapore (NUS) who spent the first half of 2026 as an applied AI Data Scientist Intern at GovTech’s Responsible AI team. Going in, I
By Rohan Jaggi
How we built Kaleidoscope: A structured workflow for realistic, scalable, and human-aligned contextual AI evaluations.
By Leanne Tan
Previously I wrote about building a harness for yourself. This one is about the environment you're building in, and why at enterprise scale, if the platform underneath doesn't exist, individual wins have nowhere to accumulate.
By Ng Shangru
Engineering Multi-Agent Architectures for Autonomous Penetration Testing.
By Watson Chua
On building your own multi-agent orchestrator, and why owning the infrastructure around AI matters.
By Ng Shangru
We tested 2026 SOTA models and found a "usability gap".
By Sumiko Teng
Ever spoken to an AI and felt like it was responding with insincere praise?
By Leanne Tan
The "hype" of robots ignore the unstructured environment problem.
By Jia Yi Goh
Measure how well your app is performing and more importantly where it's failing.
By Leanne Tan
A maturity roadmap and a cultural shift.
By AI Practice
The true value of AI agents lies in loops and self-correction rather than raw reasoning power.
By Ryan Lin
A universal "USB-C" for AI?
By Ng Shangru
A practical guide to MLOps adoption across Government teams.
By AI Practice
We tested the most effective approaches.
By AI Practice
Evaluating dimensions often overlooked by traditional benchmarks.
By Leanne Tan
Available LLMs are powerful enough. What we are missing is the knowledge to fuel them.
By AI Practice
We improved its coverage and robustness.
By Leanne Tan
Global safety guardrails are often blind to local dialects and sensitivities.
By Leanne Tan
Who Judges the Judge? At GovTech’s AI Practice, we’ve been embracing what’s known as “LLM-as-a-judge” — essentially employing LLMs as evaluators across our AI workflows. This approach has become one powerful approach in our evaluation toolkit. We use LLMs extensively across multiple areas: judging other
By Leanne Tan
Refusal by a model to answer may sometimes be more valuable.
By AI Practice
Addressing technical challenges of processing high-volume public feedback for policy-making
By AI Practice
Much a cultural shift as a technical one.
By AI Practice
Safer, faster testing of student-facing AI before real-world deployment.
By AI Practice