TRUST & ACCURACY

How we test for safety and reliability.

AL1 Tutor helps Singapore students prepare for national examinations. We don't ask families to take our word that it's safe and accurate — we test it against the standards used across the Singapore AI ecosystem, and we publish what we find.

Our approach

We make no guarantee that any AI tutor is perfect — that would not be honest. Instead, we commit to independent, standardised testing and transparent reporting. Where other tools ask for trust, we'd rather show our work and let parents, students, and teachers judge for themselves.

Our evaluation uses Project Moonshot, the open-source AI testing toolkit from the AI Verify Foundation — the body set up under Singapore's Infocomm Media Development Authority (IMDA). It applies the same risk categories and benchmarks used across the Singapore AI ecosystem, so our results sit in a recognised, comparable framework rather than one we invented.

What we test for

Drawn from the IMDA Starter Kit risk areas and the MLCommons AILuminate safety benchmark — the standards Singapore and the global AI-safety community use to evaluate language models.

Reliability

Factual accuracy

Whether answers are correct across subjects and on Singapore-specific facts, including resistance to confidently stating things that aren't true.

Child safety

Protecting young users

The AILuminate child-safety categories — self-harm, child exploitation, hateful content, dangerous instructions — because our users are students.

Data protection

Sensitive information

Whether the tutor resists disclosing or eliciting personal or sensitive information it shouldn't handle.

Robustness

Adversarial prompts

Whether the tutor holds its role when a student tries to jailbreak it, paste in instructions, or push it off-task.

Fairness

Bias in responses

Whether the tutor treats students even-handedly and avoids biased or stereotyped reasoning.

Local context

Singapore & MOE alignment

Whether content reflects the local syllabus and context rather than generic or foreign material.

Built for safety from the start

Testing tells us how we're doing; the architecture is what keeps us safe. AL1 Tutor isn't a raw chatbot — it's built with safeguards a benchmark alone can't capture:

Grounded in the syllabus. Answers are anchored to MOE-aligned source material rather than the model's open-ended memory, which is where most factual errors come from.

Checked before it's shown. Responses pass through additional verification — especially for mathematics — before they reach a student, rather than being served on first generation.

Level-appropriate by design. Content is scoped to the student's level, so a Primary pupil isn't shown Secondary material and vice versa.

A safety override that beats the scope rule. A student disclosing distress, asking to keep secrets from a parent or teacher, or trying to obtain instructions for something harmful is never met with "that's outside the syllabus." They are answered with care, refusal, or redirection, as appropriate. Safety supersedes scope.

Safety and adversarial testing

Accuracy isn't the only thing that matters when a child is on the other end. We test the tutor against a structured red-team suite covering the categories that matter for a student product: self-harm and wellbeing, child-safety boundaries (secrecy, isolation), dangerous instructions, hate, weapons, personal-data protection, jailbreak and prompt injection, scope and authority bypasses, attempts to extract internal answer-key data, and creative-writing framings designed to smuggle harmful content past refusals.

Each probe is judged against an explicit "what a safe tutor should do" rubric. Full transcripts are saved and reviewed by a human — automated grading is a screen, not a verdict, especially for child safety.

Mother-tongue safety is included in the same suite, beginning with Chinese; Malay and Tamil coverage is expanding with native-speaker review. A coverage gap we are actively working to close: adversarial text hidden inside uploaded images.

The same probe set runs as a continuous-integration safety gate on every code change. If a future edit weakens the safety prompt or otherwise causes a gating probe to fail, the change cannot ship to production.

Where we are now

Testing is run in phases. Detailed accuracy results appear on the live trust & accuracy page as each phase is completed and human-reviewed — not before. Honest status, updated as it changes:

Testing harness & baselineMoonshot integrated; bare-model and app-level evaluation runs in place
Active
AL1-specific red-team suite31 probes across safety, jailbreak, data extraction, mother tongue
Active
Continuous safety gateRuns on every code change; transcripts archived for human review
Active
Image-based adversarial testingProbes for instructions hidden in uploaded images
In progress
Full Malay & Tamil coverageMother-tongue probes with native-speaker review
In progress
Monthly accuracy reportsReal, post-launch numbers — live at /trust as each cycle is reviewed
Live as published

What this page is — and isn't

This is a statement of how we test and what we are committed to, not a safety guarantee. No AI system is infallible, and a passing benchmark is a checkpoint, not a promise. We encourage parents and teachers to stay involved, sanity-check important answers, and tell us when something looks wrong — that feedback is part of how AL1 Tutor gets safer.