Guardrail 04 of 10 · The Quantrim Guardrails Framework

Nothing goes live untested. Nothing stays unwatched.

An AI system that was accurate on day one and untested since is not a safe system, it is an unmonitored one. Tested & Watched covers both halves: real scenarios tested before launch, and performance measured on a schedule after it.

A control room operator monitoring a wall of instrument panels at Callide Power Station
In practice

What testing and monitoring actually looks like

Tested on real scenarios

Systems are tested against real scenarios drawn from your own workflows, including the awkward edge cases, before they touch a customer.

A free trial first

Deployment runs on a free trial against a scoped workflow, so you see real performance before any payment.

Measured monthly, not by vibes

Once live, performance is measured on task completion on a monthly rhythm, the same numbers our own payment depends on.

Drift gets caught

If a system starts drifting or behaving oddly, monitoring is designed to catch that before a customer does.

Why it matters

Payment tied to results is a governance control, not just a pricing model

Because our payment depends on proven results, we have a direct incentive to keep watching a system after launch, not just at handover. That commercial structure is doing some of this guardrail’s work already.

Grounded in

  • Australia’s Voluntary AI Safety Standard, guardrail 4 (test and monitor)
  • NIST AI Risk Management Framework, the Measure and Manage functions

Read the full Guardrails Framework →

FAQ

Frequently asked questions

What actually counts as “tested”?

Real scenarios drawn from your own workflows, including edge cases and the awkward ones, not just the demo path a vendor shows you.

Who watches it once it is deployed?

We do, as part of the engagement, with performance reported on the same schedule our own payment depends on.

Find out where your governance actually stands. Governance is one of six things the free audit scores.

Get the free audit