Brilliant Sector
← Journal
AI in schools 12 June 2026 8 min read

The test we use before putting AI in a district workflow

Most school AI pilots come unstuck for the same unglamorous reason: everyone agrees to try it before anyone agrees what success would look like, so what gets measured is enthusiasm. Here is the test we run before writing any code.

In short
  • A use case starts with a metric you can name and someone who already watches it.
  • The best early candidates are high volume, low stakes, and already have an adult reviewing the output.
  • Set the condition that would end the pilot before the pilot starts.

The problem with the question

Every week someone in a district asks a version of the same question: should we be using AI for this? It rarely has a clean answer, because the thing being pointed at is usually a whole workflow rather than a single task. The first move is to cut that workflow into pieces small enough to judge.

Once the pieces are on the table, most of them answer themselves. A few will hold out, and those are the ones worth a real conversation.

The test, in three questions

We ask the same three questions of any candidate task, in this order. A task that fails the first one goes no further, because there is little point discussing models for work that nobody is measuring.

First, what number moves if this works, and who already watches it? Second, when the output is wrong, does an adult catch it, and how quickly? Third, what does this cost to keep running in eighteen months, including the person who has to care about it once the grant closes?

Most failed school AI projects were good ideas that nobody agreed how to grade.

What good candidates look like

The tasks that survive this test tend to share a shape. They are usually:

  1. High volume, so a small saving per item compounds into time somebody can feel.
  2. Low stakes per item, where a wrong output is annoying and gets caught well before it reaches a student or a family.
  3. Already reviewed by an adult, so there is an existing checkpoint doing the quality control.

Novelty, board interest, and a neighbouring district’s press release are all fair reasons to look into something. They are worth keeping separate from the reasons to build it.

Questions we get about this

How long should a pilot run before we decide?

Run it long enough to see the metric move twice, which for most school operations means four to six weeks, and keep it inside a single term. If six weeks leaves you unsure, look at the measurement first.

What if the answer is that AI will not help?

Then that is the deliverable, and it is a useful one. We have written reports whose main recommendation was to fix a data-entry process and revisit the question in a year, and the district saved a budget cycle.

Want this applied to your own workflow?

Get in touch