We had criteria, and we still could not agree
The situation
The team was bringing AI into content production, and we set the right condition early: no wider rollout before structured evals. The loop ran on human judgement. Generate a set of copy, evaluate it, revise the prompting on what the evaluation found, and go again until the routine cleared its threshold.
Two things wore that loop down. Scoring: the team had built a solid first evaluation framework, and every score still passed through the evaluator's private sense of where the line sat, so two people could score the same set differently with both right by their own scale. And feedback: results travelled by instant message and calls. The communication load climbed, information slipped between threads, and nobody could trace which action a piece of feedback had caused.
The diagnosis
Criteria without bands are still opinion. A criterion tells you what to look at, not what a “6” is. Each evaluator invented a scale in silence, and the numbers could not be compared, added or trended. And a threshold you cannot score consistently is not a threshold.
The second problem was memory. An eval loop is only as strong as its records, and this one kept none: revisions went undocumented, feedback evaporated on delivery, and keep-or-drop decisions ran on recollection.
Neither is a model problem. Both are measurement problems, which better prompting cannot fix.
The mechanism
Two moves: carry the judgement, then carry the process.
The bands. I restructured the framework so scores behaved like data: comparable, sortable, ready for calculation. Then the load-bearing addition, a defined score range for each criterion with what each band means, so a rating is a reading rather than a mood. The gap between evaluators narrowed until the task could be taught, and a task that can be taught can be handed to more teammates. Review stopped queueing on the same few heads. That is where the speed came from.
The records. As the rounds multiplied, the loop's running cost surfaced, and the people closest to it named it before I did: too much of the process was manual conversation, and the colleagues building on its results wanted every revision documented, so keep-or-drop could rest on comparison. My part was the architecture: what a record must carry, which fields a form needs, and versioning principles enforced by the system rather than by discipline. The process moved into the platform where the work already happened. Every revision and every evaluation became a record; a flagged output carried its context to the people troubleshooting it, and feedback that used to live in DMs happened against the record, in the same sitting.
What changed
The evaluators who had been arguing about taste discovered they had been arguing about bands, which is an argument that can end. Scoring spread: new evaluators came on without the line moving.
The loop gained a memory. Revisions were compared on records rather than recollection, keep-or-drop became a reading, and the case for a wider rollout stopped being anyone's opinion: a threshold, cleared or not yet. Flags carried their own context, so the reliance on DM fell away and feedback could be traced to the action it caused.
What transfers
Bands are how judgement changes hands. Most teams name their criteria and stop, then wonder why two reviewers still disagree and why the standard drifts with whoever applies it. Write down what each score means, and new people can hold the same line. A rollout gate hangs from the same place: once scores are readings, "can we trust the automation with this" becomes "what score does it have to clear", a question with an answer.
Then give the loop a memory, on rails people already use. Feedback that travels by message evaporates; a record outlives the conversation and shows what changed because of it. Version the process where the work already happens, and every evaluation becomes a record the whole chain can read. Built this way, the rails have room to grow into the place where an organisation collects feedback of every kind, setting up constant improvement rather than drastic revamps that require heavy investment each time.