continuous improvement

part 03 of 03

Finished is not correct.

The unit of value is the case joined to its business outcome, not the agent’s trace. Every correction becomes a test, every incident becomes a rule, and nothing ships that fails the agent’s own cases.

§01   the measure

01 / 04

The whole workload is measured, not the easy cases.

The denominator is the original workload, with failures, escalations and repairs in it.

  • Verified completionA correct record in the system, days later, with no reopening. Not “the agent finished”.
  • Human minutesPer case of the original workload, review and repair included.
  • Reopened casesAccepted cases that came back for correction or dispute.
  • Authority failuresCounted apart. One serious one never disappears into an average.
  • Cost per accepted caseModel, review and operation, per case accepted, not per case attempted.
  • End-to-end timeFrom arrival to the accepted record, by class of case.

No fabricated percentages. The numbers enter when your baseline exists, and the method of comparison is written down before the work starts. For scale only: the Anthropic Institute’s economic scenarios (Korinek, Jones, Sacher, Cotter and McCrory, WP 2026-02, September 2026; the views are the authors’) take a gain per task of 35–57% from field trials and frontier firms as their anchor, and note that an estimate from conversation data implies about 5×. Neither is a promise about your workflow.

§01.1

The denominator is the original workload. Always.

§02   the loop

02 / 04

One loop joins operation and improvement.

A specialist’s correction becomes a test case; the owner releases; the result is measured on the next cases. Not a collection of dashboards, a way of deciding.

scene 08   the agent gets better

tests · gate · autonomy

Scene 8 · every real case is a test. A prompt change that fails one never ships; the one that passes all of them does, and the agent earns more room.owno.ai
  • 01   a case is handledThe agent runs the workflow; the case record holds every event, from intake to the outcome read back.
  • 02   a specialist correctsYour specialist fixes what was wrong, in the tool they already use. The correction is an event on the case, not an e-mail.
  • 03   the correction becomes a testThe case, with the right answer attached, joins the acceptance set. The suite grows with the agent’s history, not with someone’s imagination.
  • 04   a change is proposedA prompt edit, a new tool, a new model version. Each is a new agent until proven otherwise.
  • 05   the gateThe change runs against every real case. Fails one, does not ship. The case that caught it is attached.
  • 06   the owner signsA result compliance can co-sign replaces the review meeting. Nothing is released that nobody chose.
  • 07   measured on the next casesThe six numbers, per version, on the original workload. Better or worse is a fact on a scoreboard, not an opinion in a meeting.

i   tests from real cases

No synthetic benchmark

The cases come from the ledger: what the agent actually did, what the owner said about it, and what came of it. Each incident adds one.

ii   a gate before every change

Fails one case, does not ship

Prompt, tools, model: every change runs the suite first. The change that would have broken a real case never reaches production.

iii   autonomy that widens with evidence

Room is earned, not assumed

When the agent agrees with its owner on a long enough run and its tests keep passing, owno proposes a wider class it may handle alone. The owner signs. Incidents rise, and the room narrows again.

§03   the rulebook

03 / 04

Rules are not written once. They accumulate, and each one is signed.

A rule enters the book from one of four places. owno proposes it, replays it against last month’s real traffic so the owner sees exactly what it would have refused, and switches it on only when a named person signs.

source 1   your own incidents

The mistake becomes the rule

A duplicate send, a refund reversed, an export nobody approved. The incident becomes a test written from the real actions; the test becomes a rule that would have refused it. The same mistake cannot happen twice.

source 2   the owner’s decisions

Every yes and no is a labelled example

The approval inbox is a training set. Two hundred consistent answers become a threshold: a rule that lets the agent act alone on that class, or one that holds it. The owner signs the rule they already wrote by hand.

source 3   other fleets

Suggested before the incident is yours

Patterns, never content: a prompt-injection wave circulating this week, a model version that got worse at a task, a tool definition that changed underneath everyone. owno proposes the rule; you decide whether it applies.

source 4   model and tool changes

A new version is a new agent

A model upgrade or a changed tool is replayed against the agent’s own tests before a single production call. What passed stays; what regressed is held, with the case that caught it attached.

The rulebook growing, three of these four sources drawn, is on the home page. The rulebook grows →

§03.1

Nothing is enforced that nobody chose.

§04   the proof

04 / 04

Whether it helped is measured, not asserted.

The method is stated in advance and the result is published per customer, including when the result is no difference.

scene 10   two arms

random · by action · same instrument

Scene 10 · each action is assigned at random to one arm; both arms are recorded by the same instrument. Nulls count.owno.ai

how

Randomised by action, not by agent or by week

The two arms face the same traffic. The arm without enforcement is recorded by the same instrument, which is what makes the difference readable. A null is published with the same weight as a win.

the second loop

What we learn deploying becomes a package

Every implementation is meant to leave adapters, tests and a runbook that make the next one cheaper; the test we set ourselves is that the third deployment costs less than the first. Patterns learned in one fleet reach the others as proposed rules, never as content.

The four numbers the arms compare, cost per correct result, incidents per thousand actions, autonomy rate, releases per agent per month, are on the home page. Proof →

the first step

OWNO-WEB-1.2 · 2026-09-09

Start with one decision you can inspect.

A readiness study of two to three weeks fixes the baseline the six numbers will be measured against. Already running agents? A failure map from their own cases is where recovery starts.

owno.ai  ·  São Paulo