i tests from real cases
No synthetic benchmark
The cases come from the ledger: what the agent actually did, what the owner said about it, and what came of it. Each incident adds one.
continuous improvement
part 03 of 03
The unit of value is the case joined to its business outcome, not the agent’s trace. Every correction becomes a test, every incident becomes a rule, and nothing ships that fails the agent’s own cases.
§01 the measure
01 / 04
The denominator is the original workload, with failures, escalations and repairs in it.
No fabricated percentages. The numbers enter when your baseline exists, and the method of comparison is written down before the work starts. For scale only: the Anthropic Institute’s economic scenarios (Korinek, Jones, Sacher, Cotter and McCrory, WP 2026-02, September 2026; the views are the authors’) take a gain per task of 35–57% from field trials and frontier firms as their anchor, and note that an estimate from conversation data implies about 5×. Neither is a promise about your workflow.
§01.1
The denominator is the original workload. Always.
§02 the loop
02 / 04
A specialist’s correction becomes a test case; the owner releases; the result is measured on the next cases. Not a collection of dashboards, a way of deciding.
scene 08 the agent gets better
tests · gate · autonomy
i tests from real cases
The cases come from the ledger: what the agent actually did, what the owner said about it, and what came of it. Each incident adds one.
ii a gate before every change
Prompt, tools, model: every change runs the suite first. The change that would have broken a real case never reaches production.
iii autonomy that widens with evidence
When the agent agrees with its owner on a long enough run and its tests keep passing, owno proposes a wider class it may handle alone. The owner signs. Incidents rise, and the room narrows again.
§03 the rulebook
03 / 04
A rule enters the book from one of four places. owno proposes it, replays it against last month’s real traffic so the owner sees exactly what it would have refused, and switches it on only when a named person signs.
source 1 your own incidents
A duplicate send, a refund reversed, an export nobody approved. The incident becomes a test written from the real actions; the test becomes a rule that would have refused it. The same mistake cannot happen twice.
source 2 the owner’s decisions
The approval inbox is a training set. Two hundred consistent answers become a threshold: a rule that lets the agent act alone on that class, or one that holds it. The owner signs the rule they already wrote by hand.
source 3 other fleets
Patterns, never content: a prompt-injection wave circulating this week, a model version that got worse at a task, a tool definition that changed underneath everyone. owno proposes the rule; you decide whether it applies.
source 4 model and tool changes
A model upgrade or a changed tool is replayed against the agent’s own tests before a single production call. What passed stays; what regressed is held, with the case that caught it attached.
The rulebook growing, three of these four sources drawn, is on the home page. The rulebook grows →
§03.1
Nothing is enforced that nobody chose.
§04 the proof
04 / 04
The method is stated in advance and the result is published per customer, including when the result is no difference.
scene 10 two arms
random · by action · same instrument
how
The two arms face the same traffic. The arm without enforcement is recorded by the same instrument, which is what makes the difference readable. A null is published with the same weight as a win.
the second loop
Every implementation is meant to leave adapters, tests and a runbook that make the next one cheaper; the test we set ourselves is that the third deployment costs less than the first. Patterns learned in one fleet reach the others as proposed rules, never as content.
The four numbers the arms compare, cost per correct result, incidents per thousand actions, autonomy rate, releases per agent per month, are on the home page. Proof →
the first step
OWNO-WEB-1.2 · 2026-09-09
A readiness study of two to three weeks fixes the baseline the six numbers will be measured against. Already running agents? A failure map from their own cases is where recovery starts.
owno.ai · São Paulo