Over-automate
You remove judgment that was doing real work. The failures that follow are rare, which is what makes them expensive: you stop watching, so they surface late, where someone was counting on a catch.
JANUS / A system for instrumenting how one professional and one model reach a decision together, then revising how the next decision is run.
Almost nothing in the stack reports where the line landed, or notices when it moves again. Janus sits in the path of the work and decides, case by case, how a professional and a model reach a decision together.
Between people and the AI tools they already use — in the console, not in another tab.
A copilot here, an agent there, human review on top of both. The model's accuracy gets a benchmark. The judgment on top of it gets a policy written once.
The reasonable assumption is that this gets tidied up the way every rollout does: ship it, watch where it jams, fix it next quarter.
That assumption is where this goes wrong. Iteration works when the thing you are measuring holds still between measurements, and here it does not — because deploying the AI is what moves it. People work differently around the model. Their decisions change. The value of a review changes with them. By the time the split has been measured it has shifted again, and then the model is upgraded and part of what you learned expires.
A division of work set at rollout is a photograph.
Nothing tells you when it stopped being true.
Over-automate
You remove judgment that was doing real work. The failures that follow are rare, which is what makes them expensive: you stop watching, so they surface late, where someone was counting on a catch.
Under-automate
You keep paying scarce experts to re-check work the model already had right. The productivity the AI was bought for is partly spent on checking it, and the review is a cost with nothing to show for itself.
The expensive part is not choosing wrong. It is not knowing which of the two you are currently doing. The larger the AI's share of the work, the harder that is to see.
A person agreed with the AI. That can mean four things:
Outcome / Agreement / 01
The person had independently reached the same conclusion.
Outcome / Agreement / 02
The AI corrected a mistake they were about to make.
Outcome / Agreement / 03
The AI pulled them off an answer they already had right.
Outcome / Agreement / 04
They did not think about it and clicked approve.
From outside — a log, a screen recording, an agent watching the session — all four look identical. Reading the record harder does not separate them; neither does a better watcher.
They separate when the decision is run differently: when a person's own judgment is captured before the model speaks, when the evidence is shown and the verdict withheld, when the order is varied so its effect can be read. That is an instrumentation problem, not an observation problem — and it is where Janus works.
Show the model's answer first and you have already shaped the judgment you were about to collect.
Janus is a layer between professionals and the AI tools they work with, in the path of the work — an investigation console, a review queue — not in a separate tab. It is built to sit above any one model: swap the agent underneath and the findings about the old pairing expire, while the evidence they were formed from does not — so the next pairing starts from a record rather than nothing.
What it holds is not a policy for a role. Two professionals in the same seat catch different mistakes from the same model; a rule written for the role averages that away and calls the average a standard. Janus keeps versioned, task-bounded hypotheses about each pairing — each with its evidence, its uncertainty, and a prediction the next adjudicated outcome can break.
The machinery underneath runs today. What that proves is bounded, and the bound is on the same page as the claim.
Inspect what runs →Whatever you deploy next, and whoever builds it, one question is worth asking of it: where does the line between the person and the model fall right now, and how would anyone know when it moves?
That question is the useful thing to take from this page, whether or not you ever write to me.