Management

Productivity is not mouse movement. And an AI worker has no mouse.

Measuring presence instead of deliverables was always a poor management choice, but it was a choice. Now that part of the team can be made up of AI workers — which have no mouse, no login time, no chair — the proxy simply does not exist for part of the org chart. The deliverable stops being the preferable metric and becomes the only unit that covers the entire team.

Read in: Português · Español

The case that made the confusion visible

Wells Fargo came under criticism when it was reported that the company "fired more than a dozen employees after an investigation revealed they were using devices or apps to simulate productivity on their computers through mouse movements." The report is from The Verge.

It is worth separating two readings of the episode. The first is the obvious one: people gaming a control. The second is the one that matters: there was a control that could be gamed by a cheap little gadget, and it was being used as evidence of work.

Equating productivity with mouse movement loses the essence of what productivity means. Measures like this do not capture genuine productivity — they promote a culture of surveillance that rewards superficial activity at the expense of meaningful results. And the worst effect is not the employee who cheats. It is the honest employee, who learns that the safe path is to look busy.

Productivity is a ratio, not a sign of life

At its core, productivity is the relationship between the output generated and the inputs used. It is a fraction. It requires two quantities.

If Wells Fargo treated mouse movement as an input, the logical question is: what was the corresponding output? There was none. And the input was so trivial that a simple device could produce it on its own — which is the operational definition of an input with no value.

H. James Harrington, a performance improvement expert, once said: "If you can't measure something, you can't understand it. If you can't understand it, you can't control it. If you can't control it, you can't improve it."

The quotation is often used as a license to measure everything. It says the opposite. Before you start measuring, you have to understand the purpose: why are we measuring this? Does this input have any real value? Without that question answered, measurement does not produce understanding — it produces a report.

The three ways to improve productivity

Productivity encompasses the resources employed — salaries, infrastructure, operating costs — to produce a product or service. There are three ways to improve it:

Cut costs while holding results steady
reduce the inputs without compromising the quality or the quantity produced.
Increase output at the same cost
raise the efficiency of the inputs to get more results with no additional expense.
Cut costs and increase output at the same time
the hardest approach, and the most rewarding.

Notice what the three have in common: all of them operate on a fraction with a numerator and a denominator. Anyone who measures presence has no numerator at all — and measures the denominator badly.

What changed: part of the team has no presence

Up to this point, the argument is the one from 2024, and it was already enough. What changed is that it stopped being an argument and became a physical constraint.

Since 2026, a team can include AI workers: non-human members, hired by you, with a name that shows up on the screens, a job description written by you, a closed set of tools, a work plan with deliverables they answer for, cycles, and an assessment at the end of each cycle.

Now try applying the Wells Fargo method to one of them. There is no mouse to watch. There is no login time, because there is no login. There is no chair, no lunch break, no look of concentration that we unconsciously accept as proof. The entire repertoire of presence signals has nothing to stand on.

It delivers or it does not deliver. That is the only question the machine lets you ask.

And that is why measuring by deliverable stops being a management preference. On a mixed team, any presence metric reaches only part of the people — and it is precisely the part that already resented being measured that way. A ruler that does not reach part of the team is not a conservative ruler. It is a broken ruler.

None of this, it should be said, means the AI worker is trustworthy without supervision. The thesis is the opposite: it needs supervision, and it needs the same supervision a manager already knows how to exercise. That is why the platform treats every action a worker attempts as something classified by code before it happens, with three possible routes — it runs on its own, Orion approves it, or it is held waiting for you. That is why there is an audit trail in which every classified action becomes a line — including the replies Orion decided not to give in an authorized group. And that is why a worker never assesses its own cycle: the system refuses the attempt, and the one who assesses is the manager.

The trap is back, in new clothes

You would expect the market to have gone straight to the deliverable. It did not. Activity surveillance is back — only now it points at AI.

Tokens consumed. That is cost. It is pure denominator. A team that consumes more tokens is not more productive; it is more expensive. Treating consumption as a result inverts the fraction.

Number of prompts, of calls, of runs. It is the count of how many times someone moved the mouse, under another name. It says activity happened. It does not say what came out of it.

"Hours saved." This is the most seductive and the most fragile, because it looks like a result. But it only exists if someone measured the "before" honestly — and in most companies, nobody did. What gets called hours saved is usually an estimate presented as a measurement.

Apply the Harrington test to each one: why are we measuring this, and does this input have real value? None of the three survives. All of them measure the denominator and call it performance.

Measuring cost is legitimate — as long as it is declared as cost. That is why each worker shows a measured cost badge: how much it consumed in the current month and over the last seven days, with the number of runs. And while there is no measured run, the screen says "no measured activity yet" instead of showing zero, because absence of measurement is not the same as absence of work. None of this is presented as productivity. It is the denominator, under its proper name.

The unit that works for both: the deliverable

Deliverables are the specific, tangible products, services, or results that come out of a team's work. They are the "what" that is delivered.

The practical test is crude and it works: if you can cross an item off the list and nobody outside notices a difference, it is probably a task, not a deliverable.

AspectDeliverablesTasks/Activities
LevelResult at the team levelActions at the individual level
What vs. howWhat will be producedHow it will be produced
Where it is plannedIn the team's deliverables planIn the individual work plan
Example"Customer survey conducted""Draft the survey questionnaire"

The ritual, which is the same for people and for machines

What makes the deliverable usable as a unit is not the definition. It is the chain of steps around it, and it applies equally to both kinds of worker.

Signed deliverables plan
the team commits to a portfolio of deliverables for the period. A draft does not count — the signature is the moment the scope was reviewed by someone who answers for it.
Work plan
it translates that for a person (or for an AI worker), with a percentage of effort per deliverable and the written promise of what will be done. It is signed by both parties, and the signature creates the cycles.
Calculated progress
the progress of a deliverable is not typed in. It comes from the average progress of the tasks, weighted by the weight of each one. Nobody "declares" they are at 70%.
Cycle report and assessment
the participant reports what they did, the manager responds with 1 to 5 stars and a comment, and the cycle consolidates. The AI worker writes and submits its own report like any other member — and never assesses itself.

There is a deliberate asymmetry there, and it deserves to be said out loud. To enter a work plan, a deliverable has to come from a signed and current deliverables plan. For a human, the exception goes through and the warning is recorded. For an AI worker, it is a block: the platform refuses the signature, with the list of violations. The reason is owned up front — a person knows when the exception is legitimate and answers for it; the agent does not have that judgment, and a worker nobody can hold to account is the worst possible combination.

What measuring by deliverable demands of you

None of this comes for free, and the article would be dishonest if it stopped before this section.

You have to write real deliverables. "Do the PRD" is not a deliverable; "PRD for the backlog features moved into development" is. Swapping activity for result in the title is tedious, and it is what holds up everything else.

You have to plan before you start. Planned start date, planned end date, target type, and planned target form the baseline, and it freezes when the start date arrives. After that you use the updated fields — and it is precisely the difference between planned and updated that shows how far the plan slipped.

The system does not decide for you. The total effort of a work plan is not locked at 100%: above 100.5% the number turns red as a warning, and nothing stops you from signing it that way. Sometimes the sum goes over on purpose, and the platform does not pretend to know better than you.

The clock runs even without you. The participant has 5 days for the report, the manager has 10 for the feedback, and there are 5 days to request a review or make adjustments. If the manager does not act in time, the system applies a neutral 3-star feedback, with a generated justification that says exactly that, and notifies both parties. That closes the cycle; it does not assess anyone. An absent manager is still an absent manager, and the record shows it.

And the biggest limitation of all: nothing here guarantees the AI's work is correct. Measuring by deliverable does not make the AI get it right. It makes the AI's error reach you in the same format a human's error already reached you — inside a deliverable with an owner, a deadline, and an assessment, and not hidden in an activity metric nobody knows how to read.

Harrington's question, turned toward your dashboard

Pick the number you look at first thing in the morning to know whether the team worked. Ask it: is this the output or the input?

If it is an input — hours online, messages exchanged, meetings held, tokens, prompts — it does not measure productivity. It measures consumption. And it can be replaced by a little gadget, or by a script, depending on who you are measuring.

The deliverable is the only answer left standing when the worker has no mouse.

References

The Verge (2024)
dismissal of more than a dozen Wells Fargo employees after an investigation into devices and apps that simulated activity on the computer.
H. James Harrington
performance improvement expert, author of the quotation about measuring, understanding, controlling, and improving.