← All essays

Essay

I Am Jack’s High Cost of Delivering Measurable Value

AI can make valuable work faster, safer, and easier to measure. In a dysfunctional organization, that evidence does not necessarily earn support. It can expose the architecture, process, and authority that made the work expensive in the first place.

Published
  • work
  • organizations
  • software
  • AI
  • the recurring cost of proving that the delay was optional

I am Jack’s high cost of delivering measurable value.

I have been told to measure what I build.

Set the success line before I begin.

Set the kill line too.

Choose work worth doing.

Keep it inside the risk the organization can carry.

Use a person where the decision cannot be undone.

Run the experiment.

Show the result.

Scale what works.

Kill what does not.

This is sensible advice.

It is also advice for an organization in which evidence is allowed to matter.

That condition is doing more work than the framework admits.

In a functional organization, measurable success is information.

In a dysfunctional one, it is an accusation.

The more clearly I prove that something can be done safely in hours, the more clearly I prove that the days surrounding it were not engineering.

They were architecture.

They were queues.

They were ownership boundaries.

They were delayed decisions.

They were professional caution applied after the answer existed.

They were people adapting to conditions in which participation remained safer than decision.

I did not merely save three days.

I made three days difficult to explain.

This is the high cost of using AI.

It is also the high cost of delivering measurable value without it.

The clean model

At a recent talk on AI strategy and governance, the argument was direct.

Do not begin by asking where the company can put AI.

Begin with the business problem.

Define the result.

Run small proofs of concept.

Ask whether the result can be measured.

Ask whether a failure can be recovered.

Ask whether a human baseline exists.

Three yeses make the experiment worth trying.

Use governance to hold it inside acceptable risk.

Give each one an owner.

Measure it.

Establish the number that means scale.

Establish the number that means stop.

Most organizations keep weak projects alive because nobody defines the result that would require killing them.

The successful organizations are not necessarily using better AI.

They are more disciplined about deciding what matters and ending what does not.

This is rational.

It is more rational than the usual enterprise plan in which an executive encounters a demonstration, says the word transformation, buys licenses for everyone, and waits for operating profit to become conversational.

The model also includes good technical instincts.

Least privilege.

Sensitive data kept outside the model’s reach.

Humans reviewing irreversible decisions.

AI proposing.

A person approving the action that cannot be taken back.

Cross-checks drawing human attention to the unusual case instead of requiring a ceremonial approval for every ordinary one.

All of this can work.

The missing question is not whether the model is correct.

The missing question is whether the organization wants the result more than it wants its current arrangement.

The hidden prerequisite

A scale line assumes that success can overrule hierarchy.

A kill line assumes that failure can overrule sponsorship.

A metric assumes that the people reading it are more loyal to the stated purpose than to the decisions, systems, roles, and reputations the metric may embarrass.

Governance assumes that risk is the thing being governed.

Sometimes it is.

Sometimes governance is the respectable language through which authority delays a result it cannot openly oppose.

The proposal requires another review.

The review requires another artifact.

The artifact reveals another question.

The question expands the scope.

The expanded scope invalidates the measurement.

The missing measurement proves the experiment was insufficiently disciplined.

The process has protected the organization from the danger of learning anything that would require it to change.

Every step can be defended.

That is what makes the inversion durable.

Nobody has to say:

I do not want this to succeed because its success would reduce my authority, expose my earlier decision, or make the existing pace look unnecessary.

They can say:

We need alignment.

We need consistency.

We need to think about scale.

We need a more representative test.

We should not optimize prematurely.

We need to make sure everyone is comfortable.

The system can oppose the evidence without ever opposing value.

Value remains on the sign outside.

Any size, any seat

The presentation ended with an encouraging idea.

Any size.

Any seat.

Start with the next thing you build.

A person in almost any seat can run an experiment.

Not every person can survive the result.

The fractional executive arrives with a mandate.

Someone above the conflict has already decided that the organization should hear an inconvenient sentence.

The engagement has a beginning, a scope, an invoice, and an exit.

The outsider may be allowed to call a project waste because identifying waste is part of the purchased service.

The employee who has lived inside that waste is expected to collaborate with it.

The consultant can recommend a kill line.

The employee may become the thing killed.

This does not make the framework dishonest.

It reveals the cost hidden inside the phrase any seat.

Authority is not evenly distributed because enthusiasm is evenly available.

A person can measure local work from any seat.

The person cannot make the organization accept what the measurement means.

Evidence is not neutral

Suppose I use AI to coordinate a change across three repositories.

The change requires understanding contracts, tracing behavior, comparing implementations, updating code, checking tests, documenting decisions, and keeping each repository consistent with the others.

I can use AI to inspect the field quickly.

I can ask it to surface contradictions.

I can make it produce a plan that exposes assumptions before the first edit.

I can use it to keep the implementation, tests, and documentation moving together.

I can review every result.

I can finish safely in hours what the current delivery path would allow to occupy several days.

Now I am supposed to measure the value.

I saved three days.

Did I?

In a well-designed system, I might have completed the change in hours without AI.

The three days were not all coding time.

They were the cost of decomposition beyond the needs of the product.

They were separate pull requests whose value could not be tested until other pull requests moved.

They were deployments required to validate behavior that should have been available locally.

They were queues between people who each owned one fragment of the same outcome.

They were the time required to make the organization temporarily resemble the system it had divided.

AI did not create three days of engineering value.

AI helped me cross three days of organizational distance.

To report the time honestly is to report the architecture.

To report the architecture is to report the people and decisions attached to it.

The metric arrives carrying witnesses.

The bandage and the bridge

The official story says AI accelerates delivery.

The visible result supports the story.

A complex change crossed several repositories in hours.

The demonstration looks excellent.

Then someone asks what exactly became faster.

The answer is awkward.

AI helped reconstruct context that the architecture had separated.

AI helped coordinate edits that the repository boundaries forced apart.

AI helped preserve consistency across contracts that could not be changed together.

AI helped document a path through a system whose design made the path expensive to hold in one human mind.

The tool is valuable.

The measurable value is partly evidence that the surrounding design is not.

I am putting bandages on a missing bridge.

The bandage is sophisticated.

It has tests.

It has a rollout plan.

It has a documented rationale.

It has an executive summary describing how much faster pedestrians now reach the opposite cliff.

The bridge remains missing.

This creates an accounting problem.

If I claim the full time savings, I validate AI by normalizing the dysfunction that made the savings possible.

If I explain the dysfunction, I convert the success story into an architectural criticism.

If I say the same work should not have cost three days in the first place, the measurable AI win becomes smaller.

If I say AI made the current system survivable, the system can use my survival as evidence that repair is unnecessary.

The workaround proves the platform works.

The platform proves the workaround should be scaled.

Soon the organization has an AI strategy for applying better bandages.

The competence tax

This did not begin with AI.

AI changes the speed and visibility of the pattern.

The pattern itself is old.

In 2006, I joined a company and inherited a VB6 application that processed XML through APIs and databases.

It was not one application in one place.

It was thirty instances running across two application servers.

The servers crashed almost every day.

I wore a pager.

Several times a week, I called the infrastructure group and asked someone to restart them.

After three months at the company, still somewhat junior and working without documentation, I rewrote the system.

I reconstructed what it did.

I reconstructed what it appeared to mean.

I replaced the thirty VB6 instances with a .NET 2.0 Windows Service.

The design was not sacred.

There was a Strategy pattern in an odd place, because young developers occasionally discover a pattern and feel a civic obligation to provide it housing.

The replacement was faster.

It returned one application server to the company.

It did not crash the remaining server.

It did not fail.

The pager became quieter.

My manager had written the original VB6 applications.

The success was therefore not neutral.

It was also a comparison.

This was measurable value.

The organization measured it immediately.

It concluded that nine months of work could now be assigned to me with six weeks remaining.

Competence was not rewarded.

It was consumed.

Nine months in six weeks

Another developer had spent seven and a half months on a large integration project.

The project exchanged data through APIs, stored and transformed it through databases, and provided a Windows user interface.

It required a great deal of SQL.

The previous developer believed everything could be written in SQL.

He had spent most of the schedule demonstrating the limits of the proposition.

My manager had seen the service rewrite.

He was convinced I was good.

This conviction did not produce more authority, a realistic schedule, protection from inherited risk, or even the radical management practice of checking whether the project had progressed during its first seven and a half months.

It produced a deadline.

Complete the nine-month project in one and a half months.

So I did.

The work included sudden-death coding at the corporate office with a Java 1.4 team that regarded my presence as an architectural defect.

One person told me to leave.

I stayed long enough to finish the system.

The previous developer left the company shortly afterward.

The project went to launch.

I had written roughly fifty pages of deployment notes.

I sat beside the deployment engineer and checked all of my pieces because I understood the bargain.

If the launch succeeded, the system would have succeeded.

If the launch failed, a person would be required.

I was the most available person.

The wrong platform

The integration did not work.

For two and a half days, the corporate team blamed my components.

My manager sided with them.

I searched my code for failure points.

I exercised it more than a gym instructor.

I traced inputs.

I traced outputs.

I tested assumptions.

I tested the tests.

I tried to prove that the visible failure belonged to me because everyone with more organizational authority had already reached that conclusion.

After two and a half days, the corporate team admitted that its testing had occurred on one version of an Oracle platform and production used the next version.

The two environments handled XML differently.

The failure was on their side of the exchange.

They could not correct it immediately.

They asked me to compensate for it on mine.

I did.

In half a day, I changed my component to tolerate the difference, redeployed it, and made the integration work.

The project had consumed seven and a half months without adequate follow-up.

I completed it in six weeks.

I absorbed an environment mismatch I did not create.

I repaired the other side’s XML problem in half a day.

I received no good job.

No thank you.

No acknowledgment that the person blamed for the failure had become the recovery plan.

The organization accepted the value.

It rejected the implication.

The retrospective

The following week, there was a retrospective.

The room contained me, the business analyst, my manager, and a friend of my manager who did not appear to have an official role in the project.

The friend was a former nun.

This detail sounds invented because reality occasionally hires an editor.

I suggested that perhaps the organization should follow up with developers more often about their status.

I was referring to the seven and a half months during which an important project had produced mostly SQL and no alarm proportionate to the approaching deadline.

The suggestion was modest.

The measurement was not.

Seven and a half months had passed.

Six weeks had remained.

The difference existed in the room whether I named it or not.

My manager became angry.

He yelled.

He threw objects and chairs.

The former nun attempted to translate my sentence back into harmless management advice.

He is only saying that we should follow up, she explained.

I left the room.

A few weeks later, I left the company.

The retrospective had asked how the project could improve.

The room could tolerate the completed project.

It could tolerate the emergency repair.

It could tolerate the recovered launch.

It could not tolerate a sentence connecting those results to management.

This was measurable value without AI.

The metric was still perceived as violence.

The furniture merely completed the metaphor.

What AI changes

AI does not invent the competence tax.

It changes its exchange rate.

A capable developer can now investigate more broadly, compare more evidence, hold more context, produce alternatives, document reasoning, and implement a bounded solution faster than before.

This can improve quality.

It can also increase the contrast between the time required by the work and the time required by the organization.

Before AI, unusual speed could be attributed to experience, concentration, familiarity, or a reckless disregard for the meeting calendar.

With AI, the work can leave a detailed trail.

The plan exists.

The assumptions are listed.

The risks are named.

The tests are present.

The changes are coordinated.

The result is reproducible enough to discuss.

The evidence becomes harder to dismiss as heroics.

It also becomes easier to dismiss as the machine.

If the result is good, AI did it.

If the result is flawed, I used AI irresponsibly.

If the result is fast, the estimate was too high.

If the result reveals architectural waste, I am not being collaborative.

If I document the waste, I am overexplaining.

If I do not document it, the change lacks sufficient rationale.

AI creates more ways to produce proof.

The organization creates more ways to classify proof as a personality problem.

The proof trap

I am told to show measurable evidence of success.

Suppose I do.

I show that a change expected to occupy several days was completed in hours.

The first response is not necessarily curiosity.

It may be disbelief.

The scope must have been smaller than expected.

The work must be incomplete.

The tests must be insufficient.

The approach must not scale.

The developer must have moved too quickly.

The AI must have hidden something.

A successful experiment is examined for the failure required to restore the expected hierarchy.

Suppose the evidence survives.

Then the result becomes the new baseline.

The next change receives the compressed estimate without the conditions that made the first one possible.

The unusual effort becomes ordinary capacity.

The saved time is removed from the schedule before it can become thought, recovery, design, or life.

Productivity does not return time to the worker.

It returns expectation to the system.

Suppose I explain that the speed came from experience, careful AI use, prior investigation, and an architecture whose friction I temporarily overcame.

Now I have made the result socially expensive.

I have said that the architecture is costly.

I have said that the queue is not the work.

I have said that some accepted dependencies are avoidable.

I have said that the organization’s normal delivery time contains waste.

I have not merely demonstrated AI.

I have demonstrated a counterfactual organization.

That is the threat.

The more competent I become

The more competent I become, the less I can show it.

This sentence sounds like false modesty until the incentives are visible.

If I show only the completed work, the system sees a task finished.

It does not see the reasoning, risk reduction, prevented failures, recovered context, or compressed coordination.

AI appears to add little value.

If I show the full process, the system sees how much of its normal process was unnecessary.

AI appears to add value by exposing organizational waste.

If I measure the saved time, I establish a new expectation.

If I measure the prevented risk, I identify the decisions that created it.

If I measure the quality improvement, I invite a comparison with the approved work around me.

If I measure nothing, I become another employee claiming that AI feels helpful.

The tool asks me to make my competence legible.

The hierarchy asks me to make it nonthreatening.

Those requirements are not always compatible.

The safest demonstration is therefore a small one.

A summary.

A test.

A document.

A harmless automation that saves enough time to support the strategy but not enough to question the operating model.

The organization wants measurable AI value in quantities that do not alter anybody’s relative importance.

This is transformation calibrated to preserve the seating chart.

The high cost of using AI

The cost of AI is not only the subscription.

It is not only inference, tokens, security review, privacy, hallucination, model drift, or the additional scrutiny required when generated code enters production.

Those costs are real.

The hidden cost is that AI can make a competent person’s private understanding visible at institutional speed.

It can turn a suspicion into a diagram.

A diagram into a plan.

A plan into working code.

Working code into a comparison.

The comparison into a question.

Why did this take so long before?

Why are these pieces separate?

Why does this require three approvals?

Why can nobody test the complete behavior locally?

Why did we replace the old system if the new system routes users back to it?

Why are we measuring the experiment but not the environment that determines whether the experiment can matter?

AI lowers the cost of producing the question.

It does not lower the cost of being the person who asks it.

Measure it. Watch it. Kill it.

The advice remains good.

Measure what you build.

Watch what it does.

Kill what does not work.

But the framework must include the organization itself.

Can success be acknowledged when it contradicts a leader’s prior decision?

Can a project be killed when its sponsor remains powerful?

Can an employee define a metric that exposes architectural or process waste without being recast as disloyal?

Can governance distinguish risk from discomfort?

Can the person doing the work speak plainly about the result?

Can the organization return saved time as capacity rather than immediately converting it into a tighter deadline?

Can the metric change authority, or may it only decorate authority’s existing conclusion?

Without those conditions, strategy and governance become another layer of measurement around a system that is not permitted to learn.

The experiment can succeed.

The result can be clear.

The value can be real.

The organization can still decide that the dangerous variable is the person who made it visible.

Keep the receipts

The answer is not to stop measuring.

It is not to hide competence so completely that the surrounding system becomes the only surviving account of what happened.

Keep the plan.

Keep the assumptions.

Keep the implementation notes.

Keep the before and after.

Keep the failures prevented and the constraints absorbed.

Keep the elapsed time, but keep the counterfactual honest.

Did AI accelerate the engineering?

Did it accelerate navigation through a poor architecture?

Did it reduce risk?

Did it compensate for missing ownership?

Did it make a good system better?

Did it make a bad system temporarily bearable?

These are different forms of value.

The organization may refuse the distinction.

The career should not.

Evidence that cannot safely become internal strategy can still become professional memory.

It can become a case study.

It can become an interview answer.

It can become the reason another organization recognizes what the current one has trained itself not to see.

Voice is not always rewarded.

Exit is still a measurement.

In 2006, I delivered the project.

I recovered the launch.

I named the missing follow-up.

The room threw chairs.

A few weeks later, I got another job.

That was the first time anyone acted on the evidence.

I am Jack’s high cost of delivering measurable value.

With or without AI.

I make the work faster.

The organization makes the result dangerous.

The more competent I become, the less I can show it.

The less I show it, the less value can be measured.

The dashboard remains ready.

The bridge remains missing.

Receipts

Return to the essay library