← All posts

Stop Building AI Assistants. Start Automating the Handoff.

2026-10-068 min read
ai agentsenterprise aiai automationsystem design

Sit through enough enterprise AI demos, and you start noticing the same pattern.

The handoff tax - an AI answer passing through human interpretation, verification, navigation, and re-entry before it reaches the system of record, compared to an AI answer landing directly as a permissioned, auditable action

Someone asks an AI assistant about a claim, a shipment, an invoice, or a support ticket. The model answers correctly - genuinely correctly, often faster than the person who would have looked it up by hand. Everyone nods.

Then someone closes the chat and opens the real system.

That's the part nobody demos.

The human reads the answer, decides it's correct, and re-enters essentially the same information into the system where the actual work happens. The AI didn't automate the workflow. It automated the lookup, then handed the work right back to a human.

The handoff tax

I think of this as the handoff tax. Every time an AI system produces an answer that a human has to manually translate into an action somewhere else, the enterprise pays a tax on it:

AI answer
    ↓
Human interpretation
    ↓
Human verification
    ↓
Human navigation to the right system
    ↓
Human re-entry
    ↓
System of record

The AI may have saved five minutes of research. But the organization is still stuck with AI → human → software, instead of AI → controlled action → software. The five minutes saved on the lookup gets spent right back on the translation, and none of it shows up as a change the system of record can point to.

In that workflow, the human isn't really making a decision at the re-entry step. They're acting as an API. The model produces structured knowledge on one side; the system of record expects structured state on the other; the person in between is the translation layer connecting them. That's the handoff worth eliminating - not the judgment call earlier in the process, just the part where a person manually serializes a conclusion into a form the other system understands.

The human isn't always the decision-maker. Sometimes they're just the API between the model and the system of record.

That's one reason an AI pilot can look impressive in isolation while producing surprisingly little change in the underlying workflow. The bottleneck was never whether the model could answer the question. It's what happens after it does.

The top row is where most AI assistant projects stop. The bottom row is where the workflow actually starts changing - and it's the same five-minute answer, just wired into the system instead of into a chat pane.

The same pattern, three domains

This isn't an AI problem. It's an enterprise workflow problem that AI happens to be making visible faster than anything before it.

Insurance. The model says: "this claim has four prior claims in a short window."

Old:  AI flags pattern → adjuster reads it → searches claimant →
      opens review → types the reason → submits

New:  AI identifies pattern → creates review action, evidence attached →
      routes to adjuster's queue → adjuster approves or rejects

Customer support. The model says: "this customer is eligible for a $150 credit under policy X."

Old:  AI answer → agent reads it → opens the CRM →
      finds the account → fills out the credit form → submits

New:  AI answer → pre-authorized credit action, capped and logged →
      agent confirms in one click

Procurement. The model says: "this invoice doesn't match its purchase order."

Old:  AI flags mismatch → finance employee reads it → opens the ERP →
      finds the PO → manually flags the invoice

New:  AI flags mismatch → creates an exception, evidence attached →
      routes to accounts payable

Three different systems, three different industries, the same shape of fix. The AI wasn't the missing piece in any of them. The missing piece was a defined, permissioned way for the AI's conclusion to become a change in the system of record without a person acting as the wire between the two.

Reasoning and action are different problems

The interesting architectural boundary here isn't between "AI" and "not AI." It's between reasoning and action:

AI reasoning
      ↓
Structured proposal
      ↓
Policy / validation      →   is this allowed? does it need approval?
      ↓
Permissioned action
      ↓
Human approval (when required)
      ↓
System of record
      ↓
Audit / event

Read that chain carefully and it argues the opposite of "pipe whatever the model says straight into a write." The handoff worth automating is everything around the decision - proposing it in a structured form, checking it against policy, routing it to the right system. Whether a human still has to approve it before it takes effect is a separate question, one the "when not to automate" cases below answer case by case. Automating the handoff and removing the decision boundary are not the same move.

A model can be very good at the top of that chain - reading scattered context and reaching a conclusion a person would have taken longer to reach. It's not, by itself, a safe way to change a production system, and it shouldn't have to be. The fix isn't to make the model more trustworthy. It's to give its conclusion somewhere defined to land: a permissioned, auditable action, separate from the reasoning that produced it.

Platforms built for enterprise data tend to make this boundary explicit rather than leaving it implicit in a prompt. On Palantir Foundry, AIP Logic is the reasoning layer, while an Action is a separately defined, permissioned write against the ontology the rest of the system already trusts - Logic's output becomes an Action's input, and it never gets to skip the definition of what's allowed to change. That's one implementation of the reasoning/action split, not the split itself. The platform is evidence this separation is buildable in production. It isn't the point.

The unit of work isn't the answer. It's the action.

If you're building an enterprise AI system, the useful unit isn't "the model decided X." It's a structured action with enough context around it to survive an audit six months later:

{
  "action": "flag_claim_for_review",
  "object": "claim_123",
  "reason": "4 claims filed within 90 days",
  "evidence": ["claim_098", "claim_104", "claim_117"],
  "confidence": 0.91,
  "requested_by": "ai_claim_review",
  "requires_approval": true,
  "idempotency_key": "claim_123:review:2026-09-20"
}

A handoff worth automating answers each of these before it ever reaches a person:

Most "AI action" failures I've seen aren't the model reasoning badly. They're one of these seven fields missing - an action fired twice, a change nobody could explain later, a write with no defined scope. The reasoning was usually fine. The contract around the action wasn't.

When not to automate the handoff

None of this argues for wiring every model output straight into a write. Some conclusions should stay a suggestion:

The point isn't to remove the human. It's to stop making the human do the part that was never actually a judgment call - copying a conclusion from one screen into another.

What happens after the model answers

The next enterprise AI breakthrough probably won't look like a smarter chat window. It'll look like a workflow where the chat window disappears entirely: the model reads the context, policy validates the proposal, an action gets created, the right person approves it when necessary, and the system of record changes - with enough context for someone to reconstruct what happened six months later.

That's when AI stops sitting next to the enterprise's software and starts becoming part of it.

The interesting question was never whether the model could answer correctly. It's what the organization built to happen next.