Where Should Your Automation Be Allowed to Fail?

gpcarracedoUncategorized

Operational Excellence

Where Should Your Automation Be Allowed to Fail?

By Gabriel Pastrana·September 5, 2026·4 min read·Issue #57

Kenco is tripling the size of its Innovation Lab in Chattanooga, from 10,000 to 30,000 square feet. The facility is designed to approximate warehouse conditions, remain vendor-neutral, and accommodate larger and more complex automation systems.

The expansion itself is not the important part.

What caught my attention is where Kenco is choosing to discover whether automation works: before the technology reaches a live operation.

That distinction becomes more important as automation programs combine robotics, controls, orchestration software, sensing, and people into one operating process.

The individual technologies can all work—and the operation can still fail.

The Failure Surface Is Moving Between the Machines

Most automation validation starts with the component.

Can the robot achieve the required cycle time? Can the AMR complete its missions? Does perception recognize the required objects? Can the conveyor maintain throughput? Does the control system respond within specification?

Those are necessary questions. They are no longer sufficient.

I have seen multiple situations where the individual technologies were proven and performed as expected, but the integration of several technologies within an end-to-end operating process failed.

The problem was not necessarily the machine. It was what happened between the machines.

One system releases work faster than another can absorb it. An AMR fleet behaves correctly but creates congestion around a downstream process. A robotic cell achieves its nominal rate while exceptions accumulate outside the automation. Two systems each maintain valid internal states but disagree about what the operation should do next.

Every component can pass acceptance testing while the process misses its objective.

That is why Kenco’s expanded lab is interesting. The company says the facility will allow customers and OEMs to test automation in an environment resembling real warehouse conditions before committing it to operations. The significance is not simply access to more robots. It is the ability to test interactions without using a production site as the primary integration environment.

The same pattern is visible in physical AI.

Samsung SDS reported this week that it has been working with Samsung affiliates and robotics companies to validate physical-AI use cases tied to manufacturing, shipbuilding, plant, and precision-manipulation requirements. The company says validated applications will next move toward field demonstrations.

That is still a long way from proving repeatable production economics, but the sequencing matters: use-case validation precedes field deployment.

Caterpillar also announced a collaboration with FieldAI around physical AI, robotics, autonomy, and digital twins. There are no production-scale operating results in the announcement, but it points toward an even harder version of the same problem. As autonomy moves into less structured industrial environments, the number of operating conditions that cannot be represented by a simple component-level acceptance test grows.

Automation teams therefore need to ask a different question. Not only, “Has the technology been proven?”

But “Has the operating system around the technology been proven?”

The Failure Budget Test

Before approving an automation system for production, I would force the team to create a Failure Budget.

The objective is not to eliminate every possible failure before go-live. That is unrealistic.

The objective is to decide deliberately where each class of failure is allowed to be discovered.

I would test five layers:

  • Component failure. What happens when each machine faults, misses a pick, loses localization, produces bad data, or becomes unavailable?
  • Interface failure. What happens when two functioning systems disagree about state, timing, priority, capacity, or ownership?
  • Flow failure. What happens upstream and downstream when one subsystem degrades to 80% of expected performance rather than stopping completely?
  • Recovery failure. When the automated process breaks, who identifies the problem, who owns recovery, what information do they receive, and how long does restoration take?
  • Scale failure. Which assumptions change when volume increases, SKU variability grows, more robots enter the fleet, another vendor is added, or the architecture is replicated at another site?

Then add one column beside every failure mode:

Where will we discover this?

Lab. Simulation. Integrated test. Pilot. Production. Peak. Second site.

That classification can expose uncomfortable assumptions.

If a critical interaction has never been tested before production, the organization has effectively decided that production is its integration lab. If recovery behavior has never been exercised under degraded conditions, operators are being asked to invent the recovery process during an outage. If the architecture only works because one engineer understands the dependencies between four vendors, repeat deployment has not actually been engineered.

That is not automatically a reason to stop the project.

It is a reason to price the risk correctly.

Move Failure Earlier

The goal of pre-production testing is not to prove that automation cannot fail. It is to make failure cheaper.

A fault discovered in a controlled environment consumes engineering time. The same fault discovered during ramp may consume engineering time, operations capacity, vendor escalation, management attention, customer service, and schedule margin simultaneously.

For companies building larger automation portfolios, a permanent integration environment may therefore deserve to be treated as infrastructure rather than R&D overhead.

The relevant asset is not the lab itself. It is the organization’s ability to reproduce workflows, introduce failures, test recovery, measure interactions, and carry those lessons into the next deployment.

Before approving the next automation program, I would ask one question:

Which failures are we paying to discover before production—and which ones are we quietly asking the operation to discover for us?

Also on my radar

Nissan is expanding AMRs in its Tennessee manufacturing operation. The interesting part is not another AMR deployment, but the coordination problem: the robots transport heavy components, reroute, communicate task progress, and can request backup. It is another reminder that fleet behavior and operating integration increasingly matter as much as individual vehicle capability.

Anthropic is pushing toward a common interface between AI agents and physical equipment. Its Model Hardware Standard preview is intended to give AI agents a consistent way to interact with different machines and instruments. The operator question is whether emerging standards can actually reduce integration effort without creating a new abstraction layer that still has to be validated against physical reality.

China’s industrial-robot scale deserves more attention than its humanoid demonstrations. The Financial Times highlighted the less spectacular part of China’s robotics story this week: large-scale adoption of purpose-built industrial robots. For operators evaluating physical AI, the comparison is useful. General-purpose capability may attract attention, but specialized automation continues to set the benchmark on deployment scale and economics.

If this question is active in your organization, reply and tell me where the operating reality differs from the technology narrative.

Smart Automation

A four-minute weekly newsletter on automation, AI, intralogistics, supply chain, and operational excellence.

Subscribe free