Please ensure Javascript is enabled for purposes of website accessibility
Home AI The Execution Gap: When AI Agents Must Act on Real Devices

The Execution Gap: When AI Agents Must Act on Real Devices

headline for The Execution Gap: When AI Agents Must Act on Real Devices

Most agentic AI demonstrations run in a sandbox. The agent reads a ticket, reasons about it, calls a tool, and reports back, all inside an environment built to keep it safe. The interesting question for enterprise IT starts one step later, when the agent has to change something on a laptop in a branch office, a server in a colocation rack, or a point-of-sale terminal that hasn’t been rebooted since March. That is where the gap between planning and doing becomes visible, and it is not primarily a model capability problem.

Key Takeaways

  • Agentic AI operates in a sandbox but faces challenges when executing changes on real devices, primarily due to identity, reachability, capability boundaries, and verification requirements.
  • The transition from planning to execution highlights infrastructure concerns that most prototypes skip, leading to complications in real-world applications.
  • Traditional human approval processes become ineffective at scale, which risks turning thorough reviews into mere formalities and adding latency.
  • The security community has identified failure modes in agentic applications that stem from execution capabilities, rather than from the reasoning model itself.
  • Without a well-designed infrastructure for agentic AI, organizations may experience more rapid mistakes instead of enhancements, as the gap between intent and execution remains unchecked.

Reading Is Cheap, Writing Is Not

ai agents working on team

Language models are strong at the parts of operations work that involve interpretation. Summarizing an incident thread, correlating a symptom with a recent change, drafting a plan of action, picking the right runbook out of forty: these are retrieval and reasoning tasks, and they are reversible. If the model gets one wrong, the cost is a wasted minute and a human correction.

Acting on a real device inverts every one of those properties. The action requires credentials that carry real authority. It requires a network path to a machine that may be behind a VPN, asleep, or on a segment the agent’s host cannot reach. It changes persistent state. And when it goes wrong, the failure is not a bad paragraph; it is a device in a worse condition than before, sometimes in a way that removes the very access needed to fix it.

What an Agent Needs Before It Can Touch Anything

Four prerequisites sit between a good plan and an executed change, and none of them are supplied by the model. The first is identity: the agent must act as some principal; that principal needs scoped permissions, and those permissions need to be auditable after the fact. The second is reachability, which in real estate means agents on endpoints, a broker, or an existing management channel, because most devices are not sitting on a public API.

The third is a capability boundary. An agent that can run arbitrary commands as an administrator is not an automation system; it is a liability with good manners. The fourth is verification: the agent must be able to confirm that the change produced the intended state, rather than confirm that the command returned exit code zero. These four are ordinary infrastructure concerns, which is exactly why they get skipped in prototypes and then dominate the timeline in production.

Approval Prompts Are Not a Safety Design

The usual answer to all of this is to keep a human in the loop and require approval for consequential actions. That works at demonstration volume and degrades badly at operational volume. An engineer approving four agent actions a day reads them. An engineer approving two hundred does not, and the approval step quietly becomes a formality that adds latency while providing no actual review.

Treating execution as a first-class engineering problem produces a different shape. Real-time endpoint execution using agentic AI for IT depends far more on what the execution layer is permitted and able to do than on how capable the reasoning model is: which actions are declared in advance, which are reversible by design, which require a second factor, and what evidence gets written down at each step. The model chooses among options. The infrastructure decides what the options are.

The Security Community Already Mapped the Failure AI Agent Modes

ai agents operating in real life

This is well-trodden ground for people who think about attack surfaces. In December 2025, the OWASP GenAI Security Project published its Top 10 for Agentic Applications, the product of more than a year of research with input from over a hundred security researchers and practitioners, and reviewed by an expert board that included representatives from NIST, the European Commission and the Alan Turing Institute.  

The project cataloged the AI agents failure modes that follow from giving agents real capability, and the three the announcement highlights are instructive: agent behavior hijacking, tool misuse and exploitation, and identity and privilege abuse. Every one of those is a property of the execution layer rather than of the model.

The oversight side has been studied empirically too. 

A 2026 primer by Samir Passi and Ranjit Singh, drawing on fieldwork inside a computational biology laboratory, named four conditions that effective oversight requires: adequate knowledge of what the system can and cannot do, sufficient observation of what it is actually doing, meaningful control over its behavior, and timely intervention when it fails. Their sharper point is that intervention only becomes meaningful when the first three are already present, so a user who can technically press stop but cannot see a divergence in time has oversight in name only. The failures they describe are mundane and expensive: altered data, deleted files, misallocated resources.

Designing for the Boring Middle AI Agents

The practical consequence is that the useful work sits below the agent, in a layer most organizations have not built. Start by writing down the actions an agent is allowed to take as an explicit, finite set, rather than granting shell access and hoping the prompt holds. Make each of those actions idempotent and reversible where the underlying system permits it, and be honest about the ones where it does not, because a disk wipe has no undo regardless of how the request was phrased.

Then instrument the gap between intent and outcome. Record what the agent decided, what it executed, what the device state was before and after, and whether the two matched. That record is what makes an incident reviewable and what eventually justifies widening the permitted set. Without it, every expansion of scope is a matter of faith.

Read Next

A few related pieces worth your time:

Where the Real Bottleneck Sits for AI Agents

The industry conversation is currently focused on reasoning quality, which is understandable, since that is the part that improves visibly with each model release. The constraint in enterprise IT is somewhere less interesting. It is whether the estate is addressable, whether changes can be reversed, whether permissions are scoped tightly enough to survive a bad decision, and whether anyone can reconstruct afterward what happened and why. 

Organizations that have done that groundwork will find agentic execution a modest step. Those that have not will discover that a smarter agent pointed at an unmanageable estate produces faster mistakes rather than fewer of them.

Subscribe

* indicates required
Previous articleWhy Embedded Software Is Critical for AI-Powered Devices
Bailey 'Bails' Thomas
Bailey Thomas is a data scientist using large databases, visualization platforms and analytical tools for predictive modeling. He has experience working for Fortune 500 and other private companies. Bailey was also a professional eSports player who played Starcraft 2 competitively across the globe. He was ranked #1 of millions of players in North and South America. He travelled across North America and Europe for notable tournaments, to include DreamHack, MLG, Red Bull Battlegrounds. Bailey has a Bachelor’s degree, where he double-majored in Business Analytics and Finance from the University of Kansas.