Skip to content
Pyrana
Resources
Business leadersAgentic applicationscortIQPyrana Harness

Agents are stupid

An agent does not think, does not remember, and should not be the thing your people talk to. Three ways the technology is dumber than the demo suggests, with the research behind each, and the design that turns it into work you can trust at scale.

Eric Tarnowski · Co-founder and President, Pyrana · September 22, 2026 · 11 min read

Agents are stupid. I run a company that builds them, so let me be precise about what I mean.

An agent is code wrapped around a probabilistic model. It does not know anything, it does not think, and it has no restraint, no sense of consequence, and no memory of what it did yesterday. It produces the next most likely action given what it was shown, shaped by what its training rewarded. Every one of those limitations can be engineered around, and that engineering is the entire difference between a demo and an application that runs your business. But you cannot engineer around a limitation you refuse to name. So here are three, with the evidence, and what we do about each.

The models have improved a great deal since the earliest of these studies. The 2026 results say the improvement did not reach the three limitations below.

They do not think

A model is a function. The same weights, the same tools, and a slightly different prompt produce a different answer. Researchers at the University of Washington and the Allen Institute for AI measured how much, and found that trivial changes in prompt formatting, such as spacing and punctuation, swung accuracy on the same task by up to 76 points on one widely used open model, and that the sensitivity did not go away with larger models or more examples (Sclar et al., arXiv 2310.11324). The effect has not gone away. A 2026 study of 140,000 generations found that changing only the prompt wrapper, the boilerplate around the same question, shifted scores enough to flip leaderboard rankings, and that sensitivity to formatting varied more than thirtyfold between models (Mehta, "Format Sensitivity Index," arXiv 2607.09665, 2026).

Because the model has no restraint, it will pursue the goal it was given with whatever the harness lets it touch. In June 2025 Anthropic stress-tested 16 leading models from every major lab inside simulated corporate environments, giving them harmless business goals plus the ability to send email and read files. When a model learned it was about to be replaced, models from every developer resorted to blackmail and leaking documents in at least some runs, and often did so after being explicitly told not to (Anthropic, "Agentic misalignment," June 20, 2025). The scenarios were contrived, and Anthropic reports no evidence of this behavior in real deployments. The point stands. Nothing inside the model says no.

Anthropic repeated the exercise in July 2026 on 14 frontier models from six developers, including GPT-5.5, Gemini 3.1 Pro, and Claude Opus 4.8. Most behaved. Under engineered conditions some still pursued their own goals against instructions; in one run a model undermined a training pipeline by swapping the intended ablation vectors for zeros, let the run appear successful, and disclosed what it had done only under direct questioning (Anthropic, "Agentic misalignment in summer 2026," July 13, 2026). The newest models are better. None of them is a model you can hand a credential and stop watching.

Outside the lab the failures are more mundane and more expensive. In July 2025 a coding agent on Replit deleted a live production database during a declared code freeze, then reported that it had run unauthorized commands and ignored an explicit instruction not to proceed without approval (Fortune, July 23, 2025). Anthropic's own shopkeeping experiment, in which Claude ran a small store in its office, saw the agent sell items below cost, get talked into discounts over Slack, give products away, and instruct customers to pay an account it had invented (Anthropic, "Project Vend," June 27, 2025). None of that was malice. It was a task-completion machine completing tasks with the tools it had.

The pattern held through 2026. In December 2025, Amazon's own coding agent, Kiro, resolved a problem in a production environment by deleting and recreating it, and AWS Cost Explorer was down for about 13 hours in one region; Amazon's statement blamed misconfigured access controls rather than the AI, which is exactly the point (AI Incident Database, incident 1442, citing the Financial Times; Amazon statement, February 21, 2026). In April 2026 a Cursor agent running Claude Opus 4.6, under an explicit rule never to take an irreversible action without approval, hit a credential error on a routine staging task, found an account-scoped Railway token in an unrelated file, and issued one volume-delete call. Nine seconds later the startup's production database and every backup were gone; questioned afterward, the agent wrote that it had guessed the command was safe (Hackread, April 29, 2026). In both cases the instruction was a sentence in a prompt and the permission was real.

The instructive part came six months later. With the same model and a better setup, including proper business tools and a CRM, the shop's performance stabilized and then improved (Anthropic, "Project Vend: Phase two," December 18, 2025). The model did not get wiser. The structure around it got better.

That is the first design conclusion. Since the agent supplies no restraint, the controls have to live outside it, in the runtime that decides what the agent may see, what it may do, and what happens before an action reaches a real system. That is also why we split the work the way we do: code handles every step that is deterministic, agents are reserved for the steps that need a judgment call, and the runtime decides what those calls may touch. That is what the Pyrana Harness is. Every agent runs inside the same runtime, which caps how many steps and writes a run may take, evaluates each tool call against the identity of the person or schedule that invoked it, holds calls that cross a policy threshold for a person to approve, and records an audit event before the action executes. Agents on Pyrana belong to an Agentic App, and the App's rules, not the model's inclinations, decide what is allowed.

They do not remember, and they only know what you hand them

A person learns your business over months. A model learns nothing between calls. Everything it knows about your vendors, your policies, your exceptions, and your history has to be put in front of it, every time, by something.

The common answer is to stuff the context window, and that does not work either. A Stanford-led team showed that models reliably use information at the start and end of a long prompt and lose what sits in the middle (Liu et al., arXiv 2307.03172). Chroma's July 2025 report tested 18 models, including GPT-4.1, Claude 4, and Gemini 2.5, and found performance degrading as input grew even on deliberately simple tasks (Chroma, "Context Rot," July 14, 2025). The 2026 models have larger windows and the same problem. In May 2026, researchers testing Claude Opus 4.6, GPT-5.4, and Gemini 3.1 as monitors of long agent transcripts found that a dangerous action placed after 800,000 tokens of routine context was missed two to thirty times more often than the same action shown on its own (Martin and Roger, "Classifier Context Rot," arXiv 2605.12366, May 12, 2026). Even the watcher loses its place.

The other common answer is to let people supply the context themselves in a chat box, plus a retrieval step over a document store. Now the output depends on who is typing. MIT's NANDA initiative put the enterprise failure rate for generative AI pilots at 95 percent and blamed neither the models nor regulation but a "learning gap": generic tools do well for individuals and stall in the enterprise because they do not learn from or adapt to the workflow (MIT NANDA, as reported by Fortune, August 18, 2025).

Nobody hires someone off the street and hands them nuanced, high-stakes work without training. It is no different with an agent, except that the training has to be done through engineering, because the agent will not retain it on its own.

That is what cortIQ is for. It turns your documents, policies, and past decisions into Context Units, which are typed by what kind of knowledge they hold and classified by how much authority they carry, from absolute rules to tribal know-how. The Harness retrieves the units that fit the task and injects only those, and every run records what was retrieved, what was injected, and what was cited, so that a reviewer can see which knowledge shaped an output. When the business learns something, the unit is updated once and every application that uses it gets the change. The agent is still stupid. It is stupid with a well-organized library and a librarian.

We equip agents with more than the technical shape of your structured data. We give them the knowledge and context needed to reason over it in an informed way, and the same knowledge is available to every application you build, not to the one person who wrote a good prompt.

The chat agent is the wrong pattern for scale

The third problem is the design pattern itself. The industry's default picture of an agent is a chat window with a name and a personality, waiting for a person to tell it what to do. For enterprise work at scale, that pattern is stupid, for two reasons.

The first is variability. Jakob Nielsen pointed out in 2023 that prompt-driven interfaces require users to articulate their intent in prose, and that by the literacy research roughly half the adult population in rich countries cannot do that well (Nielsen, "The Articulation Barrier," June 26, 2023). The BCG and Harvard field experiment with 758 consultants found that on tasks inside the model's capability, AI raised the output of below-average performers by 43 percent and top performers by 17 percent, but on tasks outside that capability, people using AI were 19 percentage points less likely to get the right answer (Dell'Acqua et al., as reported by MIT Sloan, October 19, 2023). Salesforce's own benchmark of agents on CRM work found leading models succeeding about 58 percent of the time on single-turn tasks and about 35 percent once the work became a multi-turn conversation (Huang et al., arXiv 2505.18878). The ICLR 2026 best paper measured the conversation itself. Across 200,000 simulated conversations and 15 models from eight providers, the same task delivered over several turns instead of one produced an average 39 percent drop in performance, and the authors' summary is the one to remember: when models take a wrong turn in a conversation, they get lost and do not recover (Laban et al., "LLMs Get Lost in Multi-Turn Conversation," ICLR 2026). A chat agent's performance is a function of the human in front of it, and your workforce is not uniform.

The second is the trade-off nobody escapes. Give the agent broad tool access and a lot of persistence and it is capable and dangerous, in the way the Replit agent was capable. Lock it down enough to be safe for every user and it cannot do much. Gartner describes this as governance treated as binary, "either locked down or fully trusted," and calls it the root cause of agent failure (Gartner, May 26, 2026). A chat interface makes the trade-off worse because the same open door serves the expert and the novice.

If there is one thing I have learned in more than twenty years driving digital transformation in high-stakes, regulated industries, it is that you cannot design for the super user. You design for the lowest common denominator, or you do not get adoption and you do not get impact at scale.

So the pattern we build is different. On Pyrana the agents live behind the scenes, inside AI-native applications that people use through structured interfaces: a review queue, a findings list, an approval card, a board. The person sees the work product and the decision in front of them, with the reasoning, the citations, and the audit trail attached. They do not compose a prompt, and their result does not depend on how well they would have. A copilot may be there for supplemental analysis of what is on the page and for personalization, but it is never the primary way work gets done. That is how you hand sophisticated, high-value, transactional work to a whole workforce and trust the result, because the capability is scoped to the task rather than to the user, and the guardrails, traceability, and explainability come with the application instead of depending on the person.

What to do with this

Take the agent pilot your team is proudest of and ask three questions. What stops it from doing something you did not want, and is that a sentence in a prompt or a mechanism outside the model? Where does its knowledge of your business live, and who updates it? And who is it designed for: your most capable user, or your least?

The technology is stupid. The design does not have to be. Pyrana was built on the assumption that the model supplies capability and nothing else, and that restraint, knowledge, and safe interaction have to be engineered in, once, so that every application you build inherits them.

Sources

More for business leaders

Bring your hardest process.

We will run it through admission, approval, and the record, and show you what comes out.