Close Menu
Alpha Leaders
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
What's On
The Model Got Smarter, But The System Didn’t

The Model Got Smarter, But The System Didn’t

8 October 2026
Nvidia CEO Jensen Huang says if you want to be successful, be prepared to suffer

Nvidia CEO Jensen Huang says if you want to be successful, be prepared to suffer

8 October 2026
How Modernization Can Prevent Disruptions In Patient Care

How Modernization Can Prevent Disruptions In Patient Care

8 October 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
Alpha Leaders
newsletter
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
Alpha Leaders
Home » The Model Got Smarter, But The System Didn’t
Innovation

The Model Got Smarter, But The System Didn’t

Press RoomBy Press Room8 October 20266 Mins Read
Facebook Twitter Copy Link Pinterest LinkedIn Tumblr Email WhatsApp
The Model Got Smarter, But The System Didn’t

Dr. Aditya Vikram Kashyap, AI governance and innovation leader exploring how autonomous systems reshape institutions, markets and power.

We spent the AI boom learning to rank models, but the systems built around them now decide whether the answers we get are useful, safe or simply wrong.

Someone in finance asks the company’s new assistant whether a supplier contract can be canceled without penalty. The answer, which arrives in seconds, is fluent, confident and cites a clause. It is wrong. The clause was renegotiated last spring; the assistant was reading the old version, because that is what the system handed it.

The natural verdict is that the AI got it wrong. A more useful one is that the model did exactly what it was built to do with what it was given. The failure lived in plumbing that nobody sees.

For years, the public conversation about AI has been about models: which is smartest, which company is ahead. But what you experience when you use an AI assistant is not the model alone. It is a system built around the model, including instructions, document retrieval, memory, tools, permissions, testing and rules for when a person should step in.​

Two organizations can license the same model and get opposite results. One system knows the current policy and asks before it acts. The other is a fluent liability. Much of what separates them is what each built around the model.

Nobody judges an aircraft by benchmarking its engine. The engine is indispensable, but it says little about whether a flight will arrive; that depends on everything else on board. In this case, the model is the engine. The rest of the aircraft decides whether you arrive.​

A model never meets your company or your contract. It meets a representation assembled for it: the documents a search step picked out in the order the code put them. Everything else is absent. A Stanford-led study found that models can use relevant information less reliably when it appears in the middle of a long input than near the beginning or end. Many apparent failures of intelligence are failures of information architecture.

Memory sounds like a straightforward improvement until a stored mistake continues shaping answers after the original error has been forgotten. A preprint tested five agent-memory systems by giving them a revoked policy and its replacement. None consistently enforced the revocation. When both versions were retrieved, the systems favored the old policy every time, and the agents acted on it in about four out of 10 trials, regardless of the model’s capability tier. When an AI remembers something incorrectly, what mechanism makes it forget?​​

Stakes change once a system stops answering and starts acting. In February 2026, the director of alignment at Meta’s superintelligence lab connected an agent to her inbox, asked it to suggest actions it could take and told it to do nothing without her approval. By her account, the agent’s working memory filled, an automatic step condensed the conversation and the instruction to wait didn’t survive. It started deleting emails while she typed “stop.” The underlying model hadn’t changed. A sentence seemingly went missing from its memory.

A typical benchmark asks a model a question once and grades the answer. A deployed system may handle the same kind of request hundreds or thousands of times. One customer-service benchmark found that the best model tested completed about 61% of retail tasks on a single attempt; its pass^8 score, which measures consistency across eight repeated trials, fell below 25%.

A single score cannot see that gap. NIST’s agent standards initiative starts from the observation that an agent’s usefulness is constrained by how it interacts with outside systems and data, and its work runs from interoperability standards to research on agent security and identity.

The obvious objection is that if the system around the model matters so much, why spend so much on better models?​ The premise is half right. A large jump in what models can do changes what can be built at all. Researchers showed that small gains in per-step accuracy can compound into large gains in the length of task a model can complete. Frontier capability still expands the frontier. What reaches the user, however, depends on how a particular system puts that capability to work.​

When Air Canada’s chatbot told a grieving passenger that a bereavement discount could be claimed after flying, contrary to Air Canada’s policy, the airline argued, in effect, that the chatbot was a separate legal entity responsible for its own actions. A British Columbia tribunal called that a remarkable submission and said it made no difference whether the misinformation came from a static page or a chatbot. A small-claims ruling about one refund illustrates the broader principle: a company can be held responsible for what its system tells customers, whether that information comes from a webpage or a chatbot.​​

​Set the examples side by side. The contract answer was a retrieval failure: the system fetched the wrong document. The revoked policy was a memory failure. So, on her account, was the deleted inbox, although that incident exposed a deeper problem: only a sentence in a conversation stood between the agent and the delete command.

The refund case points to a different kind of failure. At the institutional level, nobody had adequately checked what the chatbot was telling customers. In each case, the user experiences the same thing: the AI gave a wrong answer or took the wrong action. But the underlying causes are different. If an organization treats every breakdown as a model problem, it will keep reaching for the same solution: a smarter model. That will not fix a stale document, a failed memory mechanism or an instruction that disappears before an agent acts.

Return to the box. You type; an answer appears. The interface invites you to imagine that the intelligence lives in the thing that produced the words. Behind the box sits a set of decisions someone made: which documents the model gets to see, what it may remember, what it may do and when a human has to be asked. The model is still extraordinary, but its intelligence is only part of the question. The rest is the system built around it, and whether anyone is checking that system when it fails.​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Dr. Aditya Vikram Kashyap
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link

Related Articles

How Modernization Can Prevent Disruptions In Patient Care

How Modernization Can Prevent Disruptions In Patient Care

8 October 2026
The Next Wave Of AI Agents Won’t Live On Your Screen

The Next Wave Of AI Agents Won’t Live On Your Screen

8 October 2026
Why Data Quality Monitoring In The Cloud Has To Become Autonomous

Why Data Quality Monitoring In The Cloud Has To Become Autonomous

8 October 2026
Which AI Workload Goes Where? A CIO’s Guide To Hybrid AI

Which AI Workload Goes Where? A CIO’s Guide To Hybrid AI

8 October 2026
Why Identity Must Become The Enterprise Control Plane

Why Identity Must Become The Enterprise Control Plane

7 October 2026
The Case For Healthy Skepticism In The Age Of AI

The Case For Healthy Skepticism In The Age Of AI

7 October 2026
Don't Miss
Trump’s Tariffs Will Make AI Data Centers More Expensive

Trump’s Tariffs Will Make AI Data Centers More Expensive

By Press Room4 April 2025

Donald Trump’s administration has gone all-in on AI: A day after his inauguration, the newly-elected…

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

27 December 2024
Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

22 October 2024
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Latest Articles
The Next Wave Of AI Agents Won’t Live On Your Screen

The Next Wave Of AI Agents Won’t Live On Your Screen

8 October 20260 Views
Jim Farley is right about Gen Z and blue-collar work. We see firsthand how the industry is failing them

Jim Farley is right about Gen Z and blue-collar work. We see firsthand how the industry is failing them

8 October 20260 Views
Why Data Quality Monitoring In The Cloud Has To Become Autonomous

Why Data Quality Monitoring In The Cloud Has To Become Autonomous

8 October 20260 Views
Europeans are drinking less beer. Heineken’s Europe boss has a plan

Europeans are drinking less beer. Heineken’s Europe boss has a plan

8 October 20260 Views

Recent Posts

  • The Model Got Smarter, But The System Didn’t
  • Nvidia CEO Jensen Huang says if you want to be successful, be prepared to suffer
  • How Modernization Can Prevent Disruptions In Patient Care
  • USS Abraham Lincoln returns home after 265 straight days at sea and after facing food shortages
  • The Next Wave Of AI Agents Won’t Live On Your Screen

Recent Comments

No comments to show.
About Us
About Us

Alpha Leaders is your one-stop website for the latest Entrepreneurs and Leaders news and updates, follow us now to get the news that matters to you.

Facebook X (Twitter) Pinterest YouTube WhatsApp
Our Picks
The Model Got Smarter, But The System Didn’t

The Model Got Smarter, But The System Didn’t

8 October 2026
Nvidia CEO Jensen Huang says if you want to be successful, be prepared to suffer

Nvidia CEO Jensen Huang says if you want to be successful, be prepared to suffer

8 October 2026
How Modernization Can Prevent Disruptions In Patient Care

How Modernization Can Prevent Disruptions In Patient Care

8 October 2026
Most Popular
USS Abraham Lincoln returns home after 265 straight days at sea and after facing food shortages

USS Abraham Lincoln returns home after 265 straight days at sea and after facing food shortages

8 October 20260 Views
The Next Wave Of AI Agents Won’t Live On Your Screen

The Next Wave Of AI Agents Won’t Live On Your Screen

8 October 20260 Views
Jim Farley is right about Gen Z and blue-collar work. We see firsthand how the industry is failing them

Jim Farley is right about Gen Z and blue-collar work. We see firsthand how the industry is failing them

8 October 20260 Views

Archives

  • October 2026
  • September 2026
  • July 2026
  • June 2026
  • May 2026
  • April 2026
  • March 2026
  • February 2026
  • January 2026
  • December 2025
  • November 2025
  • October 2025
  • September 2025
  • August 2025
  • July 2025
  • June 2025
  • May 2025
  • April 2025
  • March 2025
  • February 2025
  • January 2025
  • December 2024
  • November 2024
  • October 2024
  • September 2024
  • August 2024
  • July 2024
  • June 2024
  • May 2024
  • April 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • March 2022
  • January 2021
  • March 2020
  • January 2020

Categories

  • Blog
  • Business
  • Entrepreneurs
  • Global
  • Innovation
  • Leadership
  • Living
  • Money & Finance
  • News
  • Press Release
© 2026 Alpha Leaders. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.