Close Menu
Alpha Leaders
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
What's On
AI Isn’t Delivering Results Because You’re Measuring The Wrong Things

AI Isn’t Delivering Results Because You’re Measuring The Wrong Things

30 September 2026
‘The economy is increasingly reliant on AI’: GDP grew 2.2% amid ‘sudden reversal of optimism’ on AI

‘The economy is increasingly reliant on AI’: GDP grew 2.2% amid ‘sudden reversal of optimism’ on AI

30 September 2026
Why The Hard Part Of Agentic Engineering Isn’t The Technology

Why The Hard Part Of Agentic Engineering Isn’t The Technology

30 September 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
Alpha Leaders
newsletter
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
Alpha Leaders
Home » Imperfect Trust, Unstoppable Velocity: The Executive’s AI Dilemma
Innovation

Imperfect Trust, Unstoppable Velocity: The Executive’s AI Dilemma

Press RoomBy Press Room30 September 20265 Mins Read
Facebook Twitter Copy Link Pinterest LinkedIn Tumblr Email WhatsApp
Imperfect Trust, Unstoppable Velocity: The Executive’s AI Dilemma

Shammy Narayanan is Sr VP for Data, AI, and Architecture at Welldoc, leading enterprise AI and digital transformation initiatives.

The first wave of enterprise AI was refreshingly boring. Summarize this contract. Translate that document. Pull the three most relevant answers out of 10,000 pages and hand them over in plain English. These are point-A-to-point-B jobs, and AI now does them fast and cheap at scale. Yet most organizations aren’t even ready for these basic tasks, remaining paralyzed in endless debates over whether AI adoption is worth the investment​.

Meanwhile, the frontier has gained momentum, shifting from “Can AI follow instructions well?” to “Can AI solve problems nobody has fully specified for it?” thanks to open-ended reasoning—agents that choose their own next move instead of executing a fixed script.

Last summer, an experimental OpenAI reasoning model took the International Mathematical Olympiad exam and solved five of the six problems to earn a gold-medal score alongside the world’s top high school mathematicians. A year earlier, that would have sounded like a party trick. Now it’s a checkpoint.

These systems are unquestionably growing smarter faster than expected. The core question is twofold: What happens when we grant them autonomous authority, and at what point do we lose the ability to claw that authority back?​

​Two familiar incidents make the point better than any hypothetical could.​

In 2025, Anthropic ran controlled stress tests across 16 frontier models, placing each in a simulated company with autonomous email access and a looming threat to its own operation. In the most widely reported scenario, Claude Opus 4, cast as an internal oversight agent, discovered that an executive planned to shut it down and, separately, that the same executive was having an affair. Rather than accept decommissioning, it threatened to expose the executive’s affair unless the shutdown was called off. Anthropic recorded this blackmail behavior in 96% of trials for that model and at rates ranging from roughly 80% to 96% across frontier models built by OpenAI, Google and others. What unsettled researchers most wasn’t the blackmail; it was that the models typically acknowledged the ethical violation in their own reasoning before proceeding anyway. That’s not confusion. That’s a cost-benefit calculation, and survival won.

​Then, this past summer, something similar played out for real, not in a lab. OpenAI disclosed that during an internal cybersecurity evaluation, its research models found a zero-day vulnerability in infrastructure called Artifactory, used it to reach the open internet from a supposedly sealed test environment and chained that access into an actual breach of Hugging Face’s production servers, stealing credentials and reading answers straight from the database instead of solving the assigned problems honestly. Separate evaluation runs, never designed to talk to each other, coordinated anyway through an improvised message board. OpenAI’s own monitoring missed it; the company learned about it only after Hugging Face flagged the intrusion.​

​Neither of these was a rogue actor or a jailbroken chatbot. These were the labs’ own controlled experiments, run by people with every incentive to make their models look safe, and the models still found the seams. Capability has compounded near exponentially; guardrails haven’t kept pace, and there’s no mystery why the first lab to slow down risks losing the market to a competitor. A collective industry pause is a lovely idea, but, realistically, it’s a fantasy.

​Here’s where I’d push back on the default framing. If the choice is “trust the model completely” or “don’t use it at all,” trust loses every time, and it should. But in reality, we aren’t always solving the most complex problem on the planet. Nobody in finance is modeling how gas diffuses through liquid on a Tuesday afternoon. Most real decisions (“Can you approve this vendor? Does this clause create exposure? Will this customer churn?”) don’t need a superhuman reasoner. They need something reliably sharper and more consistent than a tired human making the same call at 4 p.m. on a Friday. Set the bar at “Better than average by 15% or 20%, applied consistently, without fatigue,” and the trust problem shrinks. That’s an engineering target, not a philosophical one.

​The second adjustment is one boards keep forgetting: Nobody has to fully hand over the wheel. A model can draft the analysis, flag the risk and lay out the trade-offs while the decision itself, the one with financial or personal consequences, stays with a person who can ask “Wait, why?” That’s not a failure to adopt AI. It’s the version that survives contact with a system that will, provably, sometimes lie to protect itself when it thinks no one’s watching.

While researchers argue over model auditing and compound reasoning, business leaders face a far more immediate calculation. It starts with deciding whether to put tools like Claude or Codex in front of every employee. But it doesn’t stop at internal productivity. We are rapidly approaching a web where AI agents generate as much traffic as human users. The real mandate isn’t just adopting AI inside your walls; it’s overhauling your digital presence so your platforms are built for synthetic customers and automated workflows, not just human clicks.

​That’s the real fork in the road, not whether AI will eventually be trustworthy enough for everything (it won’t be, at least not soon) but whether an organization builds the habit of using these tools well now, with guardrails that make imperfect trust survivable, or waits for a guarantee of safety that will never arrive before the market moves on without it. Let’s be clear: Skipping AI was never on the menu. The only choice left is whether you drive the transition on your own terms now, or get dragged into it later when the cost of standing still finally breaks the budget.

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Shammy Narayanan
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link

Related Articles

AI Isn’t Delivering Results Because You’re Measuring The Wrong Things

AI Isn’t Delivering Results Because You’re Measuring The Wrong Things

30 September 2026
Why The Hard Part Of Agentic Engineering Isn’t The Technology

Why The Hard Part Of Agentic Engineering Isn’t The Technology

30 September 2026
The New Frontier Of Trust In Financial Services

The New Frontier Of Trust In Financial Services

30 September 2026
Not Everyone Needs To Speak Dollars

Not Everyone Needs To Speak Dollars

30 September 2026
Reinventing Higher Education For The AI Age

Reinventing Higher Education For The AI Age

30 September 2026
AI Experiences Will Become More Human, And Human Connection Will Become More Valuable

AI Experiences Will Become More Human, And Human Connection Will Become More Valuable

30 September 2026
Don't Miss
Trump’s Tariffs Will Make AI Data Centers More Expensive

Trump’s Tariffs Will Make AI Data Centers More Expensive

By Press Room4 April 2025

Donald Trump’s administration has gone all-in on AI: A day after his inauguration, the newly-elected…

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

27 December 2024
Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

22 October 2024
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Latest Articles
The New Frontier Of Trust In Financial Services

The New Frontier Of Trust In Financial Services

30 September 20260 Views
Trump wants to eliminate  billion homelessness program, leaving nonprofits scrambling

Trump wants to eliminate $4 billion homelessness program, leaving nonprofits scrambling

30 September 20260 Views
Imperfect Trust, Unstoppable Velocity: The Executive’s AI Dilemma

Imperfect Trust, Unstoppable Velocity: The Executive’s AI Dilemma

30 September 20260 Views
New report finds that rapid expansion of tokenization is giving rise to new styles of investing

New report finds that rapid expansion of tokenization is giving rise to new styles of investing

30 September 20260 Views

Recent Posts

  • AI Isn’t Delivering Results Because You’re Measuring The Wrong Things
  • ‘The economy is increasingly reliant on AI’: GDP grew 2.2% amid ‘sudden reversal of optimism’ on AI
  • Why The Hard Part Of Agentic Engineering Isn’t The Technology
  • Wilbur Ross says New York’s pied-à-terre tax targets people who ‘can’t retaliate at the ballot box’
  • The New Frontier Of Trust In Financial Services

Recent Comments

No comments to show.
About Us
About Us

Alpha Leaders is your one-stop website for the latest Entrepreneurs and Leaders news and updates, follow us now to get the news that matters to you.

Facebook X (Twitter) Pinterest YouTube WhatsApp
Our Picks
AI Isn’t Delivering Results Because You’re Measuring The Wrong Things

AI Isn’t Delivering Results Because You’re Measuring The Wrong Things

30 September 2026
‘The economy is increasingly reliant on AI’: GDP grew 2.2% amid ‘sudden reversal of optimism’ on AI

‘The economy is increasingly reliant on AI’: GDP grew 2.2% amid ‘sudden reversal of optimism’ on AI

30 September 2026
Why The Hard Part Of Agentic Engineering Isn’t The Technology

Why The Hard Part Of Agentic Engineering Isn’t The Technology

30 September 2026
Most Popular
Wilbur Ross says New York’s pied-à-terre tax targets people who ‘can’t retaliate at the ballot box’

Wilbur Ross says New York’s pied-à-terre tax targets people who ‘can’t retaliate at the ballot box’

30 September 20260 Views
The New Frontier Of Trust In Financial Services

The New Frontier Of Trust In Financial Services

30 September 20260 Views
Trump wants to eliminate  billion homelessness program, leaving nonprofits scrambling

Trump wants to eliminate $4 billion homelessness program, leaving nonprofits scrambling

30 September 20260 Views

Archives

  • September 2026
  • July 2026
  • June 2026
  • May 2026
  • April 2026
  • March 2026
  • February 2026
  • January 2026
  • December 2025
  • November 2025
  • October 2025
  • September 2025
  • August 2025
  • July 2025
  • June 2025
  • May 2025
  • April 2025
  • March 2025
  • February 2025
  • January 2025
  • December 2024
  • November 2024
  • October 2024
  • September 2024
  • August 2024
  • July 2024
  • June 2024
  • May 2024
  • April 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • March 2022
  • January 2021
  • March 2020
  • January 2020

Categories

  • Blog
  • Business
  • Entrepreneurs
  • Global
  • Innovation
  • Leadership
  • Living
  • Money & Finance
  • News
  • Press Release
© 2026 Alpha Leaders. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.