Don Murray is Cofounder & CEO of Safe Software and has spent his career helping organizations bring life to data to make better decisions.
In the race to implement AI in operations, many organizations are looking outward and investing in new information, tools and data streams. But in my experience, the most valuable asset may not be something you need to go find. It may already be there, accumulated over years of operation, sitting dormant.
Most established companies own massive amounts of information that has been collected, stored or archived but never used. Gartner calls this dark data. In the age of AI, this data can now deliver real value that was difficult to realize before.
The Underdog Story
There’s a bias in tech that newer equals better. New databases, data lakes and open table formats get all the excitement. Dark data flips the script because it consists largely of historical data that has been stored but rarely, if ever, used.
Some of this dark data may not even be stored electronically, but might reside in hard copy form. Decades of operational data (work orders, maintenance logs, project files, client records, PDF reports, correspondence, etc.) are a treasure trove of valuable information. Most of it never reached a database, but it exists, and now it can be accessed.
The catch is that sitting on dark data and using it are two very different things. The challenge isn’t sourcing the data; it’s accessibility and integration.
The Challenge
Getting that information somewhere it could actually be used once meant either a lot of manual keying or OCR that choked the moment a document had a complex layout. Thanks to AI, it’s more accessible than ever.
Many of us have heard the joke that PDF was where data went to die. It’s tongue-in-cheek, but it wasn’t entirely wrong. Organizations were producing PDF reports like crazy because Adobe did a great job making it universal. A PDF opens anywhere in the world exactly as intended, but it was always a one-way street. Saving a report to PDF usually meant the underlying data either never reached a database or got cleared out of one once the report was filed.
Decades of institutional knowledge sits trapped in files that are human readable, but were unusable for analysis the moment they were created.
Before AI, unlocking that information meant hiring a small army of people to extract it by hand. Some industries did exactly that, because regulation or litigation left them no choice. Everyone else looked at the cost and decided the archive could wait.
One British electricity distributor holds more than a million service record cards, the oldest dating to 1907, carrying handwritten notes, sketches and early computer printouts. Pulling the installation dates off them by hand was estimated to require 19 years of effort, which is a polite way of saying it was never going to happen. The automated extraction ran in 26 hours. The model costs came to a few hundred pounds.
What changed is not that extraction got faster. It is that reading a document’s structure, tables, handwriting and meaning stopped being the part that broke.
From Legacy To Bleeding Edge
In the past, digitizing data meant designing database tables and deciding which fields and keys mattered, a project so convoluted that most organizations decided their old materials simply weren’t worth it. Today you can point AI at that material as it stands, and the archive becomes a living knowledge base. You can query it in plain language. It stops being an archive and starts being an asset.
Ingestion is just where the work begins. Now, you have to figure out whether the data is true and actually helpful. Archives are full of superseded figures, abandoned projects and information gathered under rules that may or may not apply anymore. The problem is that a knowledge base will serve all of it with the same confidence.
Three things separate a knowledge base people trust from one they eventually stop using:
• Every answer traces back to a source document.
• Dates travel with the data, so a 2004 estimate is never read as current.
• The archive is governed as tightly as your live systems.
All three take judgment from people who know the company, the data and what the answer will be used for. None of it is automatic.
Thirty years of annual reports, presented and promptly forgotten. Decades of operational data, filed away and never revisited. That information existed, technically. But practically, it didn’t. Now a single question reaches all of it, and you no longer need to know that a 1998 report exists in order to get what is in it.
This is where I’d push leaders to reframe their thinking. The technical barrier has not vanished, but it has dropped far enough that it is no longer what is stopping you. Imagination is the binding constraint now. We’ve been conditioned to treat archived work as finished rather than foundational, to assume that if something wasn’t useful at the time, it isn’t useful now. But decades of operational data almost certainly hold patterns nobody has gone looking for.
Where To Start
1. Lead With The Problem
Choose something your teams do every week, such as pricing a job, scoping a maintenance cycle or answering an RFP, and ask what history would make it better. Scope follows the decision.
2. Rank By Impact
The biggest archive is rarely the most valuable one. Score candidates on how often the question comes up and what a better answer is worth.
3. Digitize Widely, But Structure Narrowly
Most material only needs to be readable for AI to use it. Reserve fields, keys and validation for what feeds analytics or a system of record.
4. Classify Before You Index
Open, sensitive, restricted. Decide at the outset, because it is far harder to claw back once the data is live and people are relying on it.
Final Thoughts
One archive, one decision, one small team. That is enough to find out whether the value is there. The question isn’t whether the insights exist. They do. The question is whether you’re willing to go looking.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?






