Klaudia Zaika is the CEO of Apriorit, a software development company that provides engineering services globally to tech companies.
Predictive cybersecurity is a promising concept. Many cybersecurity vendors my company works with want to utilize big data and artificial intelligence to integrate defensive measures that can predict and mitigate incidents before they even happen.
Yet if implemented the wrong way, these measures can cost you business opportunities and waste both time and money. Based on my experience, I want to share why predictive cybersecurity initiatives fail and what you can do to change that.
AI And Big Data Lead The Shift From Reacting To Predicting
Traditional cybersecurity products are reactive by nature. They only start working when there are clear, recognizable indicators of system or account compromise.
With the integration of AI and big data, it becomes possible to shift that nature from reactive to predictive.
Predictive cybersecurity identifies connections between seemingly unrelated signals across systems: authentication logs, endpoint telemetry, network flows, identity data and even business context. These identified connections help forecast likely attack paths and flag risky behaviors, enabling security teams to strengthen defenses before any actual damage can occur.
The earlier you want to catch a cyberattack, the more your defenses depend on data. But this dependency is sometimes unforgiving.
Data Problems Behind Failed Predictions
Back in 2025, Gartner predicted that through 2026, organizations would abandon 60% of AI projects lacking AI-ready data. We’re still about to see how accurate their prediction was, although judging by the projects my company has worked on so far, I can say that this estimate isn’t far from the truth.
With enough quality data, any capable AI model has a chance to succeed in predictive analysis and incident detection. Without it, even the best model fails.
The biggest challenge is to meet both the “enough” and the “quality” criteria at once.
My team worked on projects where data would be collected across several systems: SIEM, SOAR, DLP and so on. “Having enough data” was not an issue for these projects; the quality of that data was.
If your predictive system runs on poor data, for example, it will impact your customers. Inaccurate predictions are dangerous: Too many false negatives create illusory confidence for leadership, while too many false positives cause alert fatigue for analysts.
Not to mention that security logs your product will be processing are typically collected for compliance and forensics. They have short retention windows, inconsistent formats and missing fields, leaving your model with incomplete, fragmented data to learn from. The problem is that to an AI model, isolated fragments usually look like noise and prevent the model from detecting the real pattern behind them.
On top of that, cultural context matters. A model trained on the behavior of one region, industry or user population is at high risk of misjudging anything that deviates from what it considers the norm. But behavioral baselines carry a lot of organizational and cultural context: from typical working hours and travel patterns to preferred tooling. And a model that has never seen a particular organization’s unique context may easily flag normal activity as anomalous while overlooking real threats.
This doesn’t even account for the fact that regulatory compliance may put extra constraints on dataset updates. Privacy regulations such as the GDPR, CCPA and the EU AI Act limit what telemetry you can retain and share for training. In regulated or certified environments, refreshing datasets and models is a controlled change that may even require re-auditing your product after the update.
Last but not least, it’s important to note that things can also work the opposite way. My team has had projects with perfectly annotated and stabilized data that was so limited that their model couldn’t perform accurately at scale.
Accurate Predictions Demand Quality Data
The key to increasing the accuracy of cybersecurity predictions is not getting more telemetry data but improving the quality and connections of the data you already have. I want to share some of the practices my team uses for this:
• Design for siloed, fragmented stacks. With telemetry data coming from different sources and in different formats, cross-system correlation becomes impossible. Consolidating security telemetry in a data lakehouse, for instance, gives your models a unified, queryable foundation for making predictions.
• Audit training data for relevance. Before concluding that the model underperforms, verify the quality, freshness and diversity of its training data. In my experience, an “underperforming” model turns out to be a data problem far more often than an architecture problem.
• Use synthetic data to close the gaps. Rare attack patterns and privacy-restricted records can be represented using synthetic samples, allowing your team to maintain the quality and diversity of the training dataset without crossing compliance boundaries.
• Retrain and fine-tune on current, relevant and, most importantly, diverse data. Attacker behavior changes quickly enough that a static model quietly drifts out of touch. Just keep in mind that if the affected functionality falls within a Common Criteria-certified scope, retraining may be classified as a major change that requires reevaluation of that functionality, so you may want to plan the time and budget for the impact analysis.
• Teach your model to look at the bigger picture. Never let your model learn only from fragmented data. A single anomaly may look harmless in isolation. The same anomaly, combined with the larger context of a person’s or entity’s behavior over time, may reveal an attack in progress. Feed your models entity-level context, so each event is weighed against the entity’s history.
The accuracy of your cybersecurity predictions will always be proportional to the quality and connectedness of the data behind them.
If you invest in data integration, dataset auditing and regular retraining with the same seriousness you apply to selecting AI models, your predictive cybersecurity solution can actually deliver on its promise to give your customers the knowledge and vision necessary to resolve possible incidents before they even happen.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?








