Julien Khaleghy is the Founder and CEO of SerpApi.

​The generative AI revolution is moving quickly across countless companies and organizations. Every month, there’s a new AI agentic system, a significant language model update, a new AI coding platform or another opportunity for increased productivity and product development. There’s just one problem: Bad data can break everything.​

Without an up-to-date knowledge base, even the best large language models (LLMs) can misinterpret instructions, fail at simple tasks and create endless headaches for an organization. The solution isn’t another model update or more AI agents to check each other’s work. In my experience, the true fix is providing real-time search data for the AI system.​

AI Models Are Stuck In The Past​

Every large language model and machine learning algorithm is built upon some amount of training data material. For general-purpose models like Qwen2, Meta Llama 4 and DeepSeek-R1, the starting data is a mix of public websites, published content from partnered news agencies, books and other sources. For smaller models, there’s also a distilling process where the new model is quizzed on answers by a more complex model (sometimes one from a competitor). After the core model is complete, it might be fine-tuned for specific use cases or datasets.​

The data used to train AI models, whether it’s coming from the original source or a distillation model, has to be compiled before the training can start. This creates a knowledge cutoff, where the model has no understanding of events that happened after the training data was compiled and can’t complete tasks that require a recent context of the world.​

According to data compiled by Hao Wang, Meta’s Llama 2 model started training on data from September 2022, which was a full 10 months before its release in July 2023. As newer models are developed more rapidly, the knowledge cutoff for most AI models has been reduced. Anthropic’s Claude 4.7 Opus model is trained on data that is as recent as January 2026.​

The knowledge cutoff is still a significant problem, though. If you asked a real person to complete a task with months-old knowledge of the world and no ability to look up new information, they might not be able to do it. The same is true of an AI model, but the AI wouldn’t even have the advanced thinking and reasoning skills of a real person.​

Making The Most Of Real-Time Search Data

I founded SerpApi because I needed live search engine data in my own projects, and I knew other developers and organizations would want the same access.​ In my role, I’ve been able to identify a few key best practices for utilizing this data.

First, it’s important to structure your real-time search requests as typical search queries. Search engines already know how to provide the best answers for “GM stock price,” “iPhone 17 release date” or “population of Canada.” For most real-time search providers, those simple requests will get you the best results.

Second, consider if your search queries need to come from a certain region or location for the best answers. A web search for “coffee” in New York City will have different results than the same search in London. If you need data about local events in Berlin, a German-language search (ideally originating near Berlin) might provide more relevant results than an English-language search.

Finally, you should determine which search queries actually need live data and which queries can still be useful with hours-old (or days-old) data. A search for the score of an ongoing football game needs live data, but a search for national parks in the United States does not. For answers that aren’t constantly changing, implementing a cache can improve response times and cut costs for API usage, while still providing much better data than the knowledge cutoff in most LLMs.​

You Need Better Data For Better AI​

Many companies and organizations have the wrong idea about AI readiness. It’s not just about picking the right model, how much budget to spend on tokens or which tool to use in common workflows. I believe the most important factor of all is the source data. Without that strong foundation, everything falls apart. The data needs to be high-volume, fresh and high-quality, which you can only achieve with live search abilities.

Search data is the critical infrastructure for truth in AI systems. The world is moving too quickly, and no AI model can know everything. Model upgrades are still important, especially since they have more recent knowledge cutoffs in the core training data, but they still aren’t a complete solution. Without access to live search data, AI is just motion without intelligence.​​​​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Share.
Exit mobile version