Tal Lev-Ami is co-founder and CTO of image and video platform Cloudinary, which is trusted by more than 13,000 brands and 4 million users.
No longer just for marketing, video is now used across sales, support, training, internal communications, recruitment, customer content, compliance and AI-mediated discovery, making it a core enterprise infrastructure workload.
One visible symptom is page weight: The 2025 Web Almanac found that, as the web is rapidly moving toward video-first experiences, powered by background hero videos, looping product previews and embedded clips, the amount of video transferred on the median webpage jumped 28% in a single year, from 246 KB in 2024 to 315 KB in 2025.
Most organizations didn’t budget, architect or govern for this new reality. The CTO doesn’t need to personally own every video initiative across the business, but when video infrastructure fails, the operational, financial and reputational consequences ultimately land with technology leadership.
Why Enterprise Video Is A Physics Problem
Video files are enormous compared to images: computationally expensive to process, difficult to move globally and complex to manage at scale. The enterprise video challenge isn’t simply one of storage capacity. It requires infrastructure for ingestion, transcoding, adaptive delivery, CDN strategy, access control, rights management, retention, localization, moderation, metadata, API integration and observability. It is managing a complex, simultaneous mix of versions, formats, devices, regions, permissions, languages and delivery contexts that traditional content systems were never designed to handle.
To illustrate, take a single video asset, for example: the company story video on your website. This will require multiple renditions optimized for mobile, desktop, streaming quality, localization and regional compliance. Now multiply that across thousands or millions of videos that you publish for commercial and internal use, and you soon face operational complexity that traditional content systems were never designed to handle.
Accessibility requirements introduce another layer of difficulty that extends way beyond adding subtitles. Enterprises must manage audio descriptions, multilingual captioning, transcription accuracy, localization workflows and synchronization across every delivery environment.
Then there is user-generated content. Many enterprises are no longer dealing exclusively with professionally produced media. They are increasingly ingesting videos from customers, suppliers, partners and communities into the same catalogs. That dramatically increases variability, compliance risk, moderation demands and governance complexity.
Why Governance Can’t Be An Afterthought
This staggering level of complexity almost guarantees something will go wrong. It might be a brand violation, a mis-cropped video or an explicit user-generated video that slips through the net. These happen when video governance is treated as a static legal or policy issue. Access controls, rights management, retention policies, regional restrictions, moderation and auditability cannot exist as disconnected documentation living somewhere in legal or compliance. For governance to be effective, it must become embedded directly into the workflow and operate at the infrastructure layer.
Moderation, in particular, is not feasible to scale manually. Ask someone to review a two-hour video and identify subtle violations, and they are almost guaranteed to miss something. Human review remains essential, but without automation it becomes inconsistent, expensive and difficult to apply at enterprise scale.
Why Video Is An Ideal AI Problem
By automating so many aspects of production and management, technology—smartphones, cloud platforms, generative tools and social platforms—brought down costs and made video more accessible. However, that same technology created a new problem of having to manage immense scale and complexity.
Enterprise video workflows like transcoding, metadata tagging, transcription, chapter generation, moderation assistance, rights classification and search indexing involve massive volumes of repetitive, detail-oriented, error-prone work. The same applies to watching, interpreting, classifying and moderating thousands of hours of enterprise video. These are exactly the kinds of workflows where human attention degrades over time and for which automation and AI are ideally suited. However, human expertise and oversight are still essential and a much more valuable use of their time.
The Opportunity For CTOs
Video was once just the preserve of marketing teams, but this is no longer the case. As new video-savvy generations start to dominate the workplace, video is increasingly used for internal comms, recruitment, training, product support and many other areas.
As onerous as the new demands are that “video-first” places on CTOs, it’s also an opportunity. Those that succeed will be the ones who apply the right mix of automation, IT and human skills. They can reduce infrastructure costs, improve governance, make enterprise knowledge searchable and avoid turning video into another unmanaged data-sprawl problem.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?







