TwelveLabs, a San Francisco-based startup specializing in AI models for understanding and searching video content, secured a significant $100 million Series B funding round. The investment was co-led by New Enterprise Associates (NEA) and Naver, with Amazon.com also participating as a backer. Other investors included Radical Ventures, Index Ventures, and Korea Investment Partners. This funding builds upon previous investments from prominent tech companies like Nvidia, Intel, and Samsung. The company's goal is to expand its capabilities beyond video-understanding models into a full-stack agentic system that integrates perception, knowledge, and reasoning for video analysis.

A key aspect of this development is a multiyear agreement with Amazon Web Services (AWS) to host TwelveLabs' workloads on Amazon's custom Trainium chips. This move is particularly noteworthy as it signifies Amazon's strategy to utilize its own silicon for demanding AI workloads, potentially reducing reliance on Nvidia's GPUs. The collaboration means that new TwelveLabs models are slated to launch on AWS, providing developers with tools for building video search and analysis applications. This partnership gives Amazon a strong case for its Trainium chips in the video AI sector, akin to how Trainium2 chips power Anthropic's Project Rainier cluster for chatbots and code assistants.

TwelveLabs addresses the challenge of indexing and understanding video, which is significantly more complex than text or images due to the continuous nature of motion, speech, objects, faces, and contextual information. The company's value proposition for enterprises is to enable searching video libraries in natural language, similar to searching documents, to quickly locate specific clips where events occur or products appear. This capability is highly beneficial for sectors such as media archives, retail security teams, and sports organizations that deal with vast amounts of raw, uncut video footage daily. While academic benchmarks often focus on edited video, TwelveLabs aims to bridge the gap by providing solutions for the real-world challenge of processing unedited, high-volume video data at scale.

From a business strategy perspective, TwelveLabs is focused on deepening its specialization in video processing. The company's foundation models, Marengo for search and Pegasus for video-to-text reasoning, are offered as APIs for enterprises in media, sports, and ad-tech. They provide a self-serve developer plan priced at $0.042 per minute of video indexing and $4 per 1,000 search queries, with custom Enterprise contracts for larger clients. This focus is crucial for a startup, as directly competing with big tech on broad AI solutions is impractical. Instead, TwelveLabs aims to become an indispensable component in the video processing pipeline, addressing the significant gap between academic benchmarks and the practical needs of industry professionals who generate thousands of hours of footage daily.