Amazon Bedrock Managed Knowledge Base now supports multimodal embeddings for video, audio, and image content with TwelveLabs Marengo 3.0

Posted on: Sep 11, 2026

AWS announces the availability of TwelveLabs Marengo 3.0 as an embedding model in Amazon Bedrock Managed Knowledge Base, enabling customers to create multimodal embeddings for video, audio, and image content. Amazon Bedrock Managed Knowledge Base already supports media search by transcribing audio and video to text and generating text-based embeddings—Marengo 3.0 goes further by encoding visual scenes, speech, and video cues directly into multimodal embeddings, capturing meaning that transcription alone cannot. Simply upload your media assets from data sources such as Amazon S3, sync, and search using natural language—with no infrastructure to manage.
Marengo 3.0 produces compact 512-dimensional vectors, delivering state-of-the-art retrieval accuracy. Results include segment start and end times, enabling applications to jump directly to the relevant moment in a video. This unlocks use cases across sports analytics, media and entertainment, security, education, and retail—from finding specific plays across seasons of game footage to locating lecture segments by concept rather than keywords. The model offers configurable segmentation options to match your content structure.
To learn more, see TwelveLabs Marengo 3.0 embedding model integration in the Amazon Bedrock Knowledge Base User Guide. For more information, visit the Amazon Bedrock Knowledge Bases product page.