Sample repository from a large-scale User Generated Content (UGC) video corpus featuring real-world videos across diverse creators, environments, activities, and content categories for video understanding and AI applications.
This dataset is a large-scale collection of User Generated Content (UGC) videos designed to support video understanding, content analysis, computer vision, multimodal AI, and machine learning applications.
The corpus contains authentic user-generated videos captured across diverse environments, activities, lifestyles, and real-world scenarios. The dataset reflects the variety and complexity commonly found in modern digital content, providing valuable visual data for developing robust AI systems capable of understanding real-world video content.
The collection includes videos recorded by individuals across multiple settings, capturing natural interactions, activities, events, locations, and everyday experiences. This diversity enables AI systems to learn from realistic visual patterns and user-generated media formats.
Key Use Cases
Video Understanding
Content Analysis
Activity Recognition
Scene Understanding
Human Behavior Analysis
Multimodal AI
Visual Search
Video Classification
Content Recommendation Systems
Social Media Analytics
Consumer Content Analysis
Video Intelligence Applications
Dataset Features
Large-scale UGC video collection
Real-world user-generated videos
Diverse creators and recording environments
Multiple activity and lifestyle categories
Natural visual content and interactions
Broad environmental coverage
Suitable for training and evaluation workflows
Rich contextual video information
Content Coverage
The dataset includes user-generated videos spanning a wide range of categories and scenarios, including:
Lifestyle content
Daily activities
Entertainment videos
Social interactions
Indoor and outdoor environments
Personal experiences
Community activities
Consumer-generated media
Event-based recordings
Real-world visual content
The diversity of environments, creators, and activities provides extensive visual variability for training robust AI systems.
AI & Analytics Applications
The corpus supports development of video intelligence systems capable of understanding visual content, contextual information, activity patterns, and user-generated media. Organizations can leverage the dataset for video analytics, content moderation, recommendation systems, visual search, multimodal learning, and next-generation video understanding applications.
Data Collection
The dataset consists of user-generated video content curated to represent diverse visual environments, activities, and content styles. The collection is organized to support research, evaluation, and large-scale AI development workflows.
Licensing & Access
This listing contains sample data intended for research, evaluation, and educational purposes. Enterprise licensing and access to the complete dataset are available upon request.
Large-scale User Generated Content (UGC) video corpus featuring real-world videos captured across diverse environments, creators, and content categories.
Includes authentic consumer-generated video content covering daily activities, lifestyle, entertainment, social interactions, and real-world scenarios.
Supports video understanding, content analysis, activity recognition, multimodal learning, visual search, and AI model development workflows.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
This listing offers a single pricing dimension: Product Access (Units). It grants subscribers access to the User Generated Content video dataset at no cost. There are no tiers, instance sizes, or usage add-ons to compare. The dataset itself covers user-generated video, classroom recordings, and storytelling footage for multimodal AI training. Volume and source-type mix are scoped with the vendor per engagement. Because pricing is free, the single unit dimension simply enables access rather than metering usage or scaling with consumption.
Top-of-mind questions for buyers
What does the Product Access unit actually grant me, and how is the dataset scoped?
The unit grants subscriber access to the video dataset. It does not meter usage or scale by volume. Instead, the actual data you receive is scoped per engagement with the vendor. You define your model, languages, and volume, and the vendor scopes a sample and licensing path with you.
Can I license only classroom video or only user-generated video, not the full corpus?
Yes. The volume and source-type mix are scoped per engagement. The corpus mixes STEM classroom recordings, vertical short-form user-generated video, and long-form storytelling footage. You can license a specific source type or combination rather than the entire dataset.
What delivery format and metadata come with the video dataset?
Video ships as MP4 using the AV1 codec, audited per clip. Each file pairs with JSON metadata covering source type, orientation, and refining-flag results. Every clip passes an eight-step refining pipeline, including synthetic-media detection and watermark checks, before licensing.
infobay.ai
Helpful?
Vendor refund policy
No Refunds
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
Ultimate AI Media Generation Platform With 1000+ top-tier models, WaveSpeedAI is the most powerful platform for AI image and video generation - built to help you create faster and scale without limits.
Personalized, relevant, and adaptive security awareness training for human risk management.
KnowBe4's Security Awareness Training (SAT) transforms how your organization manages human risk by combining the industry's most comprehensive security awareness training with groundbreaking AI-native defense agents - including the Orchestration Agent that fully automates program administration - all powered by 15+ years of threat intelligence and user data.
Cloudinary Video is an API-first and AI powered platform that enables over 5000 customers, including some of the world's most demanding brands such as Adidas, Bleacher Report, Etsy, Paul Smith and more to scale high-performing video experiences across their websites and apps in minutes. Visit https://cloudinary.com/products/video
Untapped Value Remains Hidden Within Video Assets:
Organizations often handle video content through separate systems that look at visuals, audio, and text independently. This disconnected approach misses the bigger picture and fails to unlock the true value of video assets, where just one minute of video can contain as much information as 1.8 million words. Companies need a holistic approach to video understanding that transforms content from an underutilized asset into a source of strategic insight.
Transforms Your Video Data into Actionable Intelligence and Personalized Experiences:
Our solution uses AWS AI to process visual, audio, and text elements simultaneously, delivering context-aware intelligence that mimics human perception. This enables automated highlight generation, personalized content recommendations, and real-time audience engagement across streaming platforms—transforming fragmented analysis into an AI ecosystem that drives viewer retention and monetization.