Listing Thumbnail

    YouTube 8 Million - Data Lakehouse Ready

     Info
    Open data
    |
    Deployed on AWS
    This both the original .tfrecords and a Parquet representation of the [YouTube 8 Million dataset](https://research.google.com/youtube8m/). YouTube-8M is a large-scale labeled video dataset that consists of millions of YouTube video IDs, with high-quality machine-generated annotations from a diverse vocabulary of 3,800+ visual entities. It comes with precomputed audio-visual features from billions of frames and audio segments, designed to fit on a single hard disk. This dataset also includes the YouTube-8M Segments data from June 2019. This dataset is 'Lakehouse Ready'. Meaning, you can query this data in-place straight out of the Registry of Open Data S3 bucket. [Deploy this dataset's corresponding CloudFormation template](https://us-west-2.console.aws.amazon.com/cloudformation/home?region=us-west-2#/stacks/quickcreate?templateUrl=https://aws-roda-ml-datalake.s3.us-west-2.amazonaws.com/YT8MRodaTemplate.RodaTemplate.json&stackName=YT8M-RODA) to create the AWS Glue Catalog entries[...]

    Overview

    This both the original .tfrecords and a Parquet representation of the YouTube 8 Million dataset . YouTube-8M is a large-scale labeled video dataset that consists of millions of YouTube video IDs, with high-quality machine-generated annotations from a diverse vocabulary of 3,800+ visual entities. It comes with precomputed audio-visual features from billions of frames and audio segments, designed to fit on a single hard disk. This dataset also includes the YouTube-8M Segments data from June 2019. This dataset is 'Lakehouse Ready'. Meaning, you can query this data in-place straight out of the Registry of Open Data S3 bucket. Deploy this dataset's corresponding CloudFormation template  to create the AWS Glue Catalog entries into your account in about 30 seconds. That one step will enable you to interact with the data with AWS Athena, AWS SageMaker, AWS EMR, or join into your AWS Redshift clusters. More detail in (the documentation)[https://github.com/aws-samples/data-lake-as-code/blob/roda-ml/README.md .

    Features and programs

    Open Data Sponsorship Program

    This dataset is part of the Open Data Sponsorship Program, an AWS program that covers the cost of storage for publicly available high-value cloud-optimized datasets.

    Pricing

    This is a publicly available data set. No subscription is required.

    How can we make this page better?

    We'd like to hear your feedback and ideas on how to improve this page.
    We'd like to hear your feedback and ideas on how to improve this page.

    Legal

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Usage information

     Info

    Delivery details

    AWS Data Exchange (ADX)

    AWS Data Exchange is a service that helps AWS easily share and manage data entitlements from other organizations at scale.

    Open data resources

    Available with or without an AWS account.

    How to use
    To access these resources, reference the Amazon Resource Name (ARN) using the AWS Command Line Interface (CLI). Learn more 
    Description
    Original YT8M *.tfrecords. [File structure info can be found here](https://github.com/aws-samples/data-lake-as-code/blob/roda-ml/docs/roda_install.md).
    Resource type
    S3 bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::aws-roda-ml-datalake/yt8m/
    AWS region
    us-west-2
    AWS CLI access (No AWS account required)
    aws s3 ls --no-sign-request s3://aws-roda-ml-datalake/yt8m//
    Description
    Lakehouse ready YT8M as Glue Parquet files. [Install instructions here](https://github.com/aws-samples/data-lake-as-code/blob/roda-ml/docs/roda_install.md).
    Resource type
    S3 bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::aws-roda-ml-datalake/yt8m_ods/
    AWS region
    us-west-2
    AWS CLI access (No AWS account required)
    aws s3 ls --no-sign-request s3://aws-roda-ml-datalake/yt8m_ods//
    Description
    Replica of the two locations above in us-east-1.
    Resource type
    S3 bucket
    Amazon Resource Name (ARN)
    arn:aws:s3:::aws-roda-ml-datalake-us-east-1/
    AWS region
    us-east-1
    AWS CLI access (No AWS account required)
    aws s3 ls --no-sign-request s3://aws-roda-ml-datalake-us-east-1//

    Resources

    Similar products