Listing Thumbnail

    Deep Lake

     Info
    Sold by: Activeloop 
    Deployed on AWS
    Free Trial
    AWS Free Tier
    Data lake for Deep Learning. Start now free production pilot.
    4

    Overview

    Activeloop Deep Lake streamlines deep learning dataset management at scale. Deep Lake maintains the benefits of a vanilla data lake, such as time traveling, SQL queries, ingesting data with ACID transactions, and visualizing terabyte scale datasets for analytical workloads with one key difference. It enables the storage of complex data types such as image, video, audio data, annotations, or tabular data, as columns with native integration to deep learning frameworks.

    As deep learning rapidly takes over traditional computational pipelines, storing datasets in a deep lake is becoming the de facto standard for shipping AI products fast. Deep Lake is optimized for cost efficient AI workflows. It allows rapid data streaming to deep learning frameworks over the network without sacrificing GPU utilization.

    With Deep Lake, teams can build a solid data foundation and iterate fast on their AI products - without bottlenecks from low quality data, underutilized compute resources, and significant labor overhead required to build and maintain large amounts of data.

    Deep Lake's features break down data silos, enable data driven decision making, improve operational efficiency, and reduce costs.

    Highlights

    • A scalable, efficient data storage system optimized for AI, built to power workloads on petabyte-scale complex multimodal data in a columnar format.
    • In-browser visualization and rapid querying engine with the full support of multimodal data types.
    • Native integration with deep learning frameworks and efficient streaming of data to models with full utilization of computing resources.

    Details

    Delivery method

    Deployed on AWS
    New

    Introducing multi-product solutions

    You can now purchase comprehensive solutions tailored to use cases and industries.

    Multi-product solutions

    Features and programs

    Financing for AWS Marketplace purchases

    AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
    Financing for AWS Marketplace purchases

    Pricing

    Free trial

    Try this product free according to the free trial terms set by the vendor.
    Pricing is based on the duration and terms of your contract with the vendor. This entitles you to a specified quantity of use for the contract duration. If you choose not to renew or replace your contract before it ends, access to these entitlements will expire.
    Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator  to estimate your infrastructure costs.

    1-month contract (1)

     Info
    Dimension
    Description
    Cost/month
    Growth
    1TB managed on Deep Lake
    $495.00

    Vendor refund policy

    Full refund during production pilot period

    How can we make this page better?

    Tell us how we can improve this page, or report an issue with this product.
    Tell us how we can improve this page, or report an issue with this product.

    Legal

    Vendor terms and conditions

    Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA) .

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Usage information

     Info

    Delivery details

    Software as a Service (SaaS)

    SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.

    Resources

    Support

    AWS infrastructure support

    AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.

    Product comparison

     Info
    Updated weekly
    By Activeloop

    Accolades

     Info
    Top
    50
    In Data Warehouses, Computer Vision
    Top
    10
    In Databases & Analytics Platforms, ML Solutions, Data Analytics
    Top
    25
    In Data Analysis

    Customer reviews

     Info
    Sentiment is AI generated from actual customer reviews on AWS and G2
    Reviews
    Functionality
    Ease of use
    Customer service
    Cost effectiveness
    0 reviews
    Insufficient data
    Insufficient data
    Insufficient data
    Insufficient data
    2 reviews
    Insufficient data
    Insufficient data
    Insufficient data
    Insufficient data
    Positive reviews
    Mixed reviews
    Negative reviews

    Overview

     Info
    AI generated from product descriptions
    Columnar Data Storage Format
    Scalable data storage system utilizing columnar format optimized for AI workloads on petabyte-scale complex multimodal data including images, videos, audio, annotations, and tabular data.
    Native Deep Learning Framework Integration
    Native integration with deep learning frameworks enabling efficient data streaming to models with full GPU utilization and rapid data ingestion without network bottlenecks.
    In-Browser Visualization and Query Engine
    In-browser visualization capabilities with rapid querying engine providing full support for multimodal data types at terabyte scale.
    ACID Transaction Support
    Data ingestion with ACID transactions and time traveling capabilities maintaining data lake benefits for complex dataset management.
    Cost-Optimized AI Workflows
    System architecture designed for cost-efficient AI workflows reducing computational overhead and labor requirements for large-scale dataset management and maintenance.
    Lakehouse Architecture
    Unified data foundation built on lakehouse architecture providing open, unified foundation for data and governance with support for open standards and formats
    Data Intelligence Engine
    Powered by Data Intelligence Engine that enables organization-wide access to data and insights across all users and roles
    Multi-Workload Unification
    Consolidates data engineering, analytics, business intelligence, data science and machine learning workloads on a single common platform
    Collaborative Development Environment
    Native collaboration capabilities enabling data teams to collaborate across entire data and AI workflow
    Open Source Foundation
    Built on open source data projects and open standards to maximize flexibility and interoperability with existing data ecosystems
    Unified Lakehouse Architecture
    Fully managed lakehouse platform integrating data storage, analytics, and AI workflows on a single unified platform
    Incremental Compute Engine
    Real-time incremental compute capabilities enabling faster model iteration and scalable experimentation
    Open Standards Support
    Built on open-source technologies and industry-leading open formats with native support for data lake standards
    Multi-Workload Integration
    Seamless integration of batch, streaming, and interactive workloads eliminating traditional data silos
    Unified Governance and Security
    Centralized governance and security controls across all data and AI use cases with vendor-agnostic architecture

    Contract

     Info
    Standard contract
    No

    Customer reviews

    Ratings and reviews

     Info
    4
    1 ratings
    5 star
    4 star
    3 star
    2 star
    1 star
    0%
    100%
    0%
    0%
    0%
    1 AWS reviews
    Gouthami

    Centralized unstructured data has reduced manual searches and saves significant analyst time

    Reviewed on Jul 15, 2026
    Review from a verified AWS customer

    What is our primary use case?

    My main use case for Deep Lake  is storing unstructured data, which can be in the form of images and PDFs, and we have some audios and videos as well. We have used it to store this data in a single place with built-in versioning. We have used Deep Lake  for searching and storing the data.

    We have used Deep Lake to store data with built-in versioning. It is used for storing and searching data plus vectors while building LLM applications and for managing datasets while training deep learning models. The typical use case is that we have used it in our support and legal team and research team who were manually digging through thousands of documents to answer questions. We wanted a bot that can read this entire data, whether it is in PDF, image, audio, or video format. We have used Deep Lake in that area.

    What is most valuable?

    The best features Deep Lake offers include the ability to store any type of data. Deep Lake stores any format of data: PDFs, images, audios, and videos in a single place. It has flexibility to store any kind of data, which is unstructured and structured. This is the main advantage.

    The flexibility to store any format of data makes the process easier. We wanted to recall everything when we needed to review instead of going through every document, opening every PDF and searching for the specific information. Deep Lake made our process easier by storing the data and giving us exactly what we are looking for at a particular point in time instead of requiring manual searching.

    Deep Lake has positively impacted our organization because it has saved our time and we can directly and indirectly see the impact on our FTEs. We have reduced one or two FTEs per month. The whole FTE reduction occurred because a few hours per week previously spent by analysis or support colleagues manually searching shared drives has been reduced. It has helped in terms of FTE reductions, faster speed, and reducing manual work that we used to do multiple times.

    I can say we have saved two FTEs per story, but a few hours per week. I cannot say this is the exact hour, but it can be a few hours a day because we roughly used to spend one or two hours previously.

    What needs improvement?

    Deep Lake can be improved by having filtering options that can give us more options to filter when we are comparing anything. Security-wise and cost predictability are areas that could be improved.

    Deep Lake could benefit from combined filtering that can handle heavy metadata filtering alongside search functionality. This is what I can recall at this point.

    For how long have I used the solution?

    I have been using Deep Lake for six months.

    What do I think about the stability of the solution?

    Deep Lake is stable and very stable compared to other solutions.

    What do I think about the scalability of the solution?

    Deep Lake handles data in larger volumes regularly. In terms of scalability, it is very scalable and reliable. Bottlenecks can be addressed effectively over time. It performs quite well.

    How are customer service and support?

    We have not reached out to customer support for major issues, but in the initial days, we raised some queries to customer support. They are available through some channels, but we need to manually follow up with them occasionally. Overall, their support is good.

    Which solution did I use previously and why did I switch?

    Before choosing Deep Lake, we looked at LanceDB . LanceDB  uses IVFPQ for ANN search, which is better for smaller datasets. Deep Lake uses linear search, so it can be used for lakh rows and HNSW based ANN beyond that. Both prioritize accuracy at a small scale and switch to approximate methods as the data grows. That is where we moved to Deep Lake.

    How was the initial setup?

    Deep Lake is not deployed as a public instance. We actually get it from the libraries. Installing Deep Lake does not require downloading from somewhere else. We can choose our own cloud for data storage. It is actually installed from the library.

    What about the implementation team?

    We use S3  for our own cloud deployment.

    What was our ROI?

    I have not seen a concrete return on investment in terms of money saved because it is an indirect impact. The direct impact was on time saving. That time saving has indirectly impacted money saving. People were actually spending a lot of time searching everything. This has been reduced to a few hours per week. Two employees were reduced if you look at the overall picture. I cannot give an exact number for the money saved, but it has made a very significant impact in terms of reducing manual work.

    What's my experience with pricing, setup cost, and licensing?

    I am not part of the licensing team, but I have an idea because I have been on calls regarding this matter. Deep Lake is open source and available from anywhere. We can download it and use it.

    Which other solutions did I evaluate?

    The switching part is the main reason we chose Deep Lake because it is open source and we can install and use it. Other tools exist which are common alternatives like LanceDB and PGVector. In terms of using and installing, we do not need to purchase anything or have any contract license-based arrangements. That is where we moved to Deep Lake.

    What other advice do I have?

    The advice I would give to others looking into using Deep Lake differs from business to business. Whether you have a small-scale or large-scale organization, if you do not have any budget to spend on these kinds of tools, Deep Lake will definitely establish itself on your platform. It depends on the use case and modeling training and multi-modal analysis. A good approach would be to acknowledge the variance rather than giving one blanket recommendation. Deep Lake is cost-sensitive, scalable, and reliable. I would recommend that all teams pilot Deep Lake with a real subset of data before full commitment since it differs from business to business. In terms of free access and everything, Deep Lake establishes itself well. I rate this product an eight out of ten.

    Which deployment model are you using for this solution?

    Private Cloud

    If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

    View all reviews