Activeloop Deep Lake streamlines deep learning dataset management at scale. Deep Lake maintains the benefits of a vanilla data lake, such as time traveling, SQL queries, ingesting data with ACID transactions, and visualizing terabyte scale datasets for analytical workloads with one key difference. It enables the storage of complex data types such as image, video, audio data, annotations, or tabular data, as columns with native integration to deep learning frameworks.
As deep learning rapidly takes over traditional computational pipelines, storing datasets in a deep lake is becoming the de facto standard for shipping AI products fast. Deep Lake is optimized for cost efficient AI workflows. It allows rapid data streaming to deep learning frameworks over the network without sacrificing GPU utilization.
With Deep Lake, teams can build a solid data foundation and iterate fast on their AI products - without bottlenecks from low quality data, underutilized compute resources, and significant labor overhead required to build and maintain large amounts of data.
Deep Lake's features break down data silos, enable data driven decision making, improve operational efficiency, and reduce costs.
Highlights
A scalable, efficient data storage system optimized for AI, built to power workloads on petabyte-scale complex multimodal data in a columnar format.
In-browser visualization and rapid querying engine with the full support of multimodal data types.
Native integration with deep learning frameworks and efficient streaming of data to models with full utilization of computing resources.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on the duration and terms of your contract with the vendor. This entitles you to a specified quantity of use for the contract duration. If you choose not to renew or replace your contract before it ends, access to these entitlements will expire.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
This listing offers one pricing dimension: Growth (TB). You pay for data managed on Deep Lake, measured in terabytes. Each unit covers 1TB of managed storage. Pricing scales with the amount of data you manage, so your cost grows as you add more terabytes. This is a contract-based commitment rather than pay-as-you-go billing. Deep Lake stores your data and supports AI workloads on that managed capacity.
Top-of-mind questions for buyers
What counts as one terabyte for the Growth (TB) dimension?
One unit covers 1TB of data managed on Deep Lake. This measures the data you store and manage, not the number of users or queries. As your stored data grows, you add more terabyte units, and your cost tracks that managed capacity.
What happens to my cost as I manage more data?
Cost scales with managed capacity. Each 1TB unit adds to your total. If your managed data grows beyond your committed terabytes, you add more units to cover the extra capacity. The charge is driven by how much data you manage, measured in terabytes.
What kind of data does this managed capacity support?
Deep Lake stores data used for AI workloads on the managed capacity you pay for. It combines vector and tensor data in one store and can stream data to fine-tuning tasks. Your terabyte units cover the data kept in this store.
activeloop.ai
Helpful?
Vendor refund policy
Full refund during production pilot period
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Scalable data storage system utilizing columnar format optimized for AI workloads on petabyte-scale complex multimodal data including images, videos, audio, annotations, and tabular data.
Native Deep Learning Framework Integration
Native integration with deep learning frameworks enabling efficient data streaming to models with full GPU utilization and rapid data ingestion without network bottlenecks.
In-Browser Visualization and Query Engine
In-browser visualization capabilities with rapid querying engine providing full support for multimodal data types at terabyte scale.
ACID Transaction Support
Data ingestion with ACID transactions and time traveling capabilities maintaining data lake benefits for complex dataset management.
Cost-Optimized AI Workflows
System architecture designed for cost-efficient AI workflows reducing computational overhead and labor requirements for large-scale dataset management and maintenance.
Lakehouse Architecture
Unified data foundation built on lakehouse architecture providing open, unified foundation for data and governance with support for open standards and formats
Data Intelligence Engine
Powered by Data Intelligence Engine that enables organization-wide access to data and insights across all users and roles
Multi-Workload Unification
Consolidates data engineering, analytics, business intelligence, data science and machine learning workloads on a single common platform
Collaborative Development Environment
Native collaboration capabilities enabling data teams to collaborate across entire data and AI workflow
Open Source Foundation
Built on open source data projects and open standards to maximize flexibility and interoperability with existing data ecosystems
Unified Lakehouse Architecture
Fully managed lakehouse platform integrating data storage, analytics, and AI workflows on a single unified platform
Incremental Compute Engine
Real-time incremental compute capabilities enabling faster model iteration and scalable experimentation
Open Standards Support
Built on open-source technologies and industry-leading open formats with native support for data lake standards
Multi-Workload Integration
Seamless integration of batch, streaming, and interactive workloads eliminating traditional data silos
Unified Governance and Security
Centralized governance and security controls across all data and AI use cases with vendor-agnostic architecture
Centralized unstructured data has reduced manual searches and saves significant analyst time
Reviewed on Jul 15, 2026
Review from a verified AWS customer
What is our primary use case?
My main use case for Deep Lake is storing unstructured data, which can be in the form of images and PDFs, and we have some audios and videos as well. We have used it to store this data in a single place with built-in versioning. We have used Deep Lake for searching and storing the data.
We have used Deep Lake to store data with built-in versioning. It is used for storing and searching data plus vectors while building LLM applications and for managing datasets while training deep learning models. The typical use case is that we have used it in our support and legal team and research team who were manually digging through thousands of documents to answer questions. We wanted a bot that can read this entire data, whether it is in PDF, image, audio, or video format. We have used Deep Lake in that area.
What is most valuable?
The best features Deep Lake offers include the ability to store any type of data. Deep Lake stores any format of data: PDFs, images, audios, and videos in a single place. It has flexibility to store any kind of data, which is unstructured and structured. This is the main advantage.
The flexibility to store any format of data makes the process easier. We wanted to recall everything when we needed to review instead of going through every document, opening every PDF and searching for the specific information. Deep Lake made our process easier by storing the data and giving us exactly what we are looking for at a particular point in time instead of requiring manual searching.
Deep Lake has positively impacted our organization because it has saved our time and we can directly and indirectly see the impact on our FTEs. We have reduced one or two FTEs per month. The whole FTE reduction occurred because a few hours per week previously spent by analysis or support colleagues manually searching shared drives has been reduced. It has helped in terms of FTE reductions, faster speed, and reducing manual work that we used to do multiple times.
I can say we have saved two FTEs per story, but a few hours per week. I cannot say this is the exact hour, but it can be a few hours a day because we roughly used to spend one or two hours previously.
What needs improvement?
Deep Lake can be improved by having filtering options that can give us more options to filter when we are comparing anything. Security-wise and cost predictability are areas that could be improved.
Deep Lake could benefit from combined filtering that can handle heavy metadata filtering alongside search functionality. This is what I can recall at this point.
For how long have I used the solution?
I have been using Deep Lake for six months.
What do I think about the stability of the solution?
Deep Lake is stable and very stable compared to other solutions.
What do I think about the scalability of the solution?
Deep Lake handles data in larger volumes regularly. In terms of scalability, it is very scalable and reliable. Bottlenecks can be addressed effectively over time. It performs quite well.
How are customer service and support?
We have not reached out to customer support for major issues, but in the initial days, we raised some queries to customer support. They are available through some channels, but we need to manually follow up with them occasionally. Overall, their support is good.
Which solution did I use previously and why did I switch?
Before choosing Deep Lake, we looked at LanceDB. LanceDB uses IVFPQ for ANN search, which is better for smaller datasets. Deep Lake uses linear search, so it can be used for lakh rows and HNSW based ANN beyond that. Both prioritize accuracy at a small scale and switch to approximate methods as the data grows. That is where we moved to Deep Lake.
How was the initial setup?
Deep Lake is not deployed as a public instance. We actually get it from the libraries. Installing Deep Lake does not require downloading from somewhere else. We can choose our own cloud for data storage. It is actually installed from the library.
What about the implementation team?
We use S3 for our own cloud deployment.
What was our ROI?
I have not seen a concrete return on investment in terms of money saved because it is an indirect impact. The direct impact was on time saving. That time saving has indirectly impacted money saving. People were actually spending a lot of time searching everything. This has been reduced to a few hours per week. Two employees were reduced if you look at the overall picture. I cannot give an exact number for the money saved, but it has made a very significant impact in terms of reducing manual work.
What's my experience with pricing, setup cost, and licensing?
I am not part of the licensing team, but I have an idea because I have been on calls regarding this matter. Deep Lake is open source and available from anywhere. We can download it and use it.
Which other solutions did I evaluate?
The switching part is the main reason we chose Deep Lake because it is open source and we can install and use it. Other tools exist which are common alternatives like LanceDB and PGVector. In terms of using and installing, we do not need to purchase anything or have any contract license-based arrangements. That is where we moved to Deep Lake.
What other advice do I have?
The advice I would give to others looking into using Deep Lake differs from business to business. Whether you have a small-scale or large-scale organization, if you do not have any budget to spend on these kinds of tools, Deep Lake will definitely establish itself on your platform. It depends on the use case and modeling training and multi-modal analysis. A good approach would be to acknowledge the variance rather than giving one blanket recommendation. Deep Lake is cost-sensitive, scalable, and reliable. I would recommend that all teams pilot Deep Lake with a real subset of data before full commitment since it differs from business to business. In terms of free access and everything, Deep Lake establishes itself well. I rate this product an eight out of ten.
Which deployment model are you using for this solution?
Private Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?