Weights & Biases AI Development Platform for AWS logo

    Weights & Biases AI Development Platform for AWS

    Weights & Biases provides AI developers with the tools needed to build models faster, fine-tune LLMs, and develop GenAI applications with confidence for enterprises of all sizes in any vertical.

    Ratings and reviews

    4.5
    63 ratings
    2 star
    1 star
    71%
    27%
    2%
    0%
    0%
    3 AWS reviews
    |
    60 external reviews
    External reviews are from G2  and PeerSpot .

    Filters

    Review type

    AWS Marketplace reviews
    External reviews
    Reviews (63)
    Joseph Kaiser

    Tracking protein runs has improved checkpointing and now saves weeks of rerun work

    Reviewed on Aug 16, 2026
    Review from a verified AWS customer

    What is our primary use case?

    My main use case for Weights & Biases is for tracking runs for protein investigation to drug target discovery targets.

    What is most valuable?

    As the administrator with Weights & Biases, I think it's an incredible piece of software. We have used it to track an incredible amount of data, and we've been able to refer back to runs that have been critical in investigating new drugs.

    It's extremely easy to administer. It has an incredible API for provisioning of users and groups and the various teams. The support has been stellar. They have been extremely responsive and they have been a real joy to work with.

    I would say the best feature Weights & Biases offers is ease of use. They have been able to work the use of Weights & Biases for tracking checkpoints easily into their workflows which are varied. They use it from the command line, they use it from dag flows via Argo. It's just been a piece of software that people have depended upon and sometimes take for granted, but it's one of those things that's always there, always available, and easy to use.

    The ease of use of Weights & Biases impacts my team's day-to-day work significantly because before they were not doing very much checkpointing at all. We had jobs that would run for days and would crash and then would be unable to return back to an earlier point in the analysis, and with the use of Weights & Biases, that's easy. Any kind of work that has to be done that gets interrupted can go back easily to a previous iteration and begin from that point rather than having to redo the entire job. This saves days and weeks of redoing work.

    Weights & Biases has positively impacted my organization in that the support is excellent both for us as administrators and for our teams. They frequently run webinars for people to get better use out of it and expand the features, and that's been helpful.

    What needs improvement?

    I don't really know how Weights & Biases can be improved; that would have to come from one of the researchers.

    From an administrator's perspective, I think one of the difficulties that we are experiencing is we have a lot of historical data, and I think we don't understand how best to easily take care of that, but I don't know if that's a Weights & Biases problem.

    There are no improvements needed for Weights & Biases that I haven't mentioned.

    For how long have I used the solution?

    I have been using Weights & Biases for about the last four years.

    What do I think about the stability of the solution?

    Weights & Biases is very stable.

    What do I think about the scalability of the solution?

    Weights & Biases handles increased workloads or more users easily.

    How are customer service and support?

    The customer support is stellar. We have an open channel with them in Slack that we can ask questions, both researchers and admins. Someone is always available to us there. The migration from self-hosted to cloud was seamless. We had parts that we had to do, and there were parts that they had to do. We had a clear roadmap, we had clear delineation of tasks on who had to do what and on which side, whether it was our side or their side. That was one of the smoothest migrations I've ever had in my software and administrative career.

    Which solution did I use previously and why did I switch?

    We started out doing everything by hand and writing things to disk; it was homemade and not a good option. Weights & Biases came along and filled a real need.

    How was the initial setup?

    Weights & Biases started out with us on-premises, but we are now in public cloud with Amazon for our deployment.

    What about the implementation team?

    I didn't purchase Weights & Biases through the AWS Marketplace; we have a private contract.

    What's my experience with pricing, setup cost, and licensing?

    My experience with pricing, setup cost, and licensing was that we received very favorable terms. The licensing was easy. The setup cost and the migration were minimal in comparison to what we got. They helped us migrate from our on-prem to cloud with just a small fee. It was amazing. They did a great deal of work.

    Which other solutions did I evaluate?

    I don't think we evaluated other options before choosing Weights & Biases, or if we did, I wasn't here at the company when that was initially made. I think people from other companies had used it and felt like it was a good fit for our organization.

    What other advice do I have?

    I can't give a quick specific example of how I use Weights & Biases for protein investigations or target discovery because I'm just the administrator for it.

    Weights & Biases is an awesome piece of software.

    I can't share any specific outcomes or metrics that show how Weights & Biases has helped my organization.

    My advice to others looking into using Weights & Biases is to absolutely use it. Figure out in what ways and what options are there in order to be able to take full advantage of it. I think sometimes we don't use all the resources it provides, but certainly, what it does provide as its core business is sufficient for us.

    Weights & Biases' governance and security capabilities are very good; we're Okta enabled on it. We don't have any issues, and the infrastructure itself only has a very few set of whitelisted IPs for access, and we're able to do most things through a service account. So it's great.

    Weights & Biases' accuracy and reliability of output are very accurate and reliable. We have had no complaints from researchers. It's just part of their everyday day-to-day and they depend on it, so its accuracy and reliability have been sufficient for us enough to move forward and not worry about having to be concerned about reliability or accuracy.

    I give this review a rating of eight out of ten.

    Which deployment model are you using for this solution?

    Public Cloud

    If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

    Amazon Web Services (AWS)
    Pawel Cislo

    Automated lineage has transformed model governance and now simplifies reliable audits

    Reviewed on Aug 12, 2026
    Review provided by PeerSpot

    What is our primary use case?

    My main use case for Weights & Biases is data lineage tracking and model registry management, as I use Weights & Biases to keep a complete version history of data sets and models, making it easy to trace every trained model back to the exact data, code, and artifacts used to create it.

    Beyond tracking and registry with Weights & Biases, the automated lineage graphs are a huge time-saver for auditability and team collaboration, meaning that if a model ever behaves unexpectedly down the line, anyone on the team can inspect the registry entry and immediately see the exact parameters, data artifacts, and code commit that produced it without having to dig through logs.

    What is most valuable?

    The standout feature of Weights & Biases is its Artifacts combined with automated data lineage graphs, which automatically track the exact inputs and outputs for every run, generating a complete directed acyclic graph that maps datasets to models seamlessly. Another top feature is the Model Registry, which gives us an organization-wide centralized hub to manage model lifecycles, assign mutable aliases such as staging or production, and trigger downstream CI/CD pipelines automatically whenever a new model version is promoted.

    On the visualization side, Weights & Biases Reports are phenomenal, as you can instantly turn dynamic experiment dashboards into interactive, shareable documents with live plots, text, notes, and code snippets. This completely eliminates the need to take static screenshots for team updates or slide decks, ensuring that anyone on the team can inspect live charts and drill down into the metrics directly.

    Weights & Biases has significantly boosted our efficiency and reliability, with the biggest impact being complete reproducibility and traceability, as we no longer waste hours trying to reconstruct how a specific model was trained or which dataset version was used. It has also streamlined our model deployment workflows through the Model Registry, making transitions from training to production much smoother and reducing human error, creating a single source of truth that saves us substantial engineering time and keeps our MLOps processes tight and auditable.

    Quantitatively, Weights & Biases has reduced our model audit and debugging time by roughly 50%, as in the past, tracking down the exact dataset commit and hyperparameter set for an older model could easily take half a day, but now it takes under two minutes in the Weights & Biases registry. Qualitatively, it has almost completely eliminated deployment errors caused by model-data mismatch or missing metadata, and having a standardized, automated lineage check before promoting a model to production gives us total confidence and saves us from costly post-deployment headaches.

    What needs improvement?

    The main area for improvement in Weights & Biases is cost predictability and pricing scaling, since as logging frequency and artifact storage scale up across larger teams, expenses can climb surprisingly fast. Therefore, more granular cost control toggles or sampling controls directly in the SDK would be a huge help. Additionally, self-hosted or air-gapped enterprise deployments can still be quite complex to configure and maintain compared to their managed SaaS version, so streamlining the Kubernetes Helm installation for private clouds and making self-hosted setups lighter on resources would make a big difference for security-conscious MLOps environments.

    On the developer experience side, the Python SDK documentation could benefit from clearer, production-grade examples, as while basic getting-started guides are great, finding detailed code patterns for advanced edge cases such as complex multi-model artifact tracking or custom orchestration setups often requires digging through community forums. Regarding integration, expanding native connectors for certain Kubernetes-native tools and GitOps pipelines would make automated model production feel more seamless out of the box, without needing as many custom webhook scripts.

    For how long have I used the solution?

    I have been using Weights & Biases for approximately two years in my current project.

    What do I think about the stability of the solution?

    Weights & Biases is highly stable, as it serves as an established, enterprise-grade industry standard for MLOps that reliably handles large-scale production workloads, high-frequency logging, and complex data tracking across large engineering teams.

    What do I think about the scalability of the solution?

    Weights & Biases' scalability is exceptional, as it seamlessly scales from individual local prototypes to enterprise workloads with millions of logged metrics, large artifact storage, and distributed multi-node GPU training clusters. Its architecture is built to ingest high-frequency logging from parallel training runs without choking, and features such as Artifacts and Model Registry scale effortlessly as data volumes and team sizes grow.

    How are customer service and support?

    The customer support experience with Weights & Biases has been very reliable, as for routine development and edge cases, their traditional documentation, API references, and active community forums such as Slack and GitHub are thorough and quickly answer most technical questions. When enterprise-level support is needed, such as troubleshooting pipeline integrations or deployment issues, their dedicated support engineers are responsive, technically competent, and work directly with MLOps teams to resolve issues efficiently.

    Which solution did I use previously and why did I switch?

    Previously we relied on MLflow along with custom in-house scripts for tracking, but we switched to Weights & Biases because MLflow required significant effort to maintain, customize, and scale on our own infrastructure. Weights & Biases provided a much smoother user experience out of the box, especially around automated data lineage visualization, a more polished Model Registry UI, and seamless interactive reporting, drastically reducing our setup overhead and improving team collaboration.

    How was the initial setup?

    During our evaluation phase for Weights & Biases, we specifically looked at MLflow, Neptune.AI, and TensorBoard, ultimately selecting Weights & Biases because of its superior automated data lineage tracking, a more refined Model Registry UI, and effortless interactive reporting, which gave us the best combination of feature completeness and low developer overhead.

    What about the implementation team?

    We use an on-premise deployment of Weights & Biases, which is managed for us by an external third-party vendor, allowing our team to leverage Weights & Biases locally while ensuring strict data privacy and security compliance within our environment.

    I'm not directly involved in the purchasing, setup, or licensing of products for Weights & Biases, as this side of things, including vendor negotiations and infrastructure management, is handled entirely by the external company managing our on-prem deployment. My focus is purely on the engineering side and hands-on usage of the platform.

    What was our ROI?

    We've seen a solid return on investment with Weights & Biases, mainly in engineering time saved and risk reduction, as quantitatively, it saves our team about 30 to 40% of time on experiment tracking and auditing. Finding past datasets or model versions now takes minutes instead of hours. Qualitatively, having an automated lineage in the registry prevents costly deployment errors from mismatched models, while also making team collaboration and handovers effortless.

    Which other solutions did I evaluate?

    From an MLOps perspective, Weights & Biases plays a central role in both AI governance and security, as its features such as Artifacts and the Model Registry provide an immutable audit trail. They automatically track end-to-end data lineage, mapping exact dataset versions, code commits, and hyperparameters directly to deployed models, making model compliance, internal audits, and reproducing past results straightforward. Additionally, Weights & Biases offers role-based access control to restrict access to sensitive datasets or production models across teams, and for enterprise setups, it supports single sign-on, encryption at rest and in transit, SOC 2, ISO 27001 compliance, and flexible deployment options such as private cloud or air-gapped VPC instances to keep proprietary data and model weights secure.

    What other advice do I have?

    What makes Weights & Biases stand out is its seamless developer experience with Artifacts and the Model Registry, which automatically builds end-to-end data lineage graphs and provides an intuitive, interactive dashboard without adding heavy code overhead. The aspects that keep it from being a perfect 10 are the pricing scaling at high data volumes and the complexity of managing self-hosted or air-gapped enterprise setups on Kubernetes.

    In terms of accuracy and reliability, it's important to clarify that Weights & Biases isn't generating model outputs itself; it acts as the system of record and evaluation infrastructure. From an evaluation perspective, its reliability is top-tier. Through toolsets such as Weights & Biases Weave, it provides a structured framework for evaluation and observability, letting you implement custom metrics and LLM-as-a-judge scoring while running standardized benchmarks to measure hallucination rates, factual accuracy, and context relevance deterministically. What makes it so reliable is traceability, as instead of giving you vague scores, every single evaluation metric or trace is tied directly to the exact model version, dataset commit, and prompt template used, eliminating guesswork and ensuring that when you measure model accuracy or failure modes in production, the data you're looking at is 100% reproducible and verifiable.

    My biggest advice for others looking into using Weights & Biases is to adopt Artifacts and standard logging conventions right from day one, as you should not treat it as a basic dashboard for plotting loss curves. Truly leverage the Model Registry and dataset lineage capabilities early on, and establish clear naming conventions for your runs, artifacts, and projects across your team, as setting up these MLOps best practices from the start saves a massive amount of cleanup time later, ensuring full reproducibility and smooth collaboration as your projects scale.

    I provided this review with an overall rating of 9 out of 10.

    Oil & Energy

    Clear ML Experiment Tracking with Easy Integration and Reliable Versioning

    Reviewed on Aug 11, 2026
    Review provided by G2
    What do you like best about the product?
    They are experiment tracking dashboard makes it completely easier for us to compare training runs, metrics and other hyperparameters in one place. Integration requires only a few lines of code and works well with popular ml frameworks. The visualisations are very clear and particularly helpful when debugging model performance. Artefact and other registry which also provide reliable versioning for datasets and models. Overall, I would say it creates a strong shade workspace for ML teams.
    What do you dislike about the product?
    This platform sometimes uh being overwhelming initially because it includes many features, dashboards and other configuration options. Organising projects also becoming bit difficult if the team does not establish consistent naming conventions early. Their interface can also Occasionally feels bit slower when loading projects contain a large number of runs. And some of their advanced collaboration along with the governance and deployment capabilities, are also limited to the paid plans. Pricing may become expensive for growing teams with extensive usage.
    What problems is the product solving and how is that benefiting you?
    This platform replaces manual spreadsheets and scattered logs with the centralized record of every machine learning experiment. It also helps us to reproduce previous results by capturing metrics, parameters, system usage and other data set models. Comparing runs allows us to identify the best-performing configuration much faster. And our teams can review progress and share findings right on the spot without repeatedly exchanging files. This overall reduced the experimentation time and improved the collaboration throughout the model development life cycle.
    Atharva S.

    Streamlined ML Experiment Tracking with Rich Visualizations and Team Collaboration

    Reviewed on Aug 05, 2026
    Review provided by G2
    What do you like best about the product?
    What I like best about Weights & Biases is how it simplifies machine learning experiment tracking, model management, and collaboration through an intuitive and well-designed platform. It makes it easy to monitor training runs, compare experiments, visualize metrics, and organize models in one place, which significantly improves the development workflow. I also appreciate its rich visualizations, seamless integration with popular ML frameworks like PyTorch and TensorFlow, and strong collaboration features for teams. Overall, Weights & Biases accelerates model development, improves experiment reproducibility, and makes managing machine learning projects much more efficient.
    What do you dislike about the product?
    One area where Weights & Biases could improve is offering more advanced customization for dashboards, reporting, and experiment organization to better support very large machine learning projects. While the platform is feature-rich, new users may experience a learning curve when exploring advanced capabilities such as artifact management and workflow automation. I'd also like to see broader integrations with additional MLOps and enterprise platforms, along with more flexible access controls and reporting options. Overall, the experience has been very positive, but greater customization, expanded integrations, and enhanced enterprise features would make Weights & Biases even more valuable.
    What problems is the product solving and how is that benefiting you?
    Weights & Biases solves the challenge of managing machine learning experiments by providing a centralized platform for experiment tracking, model evaluation, dataset versioning, and collaboration. Instead of manually recording training metrics and comparing results across different runs, it automatically logs parameters, visualizes performance, and organizes experiments in a structured way. This improves reproducibility, accelerates model iteration, simplifies collaboration among data science teams, and reduces the time spent on experiment management. As a result, it has streamlined the machine learning development workflow, increased productivity, and made it much easier to build, compare, and deploy high-performing models.
    Muhammed A.

    Essential ML Experiment Tracking with Real-Time Metrics and Team Collaboration

    Reviewed on Jul 31, 2026
    Review provided by G2
    What do you like best about the product?
    Weights & Biases has become an essential platform for managing machine learning experiments, model training, and performance tracking. The interface makes it easy to compare runs, visualize metrics in real time, and collaborate across teams, while integrations with popular ML frameworks simplify adoption. Experiment tracking, artifact versioning, and reproducibility features significantly reduce manual work, helping teams iterate faster, improve model quality, and maintain organized AI development workflows.
    What do you dislike about the product?
    Weights & Biases offers a comprehensive feature set, but new users may face a learning curve when configuring advanced experiment tracking, reports, and team workflows. Large projects with thousands of experiment runs can sometimes make dashboards feel cluttered, and premium features may be costly for smaller teams. I would also like to see more customization options for visualizations and reporting, along with additional native integrations for enterprise MLOps environments.
    What problems is the product solving and how is that benefiting you?
    Before using Weights & Biases, tracking machine learning experiments, comparing model performance, and managing training artifacts across multiple projects was time-consuming and difficult to reproduce. The platform centralized experiment tracking, visualization, model versioning, and collaboration in a single workspace, making it much easier to monitor progress and identify the best-performing models. This has reduced manual effort, improved reproducibility, accelerated model development cycles, and enabled the team to make faster, data-driven decisions throughout the ML lifecycle.
    Automotive

    Automatic Metrics Tracking, but Overall Experience Needs Improvement

    Reviewed on Jul 30, 2026
    Review provided by G2
    What do you like best about the product?
    Automatically records metrics, code versions, making results better
    What do you dislike about the product?
    Projects can get cluttered over time, and that can feel overwhelming.
    What problems is the product solving and how is that benefiting you?
    Keeps records and training for every run.
    Helps identify changes
    Biotechnology

    ML Experiment Tracking, Forward Deployment, and Open-Weight Models Made Easy

    Reviewed on Jul 28, 2026
    Review provided by G2
    What do you like best about the product?
    Makes tracking training experiments and sharing training data with my team easy, with dashboards similar to Tensorboard and low performance overhead. Easy to get started with. Backs up data to the cloud and works from a remote cluster seamlessly. Plus offers support for purchasing cloud compute for LLM fine-tuning and FAAS.
    What do you dislike about the product?
    It doesn't display large quantities of data well, and it's difficult to use some of the more complex visualizations. As a place for publishing/using models, HuggingFace has a larger library and simpler API. Cloud compute pricing is competitive but higher than competitors.
    What problems is the product solving and how is that benefiting you?
    It helps us log ML training/evaluation data (though the Experiments and Reports features) remotely as I work on an HPC cluster. I can access the data anytime through the mobile app or website, which is convenient because we don't need a secure connection to the cluster. We can also save model weights/architectures and publish them online alongside our academic papers.
    Dhruv P.

    Solid MLOps platform for experiment tracking with great collaboration features

    Reviewed on Jul 28, 2026
    Review provided by G2
    What do you like best about the product?
    Excellent experiment tracking and visualization dashboard that makes it easy to compare model runs and parameters. Strong integrations with major ML frameworks and seamless team collaboration features. The API is intuitive and well-documented, making it straightforward to log metrics and artifacts.
    What do you dislike about the product?
    Pricing scales steeply with team size, which can be a barrier for smaller organizations. The learning curve for advanced features like custom dashboards and reports is moderate, and documentation could be more comprehensive for edge cases. Occasional UI/UX inconsistencies across different features.
    What problems is the product solving and how is that benefiting you?
    Helps organize and track ML experiments systematically, reducing time spent manually managing experiment logs and parameters. Enables better collaboration across teams by centralizing model run history and results. Improves reproducibility and debugging of models by maintaining complete audit trails. Accelerates model iteration cycles and provides visibility into which hyperparameters yield the best performance.
    Jeni J.

    A Must-Have Tool for Keeping ML Experiments Organized

    Reviewed on Jul 28, 2026
    Review provided by G2
    What do you like best about the product?
    I primarily use Weights & Biases to track and compare machine learning experiments, monitor training metrics in real time, and manage model versions. I really like how it solves the challenge of keeping experiments organized and reproducible, with everything logged automatically. What I like most about Weights & Biases is how effortless it makes experiment tracking and visualization. The interactive dashboards, real-time training metrics, hyperparameter comparison tools, and artifact management are great for understanding model performance, reproducing results, and collaborating with teammates without adding much overhead to the workflow. The initial setup was developer friendly too. the AI finetuning, monitoring was very good.
    What do you dislike about the product?
    One area that could be improved is the onboarding experience for new users, especially when exploring advanced features like Sweeps, Artifacts, and Reports. While the platform is very powerful, it can feel overwhelming at first, so more guided tutorials, in-app tips, and ready-to-use workflow templates would help users become productive much faster. I'd also like to see more flexible dashboard customization and filtering options for large projects with hundreds of experiment runs. Better cost and resource usage insights, along with faster loading times for very large experiment histories, would make the platform even more efficient for teams managing complex machine learning and LLM workflows.
    What problems is the product solving and how is that benefiting you?
    Weights & Biases solves organizing and reproducing ML experiments, automates tracking metrics, hyperparameters, and versions, and aids collaboration in AI projects. It helps me monitor training metrics, manage model versions, and track experiments effortlessly.
    Muhammad O.

    A Reliable Platform for Tracking Machine Learning Experiments

    Reviewed on Jul 25, 2026
    Review provided by G2
    What do you like best about the product?
    What I like most is how easy it is to get started and keep all my experiments organized in one place. The dashboard feels clean and intuitive, so it’s straightforward to track runs, compare results, and share progress with teammates. Overall, it helps me manage model development in a more structured way without ever feeling overly complicated.
    What do you dislike about the product?
    The platform offers a lot of features, so it can feel a bit overwhelming when you’re first getting started. It took me some time to figure out where everything was and how it all fit together, but after I spent a little time exploring, it became much easier to navigate.
    What problems is the product solving and how is that benefiting you?
    Weights & Biases helps me keep machine learning experiments organized by tracking runs, comparing results, and making it easier to see which changes actually improve a model. It saves time, supports collaboration, and makes it much simpler to reproduce past experiments rather than having to start from scratch.