
Monte Carlo Data + AI Observability Platform
Monitoring data products has improved and alerts now keep our models reliable and visible
What is our primary use case?
My main use case for Monte Carlo is data monitoring, monitoring jobs that have failed, and I have also used testing fairly extensively. I usually use the pre-boxed testing that Monte Carlo offers, but on a couple of other use cases I have developed custom testing as well.
I cannot remember the exact details of how I use custom tests in Monte Carlo because the last time I did it was probably several years ago. I think that it was generally testing that was not more generic, such as uniqueness testing or volume monitoring, but I honestly do not know if I can remember or give the exact details of what I was doing.
As far as doing anything very unique or different that I have tried in Monte Carlo, I would say I am more of a standard user of it. We have it connected to one of our Slack channels that monitors our jobs and any anomalies that we have within our data products. I primarily use it to keep an eye on some of my models that I am responsible for.
What is most valuable?
I have found a lot of value in volume anomalies with Monte Carlo. I do not remember what it is called exactly, but I have really appreciated that feature because it helps identify potential issues that might be hard to identify otherwise, such as if the data source has changed without my awareness, even though the schema did not change. Monte Carlo can still detect something that changed upstream and is not as obvious. Things such as that have been very helpful with any kind of volume anomaly detection.
As far as the volume anomaly detection in Monte Carlo, this is something fairly recent. I helped develop brand new models for our Genesis phone call system data. Something had changed upstream where some of the API calls were doing something a little differently. There was no schema change, but the volume increased dramatically overnight, which prompted discussions with the upstream data warehouse team to identify what those changes were so we could accommodate them in the analytics layer.
I think improved reliability with Monte Carlo has been valuable. Our team monitors the channel that feeds us any alerts on our models that we are responsible for. While we have not gone through every single alert yet, it has been very helpful for each of us to keep an eye on the models that we personally are responsible for. The alerting has been great, as well as some of the metrics, such as being able to identify tests that are failing. It has been helpful to identify those immediately rather than having to dig into the data and take the time to do the data discovery to find out what issues there are.
What needs improvement?
Having used Metaplane and Elementary and worked with those other tools, I think Monte Carlo had a gap or something that was not as strong, specifically around custom feature development. I mentioned that I could not give specific examples, and I do not think I could do that now, but Metaplane specifically had a strong area with that. Additionally, the UI in Monte Carlo can be a little overwhelming sometimes. That is another thing to add to my comments, but overall, I think it is fairly easy to use.
I think that all those areas in Monte Carlo, the specific examples mentioned such as documentation, have been fine, as well as the support as far as I know. I have not had to work specifically with issues regarding support. I have not had to work hard with the product itself as far as getting support, but overall it has been good. I have been able to find help with issues in the past through the developer documentation. For me, the only thing that feels a little overwhelming when using Monte Carlo is the AI, which is something I appreciated about Metaplane because it felt more user-friendly and did not feel so overwhelming.
For how long have I used the solution?
I have been using Monte Carlo for probably three to four years collectively. I am using it currently at HubSpot and I have used it at two other previous positions.
Which solution did I use previously and why did I switch?
We have tried different tooling, and I was part of POC work for trying out different tools such as Elementary and other related tools. We found that Monte Carlo ended up being the best for us because it was able to do pretty much everything that the other competing tools were able to do for us, as well as being pretty entrenched already in our data architecture. Another reason I think we ended up staying, which I think says a lot about Monte Carlo in general, is that they were able to work with us on pricing and assure that we were happy customers. As a large organization, we were pretty happy with the pricing and the negotiation that Monte Carlo was willing to do.
What other advice do I have?
I would probably give Monte Carlo an eight out of 10. I think it is very usable and beneficial overall for data monitoring. I really appreciate things such as the time since last update, which is easy to find, row count changes, and even some of the orchestration, such as execution time, which has been really cool. There are some really cool monitors that come out of the box. Monte Carlo has some other cool things such as field lineage and table lineage that have been pretty cool to use. I can use it right within that tool and do not have to go to Atlan to look at that. Setting up personal dashboards has been probably a little harder to do compared to Metaplane or Elementary, but at least the functionality is there. That is kind of the reason I would give Monte Carlo an eight out of 10.
I think Monte Carlo's AI capabilities are fine overall. Monte Carlo seems to be doing what it needs to be doing as far as keeping the data secure in its product, even when using AI. I do not have much to say about that.
I personally have not used the AI tool within Monte Carlo much, so I probably cannot say I have a definitive opinion on that. I believe I might have used it once in the past and it helped me figure out what I needed to know as far as using some of the tooling in it, but I do not remember what I was trying to use at the time.
The first thing I would suggest to somebody brand new to Monte Carlo is to spend the time understanding monitors in general. That is for me the bread and butter of what the tool can do. I would suggest spending time in the developer documentation, understanding what each monitor does, some of the out of the box monitors such as volume testing, row count changes, and similar things. I think utilizing some of the custom fields or custom testing is really powerful if you know what you are doing and know how to implement that in a smart way. While the capability is there, I would suggest people really understand those two things specifically.
I would recommend Monte Carlo overall. I gave it an eight out of 10 because I think it does its job in helping us monitor our data products and making sure that nothing is breaking that we are not aware of, as well as being actively aware of any anomalies that might exist.
Effortless Setup, Fast Data Insights, and a Friendly UI
Reliable Anomaly Detection with Learning Curve
Monte Carlo Makes Data Quality Monitoring and Troubleshooting Easy
Seamless Data Monitoring and Alerting with Monte Carlo
Continuously Evolving AI Features That Keep Getting Better
Effortless Setup, Useful Alerting, But Alert Noise Needs Refinement
Automated monitoring has reduced manual checks and flags data incidents with precise alerts
What is our primary use case?
Monte Carlo is a data observability tool that can help track data volume changes and flag incidents if there is any unusual activity on a table, data models, or any jobs. For instance, if something unusual happens on an ETL job, it will raise an incident and send alerts via integrated platforms like Teams application and emails.
Monte Carlo also helps in monitoring applications like ServiceNow, Jira, and Snowflake by establishing connectivity with them. The solution is beneficial for various scenarios, such as when sudden deviations from normal data patterns occur. For instance, if a cloud data warehouse like Snowflake experiences an unexpected change, Monte Carlo flags incidents immediately. Its AI agent enhances troubleshooting by providing analytical insights. Besides, integrations with other applications, such as ServiceNow and Jira, ensure an end-to-end alerting mechanism.
On an account level, it builds monitoring capabilities across vast data sources. It also supports triggering alerts via ServiceNow by using webhooks, allowing associates to take immediate action. Recently, Monte Carlo introduced an AI agent to aid in troubleshooting, ensuring only main production layers are analyzed. This backend troubleshooting does not grant complete access to all layers but remains highly effective in problem-solving.
What is most valuable?
The most valuable aspect of Monte Carlo's observability feature is its automation of the monitoring processes, which eliminates the need for an individual to manually monitor numerous models or tables. It flags issues with precision and ensures proactive resolutions only on the affected components, thereby enhancing efficiency vastly.
Monte Carlo's scalable nature further bolsters its value proposition. Once integrations are established, future model updates are automatically captured without additional setup costs or actions. Given that the data platform's needs perpetually grow, Monte Carlo provides seamless adaptability.
The software manages data auditing and monitoring across platforms like Snowflake with its robust algorithms. By analyzing metadata over an extended period, Monte Carlo's flagging system, based on deviations from historical averages, ensures precise incident identification. Its ability to utilize custom monitors further extends its value, as users can implement logic-based rules and receive targeted alerts.
The introduction of a performance tab greatly aids optimization, visually displaying runtime graphs to identify model issues quickly. Monte Carlo's near perfection in accuracy ensures every flag corresponds to a genuine issue, attested by its consistent performance over time.
Monte Carlo's AI troubleshooting agent, which mimics human oversight through tiered analysis, provides ample support in incident resolution. This ensures incidents are well-documented, analyzed, and tackled despite limited access to all data layers.
What needs improvement?
While Monte Carlo frequently updates its UI platform, the changes might pose adaptation challenges for long-time users, as the continual evolution is not always intuitive. Additionally, occasional latency hampers efficient access during critical incidents, leading to potential misses of high-priority alerts.
In terms of data accessibility, granting read-level access to non-sensitive data layers would enhance insights for users significantly. Users currently experience limitations in data layer access, such as bronze and silver data layers containing essential business logics.
Sharing metadata with clients could bolster Monte Carlo's analytical capacity, allowing clients to draw deeper insights from shared data.
For how long have I used the solution?
I have been using this solution for four years.
What do I think about the stability of the solution?
I have never noticed something which Monte Carlo flagged that was not relevant to the issue. The accuracy is 100% from what I have noticed. I have never noticed any issues where something is actually happening on a data warehouse, but Monte Carlo still flags it as an issue. I have not seen these kinds of scenarios with Monte Carlo while working. The accuracy is consistently 100%.
What do I think about the scalability of the solution?
I can say Monte Carlo's scalability is at approximately 90%. Once we establish the integration in the future, we can enable an option to capture the future models. We do not need to work again and again. Once we establish the connectivity, we do not need to work on adding new tables. It should be a one-time effort. This approach is more scalable. Data platforms will not stop growing from their beginning state. They will always expand. Monte Carlo demonstrates scalability in adopting new models automatically, which should serve organizations well.
How are customer service and support?
We have connected to Monte Carlo regarding a few things. When we are establishing any new connection and face any issues, we reach out to them for technical support.
The technical support is fine. They will respond within two to three hours, but the solution may take some time, ranging from 24 to 48 hours. Technical support is satisfactory from them. Even though the product application team is not that much larger, they are still giving better support.
How was the initial setup?
When onboarding Monte Carlo, you can review the documentation for whatever source you want to connect from Monte Carlo. There you can find more use cases.
What other advice do I have?
When I started working with Monte Carlo, I did not see as many features as currently exist. Previously, the product did not have troubleshooting agents. Also, when I started working, it did not have a performance tab. The performance tab shows performance in a graphical way, allowing me to easily review the model and check the average run time. If there is any unexpected spike that happens for a specific day, I can see that.
For technical support, I would give it eight out of ten. Currently, for my account, we are not giving all layers access to Monte Carlo. We are only giving access to the main golden layer. We are not giving access to the bronze layer and silver layer because they contain business logics.
I am not certain about billing information because that is at an account level, and clients would be aware of billing information rather than myself.
My overall rating for this product is eight out of ten.