Using a statistic population health model generator, this data set is made up of highly realistic, but synthetic patient data that can be used for testing purposes without risk of disclosing PHI (protected health information).
This dataset is for single organization use only. Please contact us for more information on synthetic datasets for multi-partner use, FHIR server options available for testing, or hand curated datasets to meet your needs.
This data pack is a synthetic healthcare dataset comprised of 1000 patients with 1 year of longitudinal history.
Using a Monte Carlo simulation technique, each synthetic record is modeled to emulate clinically relevant treatment scenarios. During the generation synthetic patients progress through a series of healthcare encounters. It is these encounters and their events that are used to generate the dataset which is comprised of healthcare data messages across several HL7 messaging standards.
These records are highly realistic and even include gaps of information like a patient record in a real-world healthcare ecosystem.
Common conditions that may be contained in the data pack:
• Appendicitis
• Cancer
• Covid
• Deep venous thrombosis
• Diabetes
• Food Insecurity (SDOH)
• Hypertension
• Osteoporosis
• Pregnancy
• Pulmonary embolism
• STIs
• Zika
HL7 Message standards output that may be included for a synthetic patient record:
ADT
Admission, Discharge, Transfer (ADT) messages are used to communicate patient demographics, visit information and patient state at a healthcare facility.
This synthetic data set contains the following number of synthetic Admit, Discharge and Transfer (ADT) messages in HL7 messaging standard version 2.6 with the following event types:
• A01 - Admit / visit notification
• A03 - Discharge/end visit
• A04 - Register a patient
Message count in data set:
1,410 total A01
1,410 total A03
6,198 total A04
VXU
Unsolicited Vaccination Update (VXU) messages are used to receive and send patient’s vaccination information.
This synthetic data set contains synthetic Unsolicited Vaccination Record (VXU) messages in HL7 messaging standard version 2.5.1 with an event type of V04.
Message count in data set: 2,273
ORU
ORUs are unsolicited transmission of an observation message designed contain information about a patient's clinical observations and are used for transmitting patient’s laboratory results to other systems.
This synthetic data set contains synthetic Observation Result (ORU) messages in HL7 messaging standard version 2.5.1 with an event type of R01.
Message count in data set: 948
CCD
Continuity of Care Documents (CCD) are XML based markup standard built using HL7 Clinical Document Architecture (CDA) elements. CCD’s carry summary information about the patient within the broader context of the personal health record.
Current data fields in CCD’s:
• Patient demographics
• Medications
• Allergies
• Encounters
• Problem lists
• Diagnosis
• Lab results
• Immunization
• Social History
Message count in data set: 7,608
FHIR
Fast Healthcare Interoperability Resources (FHIR) is a modern standard for exchanging healthcare information electronically. FHIR leverages web standards like HTTP, RESTful APIs, and JSON to enable seamless communication between different healthcare systems, applications, and devices.
FHIR facilitates interoperability by providing a framework for representing and exchanging clinical data in a structured, standardized format, allowing healthcare stakeholders to easily access and share patient information across disparate systems, leading to improved care coordination, streamlined workflows, and enhanced patient outcomes.
The synthetic patient records generated by our statistic population health model generator are output in JSON FHIR version R4 resources.
Message count in data set: 3,769 bundles containing an average of 100 FHIR resources in each bundle (~376,900 total FHIR resources)
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on the duration and terms of your contract with the vendor. This entitles you to a specified quantity of use for the contract duration. If you choose not to renew or replace your contract before it ends, access to these entitlements will expire.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
This listing uses a single pricing dimension: Product Access, billed in units under a contract. You buy access to one synthetic data pack containing 1,000 patients with one year of longitudinal data. There are no tiers or instance sizes to choose between. Pricing does not scale by usage or add-ons. You pay for access to this fixed dataset, which is fully synthetic and contains no real patient information. To adjust the volume of data, you purchase the number of units that matches your need.
Top-of-mind questions for buyers
What does one unit of Product Access include for this data pack?
One unit grants access to a synthetic data pack of 1,000 patients with one year of longitudinal data. The data is fully synthetic, containing no real patient information. It includes clinical and claims data across several HL7 messaging standards, such as ADT, ORU, VXU, CCD, and FHIR.
How do I get more than 1,000 patients or additional years of data?
Pricing does not scale by usage or tiers. To increase volume, you purchase more units of Product Access, each granting another fixed pack of 1,000 patients with one year of longitudinal data. Each unit is billed the same way under the contract.
What clinical scenarios and conditions does the synthetic data cover?
The data uses Monte Carlo simulation to model clinically relevant treatment scenarios. Synthetic patients progress through healthcare encounters over the year. Common conditions may include appendicitis, deep venous thrombosis, food insecurity, hypertension, osteoporosis, and pulmonary embolism. This supports interoperability testing, application testing, and model training.
interoperabilityinstitute.org
Helpful?
Vendor refund policy
Refunds are not offered for this product.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
Using a statistic population health model generator, this data set is made up of highly realistic, but synthetic patient data that can be used for testing purposes without risk of disclosing PHI (protected health information).
This dataset is for single organization use only. Please contact us for more information on synthetic datasets for multi-partner use, FHIR server options available for testing, or hand curated datasets to meet your needs.
Using a statistic population health model generator, this data set is made up of highly realistic, but synthetic patient data that can be used for testing purposes without risk of disclosing PHI (protected health information).
This dataset is for single organization use only. Please contact us for more information on synthetic datasets for multi-partner use, FHIR server options available for testing, or hand curated datasets to meet your needs.
DataMasque is a data masking platform that transforms sensitive production data into realistic, fully functional and privacy-compliant datasets.
Its synthetically identical data preserves the statistical characteristics, complexity and edge cases of your original data while maintaining referential integrity and data consistency - without sensitive information ever leaving your secure environment.
DataMasque helps enterprises accelerate development, testing, analytics and AI with synthetically identical customer data. Fully functional, realistic and privacy compliant.
DataMasque helps enterprises accelerate development, testing, analytics and AI with synthetically identical customer data. Fully functional, realistic and privacy compliant.