This product allows users to perform identity resolution on their own customer data from various data sources (e.g. bookings, transactions and loyalty program). The algorithm will link those different data sources to create an accurate and complete view of their customer profiles without moving any customer data outside of their aws account.
Highlights
This product allows users to perform identity resolution on their customer data inside their own aws account.
It provides field level mapping, normalization, standardization and repair out of the box. It utilizes the recent advancement of AI.
Access real-time vendor security and compliance information through their Trust Center powered by Drata or Vanta. Review certifications and security standards before purchase.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the host hour across three activity types: training, batch inference, and real-time inference. This identity resolution product uses machine learning models, so you first train the algorithm, then run inference to match and unify customer records. Each activity type is offered on three instance sizes: ml.m5.xlarge, ml.m5.2xlarge, and ml.m5.4xlarge. Larger instances provide more compute for heavier workloads. Your cost scales with the instance size you choose and the number of hours you run each job. You can mix instance sizes across training and inference as needed.
Top-of-mind questions for buyers
What counts as one billable host hour for these dimensions?
One host hour is one running hour of a single instance of the size you select. Each active instance meters separately. If you run two instances of the same size, you accrue two host hours per clock hour. Stopped instances do not accrue software host-hour charges, though AWS storage fees may still apply.
How does batch inference billing differ from real-time inference billing?
Both meter by host hour on the same instance sizes. Batch inference runs jobs against groups of records and bills only while the batch job runs. Real-time inference keeps an instance available to score records on demand, so it accrues hours whenever the endpoint stays active, even between requests.
Do I pay for both training and inference, and how do they combine?
Yes. Training and inference bill independently, each by host hour. You first run training to build the model, accruing training host hours. You then run inference to match records, accruing separate inference host hours. Training is periodic, so inference usually drives most ongoing cost for continuous identity resolution.
harpin.ai
Helpful?
Vendor refund policy
This product is offered for free. If there are any questions, please contact us for further clarifications.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker algorithm is a machine learning model that requires your training data to make predictions. Use the included training algorithm to generate your unique model artifact. Then deploy the model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Algorithm training
Before deploying the model, train it with your data using the algorithm training process. You're billed for software and SageMaker infrastructure costs only during training. Duration depends on the algorithm, instance type, and training data size. When training completes, the model artifacts save to your Amazon S3 bucket. These artifacts load into the model when you deploy for real-time inference or batch processing. For more information, see Use an Algorithm to Run a Training Job .
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
Version release notes
Update the underlying similarity model
Additional details
Inputs
Outputs
Channel specifications
Usage instructions
Sample notebooks
Inputs
Summary
Either CSV or avro or parquet type is allowed for the clustering process. If CSV input files are used, each CSV file should be comma-delimited (,) and contain a header line at the top. Each row of a CSV file represents a single record, while each column represents a field. The following are the recommended fields:
sourceRecordId, firstName, middleName, lastName, dateOfBirth, emailAddress, mobilePhone, homePhone, workPhone, postalCode, streetAddress, city, governingDistrict, ipAddress, accountId.
Limitations for input type
CSV, avro, or parquet
Input MIME type
csv, avro, parquet
Real-time inference sample input data
record_id,given_name,sur_name,dob,email,phone,zip,street_address
101,John,Smith,19901010,john@gmail.com,5051234567,92128,123 main street
202,Joe,Matthew,20001010,joe@gmail.com,8581234567,92101,456 ace street
The following table describes supported input data fields for real-time inference and batch transform.
1
2
Field name
Description
Constraints
Required
firstName
firstName: given name;
middleName: middle name;
lastName: surname
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
middleName
firstName: given name;
middleName: middle name;
lastName: surname
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
lastName
firstName: given name;
middleName: middle name;
lastName: surname
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
dateOfBirth
dateOfBirth: date of birth
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
emailAddress
emailAddress: email address
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
mobilePhone
mobilePhone: mobile phone;
homePhone: home phone;
workPhone: work phone
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
homePhone
mobilePhone: mobile phone;
homePhone: home phone;
workPhone: work phone
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
workPhone
mobilePhone: mobile phone;
homePhone: home phone;
workPhone: work phone
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
postalCode
postalCode: postal code;
streetAddress: street address;
city: city
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
No
streetAddress
postalCode: postal code;
streetAddress: street address;
city: city
Default value: BLANK
Type: FreeText
Limitations: None of the above fields are required, but there is a minimum information required in order to be able to uniquely identify each record.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Agent IAM is an Agentic AI–powered Federated Identity and Access Management solution that unifies human, machine, and AI identities across AWS, Azure, SaaS, and on-prem environments. Built natively on AWS, it centralizes authentication, authorization, and lifecycle management while automating Joiner–Mover–Leaver workflows through HRMS/ITSM integration. The platform supports modern apps (OIDC, SAML, SCIM, APIs) and legacy systems via Computer Use Agents for automated provisioning. With policy enforcement through OPA, short-lived AI access tokens, and complete audit logging, Agent IAM provides secure, compliant, and scalable identity governance for hybrid enterprises
This product has charges associated with it for seller support. This VM offers a hassle-free setup with Jupyter for projects, Jupyterhub for multiuser collaboration, a Jupyter AI extension for for LLM and Generative AI development, and preloaded popular libraries like TensorFlow and PyTorch. It also comes with pre-configured NVIDIA GPU drivers and CUDA libraries.
This product has charges associated with it for seller support. This VM offers a hassle-free setup with Jupyter for projects, Jupyterhub for multiuser collaboration, a Jupyter AI extension for for LLM and Generative AI development, and preloaded popular libraries like TensorFlow and PyTorch. It also comes with pre-configured NVIDIA GPU drivers and CUDA libraries.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.