Skip to main content

Intelligent Document Processing

GenAI Intelligent Document Processing (GenAIIDP) on AWS - FFP

Intelligent Document Processing

AWS will advise and assist Customer with the following activities, in Customer’s non-production environment. AWS will implement a generative artificial intelligence (“GenAI”) solution for Intelligent Document Processing (“IDP”), designed to automate document classification and field extraction from unstructured documents leveraging Customer’s existing assets, as set forth in more detail below (the “Solution”).

GenAI Intelligent Document Processing (GenAIIDP) on AWS

The Solution will be developed based on Customer’s-identified requirements and AWS general best practices as defined through the work-streams below (each a “Workstream”, collectively the “Project”). AWS will assist Customer to:

Workstream 1 - Project Initiation

Discovery & Planning Workshop:

      • Review Customer's document processing requirements, data quality, environment, and infrastructure
      • Demonstrate existing GenAI IDP Solution assets that will be used in the engagement
      • Identify and prioritize Customer’s business requirements
      • Identify and mutually agree on the document types, classes, and fields (In-Scope Document(s)) for the Solution to process during the Project
      • Align on Project governance to manage time and cost, and communication
      • Validate Solution deployment pre-requisites (including access to AWS accounts and services)

Workstream 2 - Project Execution

Data Assessment:

      • Review data sources, quality, and availability for one In-Scope Document
      • Specify requirements to improve data quality and coverage
      • Support Customer in establishing Customer’s data governance and security for production deployment
      • Specify requirements for document page/section classification, packet splitting, and information extraction and normalization for In-Scope Document
      • Define the evaluation criteria and metrics for user acceptance testing

Solution Implementation:

      • Deploy the Solution in Customer’s non-production environment and adapt (as necessary) to address the prioritized requirements, to include as needed:
        • Prompt engineering
        • Model fine-tuning
        • Evaluation logic for each document section class and extracted field
      • Evaluate the deployed Solution in Customer’s non-production environment against defined evaluation criteria and metrics for user acceptance testing for In-Scope Document.
      • If needed, integrate with Customer's identity provider, monitoring tools, security and governance mechanisms, and deployment processes and pipelines
      • Advise and assist Customer on their efforts to incrementally deploy Solution to production

Workstream 3 - Solution Deployment Support (done in parallel with Workstream 2):

      • Provide [XXX] training and knowledge transfer sessions to Customer teams on the Solution
      • Provide prescriptive guidance to customer on deployment of accepted sprint deliverables to production throughout the Project

Intelligent Document Processing (continued)

GenAI Intelligent Document Processing (GenAIIDP) on AWS - FFP

Workstream 4 - Project Conclusion

    • Knowledge transfer
      • Conduct knowledge training session with Customer to review features delivered
      • Discuss potential roadmap for the Solution, and show Customer how they can upgrade to subsequent releases of the Solution

Alignment

The parties will assess In-Scope Document(s) on the three criteria below to determine the complexity of such documents and ensure that any complexities are consistent with the Engagement Assumptions section below. The Solution requires a clear definition of the document taxonomy to drive efficient and effective document processing requirements. The taxonomy is defined as follows:

  • Document Modality: This refers to the format of the document, such as PDF, audio, video, or images. Knowing the modality is important to ensure the appropriate processing techniques are applied.
  • Document Type: This specifies the type of document, such as a loan application, invoice, or other business document. Understanding the document type helps determine the relevant data fields and processing requirements.
  • Document Classes and Fields: Within a given document type, there may be different sub-sections or Classes of information (as defined below). For example, in a loan application, the document Classes could include asset information, proof of address, other financial records. Identifying these document Classes, and the fields within, are crucial for extracting the right information.

Deliverables and Payment Milestones:

Subject to the acceptance process, AWS will invoice Customer, and Customer will pay, for the charges corresponding to the Payment Milestones as specified in the table below.

Payment Milestone Deliverable Deliverable Specifications Charges
Payment Milestone 1 (at the conclusion of Workstream 1) Documented Business and data requirements, and evaluation criteria Document summarizing Customer’s prioritized Business requirements and expected outcomes.
Document summarizing Solution’s data requirements for document page/section classification, packet splitting, information extraction and normalization, based on AWS general best practices.
Defined evaluation criteria and metrics for Acceptance testing
30% of Total Price
Payment Milestone 2 (at the conclusion of Workstreams 2 and 3) Solution (as Deployed in Customer’s non-production Environment) Solution pipeline deployed for up to five In-Scope Document(s) in Customer’s non-production environment [60%] of Total price
Functional/non-functional testing results and evaluation report passed
Payment Milestone 3 (at the conclusion of Workstream 4) Advisory support for deployment / adoption of the Solution Advisory support to assist customer with implementing the solution (up to 3 mo) 10% of Total Price
Document Solution source code
Runbook for the Solution covering operations and troubleshooting, based on AWS general best practices.

Assumptions:

  • In-Scope Documents will involve one document type, no more than 5 document classes and document fields, and no more than 2 document modalities.
  • Training and targeted validation data must exist, owned by the Customer, and be available already in Amazon Simple Storage Service (“Amazon S3”) buckets
  • Amazon S3 is used as data source and destination. The Solution will get the data for processing from source Amazon S3 bucket and will store the results in Amazon S3 destination bucket. The Customer will develop any required upstream or downstream integrations for the Solution.
  • Necessary AWS Services are approved for use in Customer's accounts, including AWS CloudFormation, Amazon Bedrock (all features), Amazon SageMaker, Amazon Textract, AWS Step Functions, AWS Lambda, Amazon Cognito, AWS AppSync, Amazon CloudWatch, and more as needed.
  • AWS will deploy the Solution in the Customer’s non-production AWS environment in a single AWS region.

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages