AWS for Industries

How researchers can analyze biomedical literature in minutes using Amazon Bedrock

Thousands of biomedical research papers are published every day, and no single researcher can keep pace without automated analysis. For researchers in genomics and biotechnology, staying current means manually reading through dozens of dense PDFs, extracting relevant findings about genetic variations, biomarkers, and disease associations one paper at a time. It’s slow, error-prone, and ultimately limits how much literature a researcher can realistically incorporate into their work. When the goal is discovering life-saving treatments, that bottleneck matters.

This open-source solution automates the entire process. A researcher uploads a PDF, and within minutes the system determines whether the paper is relevant, extracts biomarkers and gene variants, answers predefined research questions, and produces an evidence-based summary with citations mapped back to the original document. The project is deployable to any AWS account using the included CDK application.

Explore the full source on GitHub

The Challenge

Scientific papers are inherently complex. A typical genomics paper spans 15–20 pages of multi-column layouts packed with specialized terminology, data tables, statistical analyses, and references that can number in the hundreds. These papers discuss multiple genes and their variants, intricate biological pathways, and disease associations that require deep domain knowledge to interpret. Researchers don’t just read them. They must extract specific data points, evaluate methodology, and synthesize findings across many sources.

You need a system that can handle this complexity end to end: ingest raw PDFs, parse their structure reliably, determine relevance to a target condition, pull out biomarkers, answer specific research questions, and generate summaries grounded in evidence from the text. All without manual intervention, and flexible enough to adapt to different research domains through configuration rather than code changes.

How the Pipeline Works

The pipeline is orchestrated by AWS Step Functions and flows through five stages. When a researcher uploads a PDF through the web interface, Amazon Bedrock Data Automation (BDA) extracts structured content from the document, preserving tables, section boundaries, and text flow as clean markdown. The extracted content is indexed in Amazon OpenSearch Serverless for semantic search and stored in Amazon Simple Storage Service (Amazon S3) for downstream processing. In production, document ingestion would typically be automated from sources such as PubMed feeds, institutional repositories, or shared drives.

Next, a filtering step uses Anthropic’s Claude via Amazon Bedrock to determine whether the paper is relevant to the researcher’s target condition. Papers that pass the filter move into parallel analysis branches: one extracts biomarkers and gene variants, while two others run concurrently to analyze study outcomes and generate comprehensive answers to predefined research questions. The prompts driving each step are stored in Amazon DynamoDB, so researchers can tune the analysis for different conditions and genes without touching code.

Finally, the system maps quoted evidence back to locations in the original PDF, giving researchers a direct path from any summary claim to its source material. All results are accessible through a React-based frontend, which also supports browsing by gene, bookmarking papers, and tracking analysis history.

Figure 1 Pipeline architecture for automated biomedical literature analysis

Figure 1: Pipeline architecture for automated biomedical literature analysis

How Amazon Bedrock Data Automation simplified document extraction

Early versions of this pipeline used custom Python code to extract text from PDFs. PDF text extraction is notoriously unreliable. Multi-column layouts break text ordering, tables lose their structure, figures are ignored entirely, and every journal formats things differently. Significant development time went into writing and maintaining format-specific parsing rules, and the results were still inconsistent.

Switching to Amazon Bedrock Data Automation (BDA) eliminated that entire problem. BDA is a managed service that extracts content from documents at both page-level and element-level granularity, producing structured markdown output. Tables come through as actual tables. Section boundaries are preserved. Bounding box coordinates are attached to every extracted element, which is what makes the evidence-mapping feature possible. The system can point a researcher to the exact location in the PDF where a claim originates.

BDA also generates AI-powered descriptions for complex elements like figures and charts, enriching the content available to downstream LLM analysis. The practical impact was significant: hundreds of lines of brittle, format-specific parsing code were replaced with a single managed API call, and extraction quality improved markedly across the journal formats tested.

The BDA workflow runs as its own Step Function, submitting documents, polling for completion, and routing output to a function that creates embeddings and indexes content in OpenSearch. This clean separation means the extraction stage can evolve independently from the analysis logic.

The generative AI approach

For the AI analysis components, two approaches were evaluated: an iterative RAG-based design using agent frameworks like LangGraph, and a simpler single-step approach where the full document, prompts, and questions are sent to the model in one call. The single-step approach was chosen. For the average 16-page papers in the corpus, it delivered high-quality results without the complexity of managing retrieval chains and agent loops. Note that this approach is subject to the model’s context window limits. Currently approximately 200K tokens for Claude, which accommodates most individual research papers but may require chunking or a RAG-based fallback for exceptionally long documents. The Amazon Bedrock messaging API is called directly. No additional frameworks needed.

Each analysis step (filtering, biomarker extraction, outcome analysis, and summarization) runs as a separate function with its own prompt template pulled from DynamoDB at runtime. The summarization step is the most involved: it groups research questions into batches of four, makes multiple LLM calls to stay within token limits, and assembles the results into a structured response. Anthropic’s Claude powers all analysis steps, accessed through an Amazon Bedrock cross-region inference profile.

Results

Researchers report that manually processing a single paper, reading, extracting biomarkers, and synthesizing findings, typically takes 3–5 days. This pipeline completes the same work automatically in under 5 minutes. Filtering produces reliable, evidence-based relevance decisions. Amazon Bedrock Data Automation (BDA) extraction captures document structure with high accuracy, including tables, figures, and section boundaries that custom parsers struggled with. Summarization generates accurate, evidence-grounded answers to research questions. Once a paper is uploaded, the pipeline runs end to end without manual intervention, freeing researchers to focus on interpretation and decision-making rather than data extraction.

Getting Started

The project is open source. The repository includes the full CDK infrastructure, all Lambda functions, Step Function definitions, prompt templates, and a React-based frontend. To deploy it to your own AWS account, you’ll need access to Amazon Bedrock (Claude models and Data Automation), the AWS CDK CLI, and Python 3.14 or later for the Lambda runtime. The entire solution is deployed as a single CDK application and runs fully serverless.

See the deployment guide in the GitHub repository

Conclusion

By combining Amazon Bedrock Data Automation for document extraction with Amazon Bedrock foundation models for intelligent analysis, this pipeline transforms how researchers interact with scientific literature. The serverless architecture keeps operational overhead minimal. The prompt-driven design makes it adaptable. Swap the prompts and questions in DynamoDB, and the same pipeline works for a different research domain.

This pipeline was built for genomic research, and its greatest value lies in accelerating the kind of literature analysis that directly supports drug discovery, biomarker validation, and clinical decision-making. The underlying architecture, serverless extraction paired with prompt-driven analysis, is adaptable to other document-heavy domains as well. Contributions and feedback are welcome in the GitHub repository.

Deepansha Tiwari

Deepansha Tiwari

Deepansha Tiwari is a Prototyping Architect on the AWS PACE team, where she helps customers tackle complex, high-impact challenges by rapidly building AI/ML solutions and applications on AWS. With 10+ years of experience spanning software engineering, cloud architecture, and AI/ML, she brings a hands-on approach to designing and building solutions that turn ambitious ideas into working applications. She leads innovation across customer engagements, partnering with teams across industries to take complex business challenges from concept through prototyping and toward scalable solutions.

Jeff Harman

Jeff Harman

Jeff Harman is a Senior AI Engineer at AWS specializing in Generative AI and AI-assisted software development. He is an open-source leader on AWS’s AI-DLC framework (awslabs/aidlc-workflows, ~5K GitHub stars, GA across seven agentic coding harnesses) and leads the AI/ML Knowledge Worker Persona Track, turning complex business challenges into production-ready AI systems with measurable outcomes. He combines 30+ years of enterprise architecture experience with hands-on engineering, transforming complex business requirements into scalable AI-powered systems.