AWS for Industries
Trust-Earned Autonomy: How Texas Capital Bank demonstrates an AI agent that commits to the core banking ledger, then reverses itself
To write, or not to write: a bank’s dilemma
Texas Capital Bank is a full-service financial services firm headquartered in Dallas, Texas, serving commercial, consumer, and institutional clients. Its technology organization builds internal platforms across several areas, including document intelligence for lending. Processing a commercial loan package is slow, manual work: an analyst reads a 200-plus-page package and extracts financial obligations, servicing requirements, and covenant terms by hand, a task that can take hours per package and is easy to get wrong. Texas Capital Bank already runs a production capability that automates the extraction itself. The question this post explores is the next step beyond extraction: once the software can read the document, when should it be allowed to act on what it found, and how do you take that permission back if it stops being trustworthy?
The gap between an AI that recommends and an AI that commits
Most organizations are comfortable letting a language model draft a document, summarize a contract, or suggest a value. They get nervous when the same model presses the button, the one that commits a transaction, marks a loan approved, or moves money.
The moment an AI commits to a system of record, the stakes change. A reversible suggestion costs a re-run if it is wrong. A committed transaction, a loan marked approved or a wire sent, costs an audit finding, a customer remediation, or a regulator conversation. That is the line this pattern is built around. The demonstration in this post commits to a real ledger: the hard part only shows up once the write is real.
The usual responses are unsatisfying. Route every AI output to a human queue, and the result is a faster typewriter, not automation. Gate on model self-confidence (“commit if the model is 95% sure”), and the decision rests on a self-reported number whose failure mode is silent: confidence stays high while accuracy drifts, and nobody notices until something breaks.
Texas Capital Bank started from a different question, one Vibhu Bhatnagar frames this way:
“How does an automated system earn the right to act on its own, and how do you take that right back, cleanly, when it no longer deserves it?”
This post describes the pattern the bank built to explore that question, which its team calls Trust-Earned Autonomy, and the AWS architecture that runs it. Texas Capital Bank already has a document-processing capability in production; the Trust-Earned Autonomy pattern described here is the next iteration, and at this stage it is a demonstration rather than a production workload. Its purpose is to validate the governance pattern, how an automated decision earns the right to act and can be safely reversed, so that it can inform a future production rollout. To show the pattern safely, the team runs it against a representative open-source core banking system (Apache Fineract) using synthetic data on demo-scale infrastructure, so that ledger writes, and their reversal, can be tested without touching production data. Letting an agent produce an output, a draft, a summary, a recommendation, is by now well-trodden. Letting it commit that output to a system of record, and then revoking that authority and undoing what it committed, is not. That harder, less-demonstrated half is the focus here.
How it works
The system reads an incoming document and applies a recipe, a set of approval rules written in plain English. Each recipe carries a trust score that rises as its decisions agree with human reviewers over time. A recipe with a low score can only suggest; a human still approves every commit. A recipe that has consistently agreed with reviewers earns the right to commit on its own, but only behind a gate of safety checks. The system recomputes trust every night. If a recipe’s score falls, it is demoted and loses that right, and the system reverses the commits it made while it held authority it no longer deserves. That reversal, on a real ledger, is the part this post focuses on.
For example, a 26-page commercial credit agreement arrives for a $1.5 million facility. A recipe reads the borrower name and the facility amount, looks the borrower up in the core banking system, confirms the amount matches the loan on record within a small tolerance, and proposes to approve. Whether that approval commits on its own, waits for a human, or is later reversed depends entirely on how much trust the recipe has earned. We follow this one agreement through each stage below.
Figure 1. How the pattern works, end to end, for a single document.
Key components
Figure 1 showed the five stages a document moves through. Four components make those stages work, and none is individually novel; the contribution is the combination, and the rollback behavior that this combination makes possible.
1. Versioned prose recipes
The unit of decision logic (the set of rules that decide whether a document is approved) is not code. It is a prose recipe: the rules are written in plain English by a bank analyst, not programmed by a developer. Each recipe is versioned, meaning every published revision is tracked, so any decision the system makes can be traced back to the exact wording of the rule that produced it. A recipe targets one document type (for example, a credit agreement). For example, a recipe for credit agreements reads the borrower name and facility amount from the document, looks the customer up in the core banking system, pulls the matched loan, checks that the document’s facility amount agrees with the loan principal within a 1% tolerance, and approves only if every step passes.
Expressing the rules as prose means a credit or risk analyst, the person accountable for the lending decision, can author and own the recipe without needing a developer. Because it is versioned and immutable, a verdict can always be traced back to the exact rule text that produced it.
2. A closed capability set
The second piece limits what the agent is able to do. Whatever a recipe’s rules say, when it runs, the agent can call only a fixed, pre-approved set of ten tools and nothing else. Seven of them only read information (for example, read a field from the document, look a customer up in the core banking system, or check a compliance requirement); one evaluates a math expression in a sandbox; and two record the outcome (emit a finding, or decide the verdict). The full set is: extracted_field, peer_doc_field, find_field_anywhere, policy_lookup, screening_lookup, banking_lookup, compliance_requirement, evaluate_expression (sandboxed, with no ability to run arbitrary code), emit_finding, and decide_outcome.
A recipe author cannot register an eleventh tool. A prompt cannot expand the set. Adding a capability is a framework code change that goes through review. The payoff is an audit property most agent stacks cannot offer: the question “could the agent have done X?” is answerable from the capability registry at a specific git commit, not by reconstructing which tools happened to be attached on a given day.
The closure is over what the agent can do, not which model does the reasoning. The same recipe runs against any model that can reason over the tool surface, which keeps the door open as models evolve.
3. A nightly trust score
For every recipe and version, the system maintains a score from 0 to 100, recomputed nightly from the recipe’s measured track record:
trust = 20 · golden_coverage
+ 30 · golden_replay_pass_rate
+ 40 · reviewer_agreement_rate
+ 10 · min(1, production_runs / 200)
Here, golden_coverage is the share of a curated benchmark set the recipe handles, and golden_replay_pass_rate is how often it reproduces the correct decision on that set. Reviewer agreement carries the most weight (40 points) by design: it is the only term grounded in human judgment rather than the system’s assessment of itself. The volume term (production_runs) prevents a recipe from earning high trust on a handful of lucky cases.
4. A multi-predicate gate
Reaching the autonomous tier does not mean a run commits. Every would-be autonomous commit passes a function gate, evaluated per run. Eight predicates are checked in engine order. The recipe’s tier must still be autonomous. The per-recipe kill switch must be off. The verdict must be approve. The run’s confidence must clear a per-recipe floor (for example 0.95), used here as one predicate among several rather than as the sole basis for the decision. There must be no open blocking or needs-review findings. The recipe must be under its daily commit cap. The outcome amount must be under the recipe’s amount cap, with a Cedar-enforced ceiling at the gateway regardless of what any prompt asks for. And the run must not be sampled for mandatory human review. If any predicate fails, the run routes to a human with the failure reason attached. The failure mode is always “a human looks,” never “the commit silently proceeds.”
Note the relationship to the confidence gating dismissed earlier: confidence appears here as one predicate among eight; it is never the only decider.
Autonomy is per run, not per recipe. Figure 2 shows the gate in full.

Figure 2. The multi-predicate gate. Reaching the autonomous tier does not mean a run commits.
No single predicate is load-bearing. Each one prevents a different failure mode, and a failure of any one degrades safely to human review.
Autonomy as a state machine
A recipe’s autonomy is a function of measured history, not a configuration flag. The three tiers and the transitions between them are shown below:
Figure 3. Recipe trust tiers and the promotion and demotion transitions.
| Tier | What it can do | How it gets there |
|---|---|---|
| Learning | Processes documents; every output goes to human review. No autonomous commits. | Entry tier, or demotion target after a sharp regression. |
| Competent | Processes; commits still require human review. | From Learning: score ≥ 70 and ≥ 50 production runs. |
| Autonomous | May auto-commit, but only behind the per-run gate. | From Competent: score ≥ 90 and ≥ 200 production runs. |
Promotion is deliberately slow: a recipe climbs one rung at a time, and only with both a high score and a long track record. Demotion is deliberately fast:
- Score falls below the autonomous floor of 80 → drop to Competent.
- A single-cycle regression of 15 points or more → drop straight to Learning.
- A golden-replay failure → drop to Learning, unconditionally.
The asymmetry is the point. As Bhatnagar puts it:
“Trust is slow to grant and fast to revoke.”
The hard half: demote-with-rollback
Losing autonomy going forward is necessary but not sufficient. By the time a recipe regresses, it has already committed decisions to the system of record under authority it no longer deserves. Going-forward demotion leaves those commits in place.
So the moment a recipe is demoted, the system does something few agent stacks do: it finds the commits it made under the now-revoked tier and reverses them.
The recipe had earned the top tier. A banking cross-check recipe sat at a trust score of 87, autonomous. A 26-page commercial credit agreement arrived. The system pulled the $1.5 million principal, the Term SOFR pricing grid, and the covenant thresholds off pages scattered through the document, for cents of compute.
It acted. The recipe resolved the borrower against the core banking system, cross-checked the facility, and, because every gate predicate passed, committed. The commit is not an opinion or a queue event: the engine called Fineract’s disburseLoan, and the loan moved from Approved to Active. Real money, on the books. Every banking_lookup, the evaluate_expression cross-check, and the engine-invoked commit are on the agent’s audit trail with full gateway and Okta JSON Web Token (JWT) wire detail.
Its trust caught up with it. Reviewers had been quietly disagreeing with the recipe’s recent calls. The nightly scorer recomputed the score from that record, and it fell from 87 to 79, crossing the autonomous floor of 80. The recipe was demoted one tier, from autonomous to competent, and lost the right to commit on its own, immediately.
Then it undid itself. On demotion, the recipe engine ran a rollback sweep in-process, deliberately reusing the same gateway credentials, wire tracing, and AWS Identity and Access Management (IAM) boundary as the forward path. It found the one commit made under the revoked tier and posted a full-principal compensating repayment to Fineract (banking_reverse → makeLoanRepayment), with the original disbursal transaction id in the memo so an auditor can trace the pair. The loan settled to Closed, balance zero. No human pressed either button.
The live ledger now carries two symmetric, real writes: the disbursal on commit and the repayment on rollback, both traceable transaction to transaction. Figure 4 shows the rollback as five ordered steps.
Figure 4. Demote-with-rollback sequence. Real ledger writes; synthetic data.
Where rollback stops
Rollback is the system’s to issue only while the action is still inside its own control surface. A banking transfer reversal, a dossier disposition marked withdrawn, a compliance verdict invalidated and re-enqueued: all reversible. Once an action crosses an external settlement layer (a FedWire wire, a SWIFT instruction, a payment under network irrevocability rules) the system can record a reversal request but cannot unilaterally undo it. The pattern’s promise is that no autonomous commit crosses that boundary unreviewed. That is what the amount cap and the daily ceiling are for.
The AWS architecture
The pattern is abstract by intent. It is grounded in a specific system Texas Capital Bank built and operates end to end on AWS. Figure 5 shows the end-to-end system: a document platform VPC that hosts the processing pipeline, the agent runtimes, and the data layer, and a separate banking VPC that hosts the Apache Fineract core banking system. The two estates meet only over AWS private networking.
Figure 5. Trust-Earned Autonomy system overview. Synthetic data; demo-scale infrastructure.
How a document flows through the system
The numbered steps below correspond to the flow annotations in Figure 5.
- A user request reaches Amazon Route 53 and is routed to an Application Load Balancer.
- The Application Load Balancer performs OpenID Connect (OIDC) authentication against Okta before any AWS Lambda function runs. The single-page application is served by a frontend Lambda function rather than from Amazon Simple Storage Service (Amazon S3) or Amazon CloudFront.
- An authenticated user uploads a document to Amazon S3, the entry point for the pipeline.
- A trigger AWS Lambda function content-addresses the document with a SHA-256 hash to deduplicate it, so the same document is never processed twice.
- A single 10 GB AWS Lambda function, built with Lambda Durable Functions, orchestrates the pipeline end to end, from routing the document to a recipe, through extraction, page indexing, compliance evaluation, and normalization, with checkpoint and replay.
- Amazon Textract performs optical character recognition, and Amazon Bedrock runs the models: Claude Haiku for routing, extraction, and normalization, and Claude Sonnet for compliance evaluation and agent reasoning. The model is selected automatically per run by a complexity heuristic.
- The pipeline writes transactional state to Amazon DynamoDB (19 tables), with audit artifacts in Amazon S3.
- Streams on four DynamoDB tables feed a stream-processor AWS Lambda function that publishes domain events (document.processed, compliance.scored, reviewer.override, baseline.updated) to Amazon EventBridge, decoupling the command side from the agent layer.
- An EventBridge adapter invokes the recipe_engine runtime on Amazon Bedrock AgentCore, which executes the published prose recipe for the document type.
- The recipe engine binds the recipe to the closed ten-tool capability set and calls the AgentCore Gateway (Model Context Protocol) for the banking lookup.
- The gateway authenticates the call with a custom Okta JSON Web Token and applies a Cedar policy in ENFORCE mode (an operator-set amount cap), then the request rides Amazon VPC Lattice into the separate banking VPC.
- An internal Network Load Balancer routes to Apache Fineract on Amazon Elastic Kubernetes Service (Amazon EKS); on an autonomous-tier commit the engine calls disburseLoan. Fineract is backed by Amazon Aurora PostgreSQL and Amazon Managed Streaming for Apache Kafka (Amazon MSK).
- In parallel, DynamoDB Streams on seven tables feed a graph-sync AWS Lambda function directly (no EventBridge hop) that materializes a relationship view into Amazon Neptune Serverless, the Command Query Responsibility Segregation (CQRS) query side, which agents query over openCypher.
The document pipeline
Texas Capital Bank moved this orchestration off AWS Step Functions in an earlier version. The pipeline is CPU-bursty (Amazon Textract fan-out across approximately 30 workers) and benefits from 10 GB / 5.7 vCPU in one function; checkpoint/replay gives the team Step Functions semantics without per-transition cost or split logs. The trade-off the team accepted: orchestration logic now lives in Python, so replay-determinism is enforced by convention and tests rather than by the state machine definition.
The agent layer
The reason a bank can let this run at all comes down to how tightly the agents are boxed in. Three agents run the system, and each is confined to its own permissions so that a fault or a prompt-injection in one cannot reach the others or the data they do not need. They are packaged in a single container image, with each agent selected at startup; the point for a security reviewer is the isolation between them, not the packaging.
- The three runtimes have distinct, least-privilege roles: recipe_engine executes recipes and runs the rollback sweep (and is the only one with Fineract credentials and gateway access); plugin_maturity tracks per-plugin accuracy and drives tier changes; and decision_synthesizer aggregates findings into reviewer-ready narratives.
Per-runtime IAM isolation is the design point to probe: the blast radius of a prompt-injected recipe run is bounded by the recipe_engine role, which can read bounded secrets and invoke one gateway but cannot touch other tables’ write paths, Amazon S3 buckets, or the other runtimes.
The banking path
The recipe engine reaches the banking system through a defense-in-depth chain already traced in the flow above: a closed capability boundary (the recipe never sees a URL or credential), the AgentCore Gateway with Okta workload-identity auth, and Amazon VPC Lattice into the banking VPC. The one detail worth drawing out is the Cedar policy engine at the gateway: in ENFORCE mode it permits reads and caps loan mutations at an operator-set amount (for example, $2 million) regardless of what any prompt asks for. The cap is a policy parameter, not a model decision.
The banking environment is a real deployment in its own VPC, provisioned by five phased AWS Cloud Development Kit (AWS CDK) stacks, with Amazon Aurora PostgreSQL and Amazon MSK backends, mirroring a bank’s separation between an application estate and a core banking estate. An Istio waypoint enforces a dual-hostname auth split: the machine-to-machine hostname the agent uses is Okta-JWT-enforced, while the human loan officer’s Mifos X console reaches the same Fineract backend under Fineract’s own credential auth. One cluster, one backend, two front doors with different identity models.
Command and query sides (CQRS)
Amazon DynamoDB owns every transactional write. A separate query side materializes a relationship view into Amazon Neptune Serverless, so the agent can answer questions the command store cannot cheaply serve, such as which documents reference a given Fineract client. Splitting the write path from this read view (CQRS) keeps each side simple.
Every egress path is an AWS PrivateLink VPC endpoint; there is no NAT gateway anywhere. The trade-off the team accepted: every new AWS dependency needs an endpoint added in AWS CDK.
What the team measured
Measured on a labeled 8-document benchmark, reproducible from the eval harness:
| Metric | Result |
|---|---|
| Classification accuracy | 100% (8/8) |
| Extraction precision (surfaced fields) | 92.6% (88/95) |
| Extraction coverage | 66.9% |
| Banking-verdict accuracy | 50% (3/6), active work |
| Cost per document (marginal) | $0.11–$0.61 (approximately $0.36 typical for a 3-page agreement) |
| Fixed floor | approximately $115/month (Neptune Serverless) |
The team is deliberately reporting the banking-verdict number while it is still low. It is the metric the whole pattern depends on, and it is exactly the case the trust model is built to handle: a recipe with 50% verdict accuracy does not reach, and would not stay in, the autonomous tier, because reviewer disagreement pulls its score below the floor. The system working as designed means not trusting that recipe yet.
Routing the high-volume stages (classification, extraction support, normalization) through Claude Haiku rather than Claude Opus 4.8 is what drives the roughly 10x lower per-document cost, while reserving Claude Sonnet for the reasoning-heavy compliance and agent steps.
Why this matters now
Two things converge that make the pattern worth naming.
The cost of inference has crossed a threshold. Running a recipe engine for a fraction of a cent per execution means the trust score is computable across the whole document population, not just the borderline cases. Volume is what makes the measurement real.
The regulatory environment is hardening. The EU AI Act expects high-risk systems to keep logs sufficient to backtrack decisions. US Treasury BSA/AML modernization has surfaced explainability expectations for automated decisions. ISO 42001 wants documented mechanisms for monitoring AI performance over time. A pattern that provides a versioned rule for every decision, a measurable trust score with explicit components, a documented gate, and a defined rollback procedure is aligned with what these frameworks ask. Regulators will not ask whether the system is intelligent; they will ask how the institution knew when to trust it, and what it did when the trust was wrong.
How to think about this pattern
- Human judgment stays in the loop. A fraction of cases is routed to humans on an ongoing basis, even at the autonomous tier. Those cases are what the trust score is measured against. Remove the carve-out and drift becomes undetectable.
- Safe only in bounded environments. The pattern relies on a closed capability set; it works only where side effects can be enumerated in advance. A general-purpose assistant does not have this property.
- Applicable beyond document processing. The pattern generalizes to any narrow decision domain where an agent produces a verdict that touches a system of record: fraud screening, service refunds, infrastructure remediation, medical claims, supply-chain reordering.
Where this is going
The system runs today on demo-scale infrastructure with synthetic data. The road to production will require: multi-tenant isolation, production networking, throughput sizing, and, the item the team cares most about, a model-approval layer that scores a candidate model before it ever enters a recipe, complementing the recipe-behavior trust score the platform already has.
Conclusion
Trust-Earned Autonomy answers a question most production AI systems leave implicit: how does an automated decision earn the right to act, and how do you take that right back when it no longer deserves it. The pattern’s parts (versioned prose recipes, a closed capability set, a measured trust score, and a per-run gate) are individually familiar. What makes the combination useful in a regulated setting is that autonomy is granted slowly by measured behavior, revoked quickly on regression, and, the part Texas Capital Bank set out to demonstrate, reversed on the ledger when it is revoked. On a live Apache Fineract deployment with synthetic data, the system committed a disbursal under an autonomous-tier recipe and then, once that recipe’s trust fell below the floor, posted a compensating reversal on its own, with both writes traceable transaction to transaction.
For organizations building AI systems that touch a system of record, the questions transfer even where the answers do not: what is the closed set of things the agent can do, how is drift measured against human ground truth, and what is the compensating action when an autonomous decision turns out to be wrong. To explore these building blocks for your own workloads, see the resources below.



