Containers

How athenahealth modernized healthcare workloads with Amazon EKS Hybrid Nodes

If you run latency-sensitive workloads that must stay on premises (whether for data residency, compliance, or proximity to existing databases), you face a familiar trade-off: keep the workload close to the data or adopt a cloud-native operating model. This post shows how you can do both. athenahealth used Amazon Elastic Kubernetes Service (Amazon EKS) Hybrid Nodes to cut response times in half, reduce hardware and operational costs by 50 percent, and maintain a single Kubernetes operating model across its data center and the cloud.

athenahealth serves more than 160,000 healthcare providers through electronic health records (EHR) and patient portal applications. Because regulated patient data must remain on premises, the company ran these workloads on AWS Outposts. As athenahealth standardized on Kubernetes across its application portfolio, the team looked for a way to extend its Amazon EKS operating model to applications running in its data centers. Amazon EKS Hybrid Nodes provided that consistency while preserving data residency and compliance requirements.

In this post, we walk through athenahealth’s architecture, the decisions that shaped it, and the measurable outcomes. Whether you’re evaluating hybrid Kubernetes options or planning an on-premises modernization, this architecture provides a validated reference.

Challenge

athenahealth’s EHR and patient portal workloads run adjacent to sensitive patient data that must remain on premises. As part of its platform modernization, the company decided to standardize on Kubernetes and Amazon EKS. The team therefore needed an architecture that could bring the same Kubernetes tooling, deployment practices, and operational processes to its on-premises workloads while continuing to meet its data residency requirements.

The team defined three requirements for the next architecture:

  • A single Kubernetes operating model across cloud and on-premises environments: same kubectl commands, same deployment manifests, same observability stack.
  • Lower cost than dedicated on-premises hardware.
  • Flexible node placement in specific network security zones to satisfy a HITRUST security audit, rather than accepting a fixed topology.

In short, athenahealth needed on-premises compute with cloud-consistent operations. Amazon EKS Hybrid Nodes met that need through a single-cluster model: the EKS control plane runs in an AWS Region while on-premises worker nodes join the same cluster. The control plane treats each on-premises node the way it treats an Amazon Elastic Compute Cloud (Amazon EC2) instance. It schedules pods and reports node health through the same Kubernetes API. For cloud-based nodes, upgrades are driven through the EKS API or Karpenter, while hybrid node upgrades are performed using the nodeadm CLI. Because one control plane manages both environments, the team uses identical tooling everywhere, with no second control plane to operate and no separate runbooks to maintain.

Workloads on hybrid nodes integrate with the same AWS services as workloads in the cloud, including AWS Identity and Access Management (IAM) for authentication, Amazon CloudWatch for monitoring, and AWS Systems Manager for node management. With one cluster spanning both environments, athenahealth places each workload where it belongs: latency-sensitive components on premises, close to the data they consume, and supporting services in the Region, where elastic capacity costs less than standing hardware.

Architecture

The architecture separates the control plane from the data plane. The EKS control plane runs in a Region and hosts the Kubernetes API server, scheduler, and controller manager. AWS manages upgrades, patching, high availability, etcd backups, and API server scaling. This work otherwise falls to a platform team. The control plane attaches elastic network interfaces (ENIs) to subnets in athenahealth’s virtual private cloud (VPC), and the on-premises worker nodes connect through these ENIs over private connectivity using AWS Direct Connect or AWS Site-to-Site VPN.

Hybrid nodes initiate outbound connections to the VPC through Cilium Clusterwide Network Policy, while ingress traffic passes through a NetScaler appliance at the data center edge. This architecture limits the external exposure and reduces the scope of a security review. Systems Manager handles node activation and node operations, while IAM Roles Anywhere provides credential management for hybrid nodes (see Figure 1).

Workload placement. The central design decision was where each workload runs. Latency-sensitive EHR transactions run on hybrid worker nodes in the data center, adjacent to existing databases, so reads and writes stay on the local network rather than crossing to a Region. Supporting services such as logging, monitoring, and continuous integration and continuous delivery (CI/CD) run in the Region, where on-premises proximity is unnecessary and elastic capacity is readily available.

Traffic management. The architecture runs a high availability (HA) pair consisting of one “Region-only cluster” and one “on-premises hybrid cluster”. Failover happens automatically based on health checks or can be triggered manually during incidents.

Infrastructure as code. To keep configuration consistent across accounts, athenahealth uses Puppet for operating system configuration management and Crossplane to manage infrastructure declaratively, integrating it with GitOps workflows through Flux CD. A change lands as a pull request and reconciles to every account automatically, replacing manual per-environment application.

Architecture diagram showing the Amazon EKS control plane in the Region connected to on-premises hybrid worker nodes over private connectivity

Figure 1: athenahealth EKS Hybrid Nodes architecture, with the control plane in the Region and hybrid worker nodes on premises connected through a private link

Data plane layer

Most of the design work sits in the node layer, because athenahealth owns the hardware beneath the hybrid nodes. The nodes are physical or virtual machines that run the EKS-optimized custom image, register with the control plane, and receive workloads like any cloud node, while the team retains control of hardware, network, and storage.

In athenahealth’s deployment, Clinicals Macro Service (large Java Spring application pod tied to monolith/database) runs on eight Dell 16th-generation bare-metal servers with AMD EPYC processors. Each server provides 48 cores and 384 GB of RAM. The team validated Border Gateway Protocol (BGP) routing on the bare-metal servers. This lets pods on hybrid nodes communicate directly with other on-premises pods and services without an overlay network hop. Puppet handles configuration management of the operating systems and core configuration, implementing Kubernetes and Hybrid Nodes dependencies at provision time and validating the correct configuration across network zones.

Bandwidth planning. The connection back to AWS carries wide-area bandwidth, so athenahealth planned capacity early. Nodes connect to Amazon Elastic Container Registry (Amazon ECR) to pull images and to CloudWatch for metrics and logs, two flows that grow with the fleet. To keep bandwidth tied to real change rather than node count, the team uses an Amazon ECR pull-through cache for frequently pulled images and aggregates logs before shipping them to the Region.

Running compute next to the data avoids a round-trip to a Region on every transaction. It also supports incremental modernization: the team can move one service onto Kubernetes at a time, without first migrating data or refactoring applications.

Results

After migrating the Clinicals Macro Service from AWS Outposts to EKS Hybrid Nodes, athenahealth measured the following gains:

  • Response times dropped by roughly half. Measured at the HTTP layer, response times fell by approximately 50 percent. Calls to Redis and Amazon DynamoDB (the busiest backend dependencies) improved by up to 70 percent.
  • Production scale held throughout the migration. Request volume peaks near 40 million requests per daily business processing cycle. Errors stayed near zero, and lower response-time quantiles decreased and held after cutover.
  • CPU utilization stayed below 10 percent across the eight-node fleet, even at daily peaks. This headroom gives the team room to consolidate workloads and plan for growth without adding hardware.
  • Pod count halved for the same workload. Average available replicas fell from approximately 48 to 24 while unavailable replicas stayed near zero, and peaks still reached the low 50s when traffic required it.

Meeting the security audit

Compliance shaped the design as much as performance did. The architecture enforces data classification tiers aligned to healthcare requirements: development environments hold no protected health information (PHI), pre-production environments use masked patient data, and only production clusters handle live PHI. Enforcing these tiers at the cluster boundary prevents a workload in a lower environment from reaching live data. Developers iterate quickly where data is not sensitive, and production stays HIPAA-eligible.

This tier model mattered for the HITRUST security audit that drove the migration timeline. EKS Hybrid Nodes gave athenahealth the ability to place worker nodes in the correct network security zones rather than accepting a fixed topology. The same Puppet-based configuration that provisions the fleet produces an identical, testable artifact in every zone, giving auditors a repeatable demonstration of consistent controls. athenahealth passed the audit and retained its HITRUST certification.

Across both environments, athenahealth applies uniform security controls. IAM handles authentication while Kubernetes role-based access control (RBAC) handles authorization, with attribute-based access control (ABAC) through resource tags for fine-grained permissions. Secrets reside in AWS Secrets Manager or Systems Manager Parameter Store and reach pods through the Secrets Store CSI Driver, which mounts them as volumes rather than embedding them in manifests. Images are scanned for software vulnerabilities before deployment, and Pod Security Standards prevent non-compliant containers from starting.

Conclusion

By adopting EKS Hybrid Nodes, athenahealth unified its Kubernetes operating model across the cloud and the data center. Response times dropped by half, CPU headroom expanded, pod density improved, and projected hardware costs fell by 50 percent, all while meeting a HITRUST security audit and keeping regulated healthcare data on premises.

If your workloads share similar constraints (latency sensitivity that demands on-premises proximity, data residency requirements that prevent full cloud migration, or a need to modernize incrementally without refactoring), EKS Hybrid Nodes offers a path forward. You keep ownership of the hardware and network, gain a fully managed Kubernetes control plane, and operate a single cluster that spans both environments.

To get started, see the EKS Hybrid Nodes documentation and the EKS Workshop hybrid nodes module for a hands-on walkthrough.


About the authors

Sai Charan Teja Gopaluni

Sai Charan Teja Gopaluni

Sai is a Senior Specialist SA at Amazon Web Services (AWS), focusing on Agentic AI workloads and Container Services. He brings deep expertise in Kubernetes, AWS container services, GPU-accelerated computing, and scalable model inference.

Mallory Quaintance

Mallory Quaintance

Mallory is a Principal Site Reliability Engineer at athenahealth, focused on Kubernetes platform reliability, Amazon EKS, and hybrid infrastructure. She works on production Kubernetes operations, GitOps-based platform delivery, Cilium networking, and EKS Hybrid Nodes architecture supporting athenahealth’s modernization across AWS and on-premises environments.

Hassan Mousaid

Hassan Mousaid

Hassan Mousaid, PhD is a Principal Solutions Architect at Amazon Web Services supporting Healthcare and Life Sciences (HCLS) customers to accelerate the process of bringing ideas to market using Amazon’s mechanisms for innovation.

Daniel Leich

Daniel Leich

Daniel is Architect, Site Reliability Engineer at athenahealth, leading design and architecture for the athenaOne Reliability Engineering organization. He is an experienced operations leader who guides teams from inception to execution, with skills focused on container orchestration, configuration management, and AI enablement.