What Is Latency Sensitive?
- What is latency sensitive?
- Why is latency sensitivity important?
- What are latency-sensitive applications?
- What causes latency in systems?
- What are the characteristics of latency-sensitive applications?
- How do you measure latency sensitivity?
- What are some best practices for countering latency-sensitive applications?
- How can AWS support your latency-sensitive requirements?
What is latency sensitive?
Latency sensitive refers to the degree to which an application or workload’s performance is affected by longer-than-usual request and response cycle times. Latency time includes network, processing, and storage times. For example, applications with high latency sensitivity, such as real-time video streams, can have interruptions in service due to network congestion, making the application unusable. Low-latency-sensitive applications, such as file system backup services, might take 10 times as long as usual to complete without having an impact on the outcome. Latency-sensitive applications benefit from latency optimizations, such as applying higher bandwidth networking and caching strategies.
Why is latency sensitivity important?
It is important to determine an application or workload’s latency sensitivity before its general deployment for several reasons.
User experience impact
If an application is latency-sensitive, it can have a significant impact on the user’s experience. An application might have interruptions, experience wait times, operate more slowly, or sometimes exit unexpectedly, all of which can be frustrating for a human user.
Business outcomes
Applications that are affected by latency tend to have a slower time-to-result for businesses. In these cases, latency affects productivity within the business, regardless of whether a human user is interacting with the system.
System architecture requirements
A latency-sensitive application usually requires a specific system architecture to avoid degradation of service. For example, older mobile phones can usually still run new versions of mobile operating system (OS) software, but the user will notice a degradation in responsiveness. It is important to use systems that correspond with an application’s latency requirements.
What are latency-sensitive applications?
Latency-sensitive applications or workloads are those applications that have a negative impact on users or businesses in the case of deviation from normal latency levels. Grading an application’s level of latency sensitivity can help you determine which latency optimizations to deploy.
High latency sensitive applications
Applications with a high degree of latency sensitivity are applications that experience service quality degradation even with very minimal latency changes. These applications will require latency optimizations across all levels of the system.
An example of an application with high latency sensitivity is a video conferencing app, where minimal delays can lead to missed or garbled words or video stream drops, defeating the purpose of the application.
Medium latency sensitive applications
Applications with medium latency sensitivity have a relatively strong resilience to latency within the system. In very high latency time periods, the application might experience service interruptions; otherwise, the effects of latency fluctuations will be minimal.
An example of an application with medium latency sensitivity is an internal messaging app, where small delays in message delivery aren’t impactful on the user experience.
Low latency sensitive applications
Applications that have low latency sensitivity are not affected by system latency, or the effects do not impact user or business outcomes.
An example of an application with low latency sensitivity is a weekend file system backup, where it doesn’t matter whether the backup takes minutes or hours to complete, and the process can resume after an extended interruption, even though the system processes critical data.
What causes latency in systems?
Application or workload latency is the time between a request and a response within a system. Many components comprise the overall latency of a system. Here are three of the most important elements that impact overall latency.
Network latency
Network latency is the time taken to transmit the message over the network. Network latency is impacted by queuing delays from inadequate network bandwidth, network traffic, packet sizes, routing selection, and other network communication and network infrastructure elements.
Processing latency
Processing latency is the time it takes for an application to process a request and output a result. Data processing latency can be affected by the programming language of the application, its underlying hardware, or the way the code is written.
Storage latency
Storage latency refers to the time taken to shuttle data to and from storage within the application. Storage latency is impacted by the types of storage in use within the application, such as physical or virtual RAM and SSD, and the communications methods between the processor and storage.
What are the characteristics of latency-sensitive applications?
Several characteristics tend to apply to latency-sensitive applications:
-
High volume of system requests
-
Real-time performance requirements
-
Stream or event-driven processing
-
High bandwidth consumption
-
Stateful systems
While not all latency-sensitive systems share all these characteristics, at least one or two will apply.
How do you measure latency sensitivity?
Measuring latency is a time-based activity, where you measure latency distribution using percentiles. For example, you would typically measure the full request and response cycle of the system at the 50th, 95th, and 99th percentiles. The 50th percentile represents the average latency of the system overall, to any user or system.
However, latency sensitivity must be considered mainly from an impact perspective: on the user or on the business. This means that to measure the latency sensitivity of an application, to grade it, and decide which latency optimizations to apply, you need to measure both user impact and business impact. For instance, you could take a user survey of the system operating at the 50th, 95th, and 99th percentiles to determine satisfaction. Similarly, you could run the system at the 50th, 95th, and 99th percentiles to assess the impact on business systems throughput.
What are some best practices for countering latency-sensitive applications?
There are many different methods for reducing latency across a system. Here are some of the most common techniques for reducing latency.
Edge computing
Edge computing moves applications to locations that are closer to end users to decrease network-related latencies.
Caching strategies
Caching strategies for frequently accessed data, such as storing data locally, at the edge, or in the preprocessing stage, help to reduce processing latencies.
Connection optimization
Network engineering techniques to gain more bandwidth or speed from the network, such as dedicated connections and higher-speed internet help to reduce network latencies.
Geographic distribution

Deploying applications in multiple geographic locations closest to users, rather than a central location, such as using a content delivery network, helps reduce network latency.
Load balancing
When an application can distribute processing across multiple machines or CPUs, this helps to reduce the time spent in processing.
How can AWS support your latency-sensitive requirements?
AWS offers a range of instance types, networking services, and delivery architectures to help make sure that your high-latency sensitive applications run smoothly in any region:
-
Amazon CloudFront is a content delivery network that helps you securely deliver content with low latency and high transfer speeds. CloudFront reduces latency by delivering data through 700+ globally dispersed Points of Presence (PoPs) with automated network mapping and intelligent routing.
-
Amazon EC2 offers latency-optimized instances to meet demanding business needs across all regions.
-
Amazon ElastiCache is a serverless, fully managed caching service delivering microsecond latency with Valkey-, Memcached-, and Redis OSS-compatibility.
-
AWS Global Accelerator allows you to onboard your user traffic at one of the Global Accelerator edge locations. AWS Global Accelerator helps to improve network performance for your applications by up to 60%.
-
AWS Local Zones allow you to run applications on AWS infrastructure closer to your end users and workloads, while meeting data residency requirements for regulatory and compliance-sensitive workloads.
Get started with latency-sensitive workloads on AWS by creating a free account today.
Browse all cloud computing concepts
Browse all cloud computing concepts content here:
Did you find what you were looking for today?
Let us know so we can improve the quality of the content on our pages