Skip to main content

What Is a Time Series Database?

What is a time series database?

A time series database (TSDB) is a type of database management system that is structured for storing and analyzing time-ordered data. TSDBs are capable of handling large amounts of continuous event stream data, such as Internet of Things (IoT) sensor data from a fleet of devices. TSDBs are designed to support fast ingestion, compression, and optimized time-based querying.

Why are time series databases important?

Time series databases deliver real-time insights and applications that output up to sub-millisecond updates. Most IoT devices, application monitoring systems, and live streaming applications all create and deliver real-time data. Time-stamped data can be critical in areas such as financial trading and data observability.

Traditional databases are not optimized for large-scale write throughput and querying based on data’s timestamps. TSDBs are optimized for high-volume writes and to query data within specific time windows instead of throughout an entire dataset. This time-based, narrowed approach to querying is designed to provide better performance.

Due to their specific purposes, TSDBs can also implement additional features to improve efficiency. As they mainly store repeats or values with only slight changes, TSDBs can use compression techniques to reduce the total size of a dataset without reducing the total amount of information stored. For example, a timestamp may be stored as the small-value delta between two times instead of the lengthy timestamp itself. Especially for long-term real-time database tracking, this helps save on storage costs and reduces the amount of raw data you hold.

Additionally, you can also create data retention policies that delete or archive data based on its age, reducing clutter and keeping your database as small as possible. A TSDB solves the common issues in storing time series data in a regular database.

How does a time series database work?

Time series databases use an append-only write model where new information is added, but existing records are only rarely edited. By simplifying the write process with this strategy, TSDBs can ingest lots of data quickly and efficiently.

Here is how the process works.

Data ingestion

TSDB data infrastructure is optimized for high-volume ingestion, helping to collect lots of data points without performance issues. You typically write data in batches or streams, with the system being able to handle massive volumes of data points per second. Ingestion pipelines will support batching and buffering to create a scalable, high-throughput system.

Indexing and partitioning by time

A timestamp is the primary index in a TSDB. Historical data is partitioned into time-based segments, such as blocks by hour or by day. When you want to query specific information for data analysis, you can outline the time period you want information about, with the database only searching through the relevant blocks.

Data compression techniques

TSDBs use specialized compression algorithms to store large volumes of incrementally changing data in an efficient manner. One of these is delta encoding, where only the difference between two sequential values is stored. Other strategies, such as the Gorilla XOR-based compression scheme, store floating-point values without losing any precision.

Data retention and downsampling policies

Retention policies outline how long your business will store certain types of data. You can move older data to store or delete it entirely once it reaches a certain age.

Alternatively, you can downsample data by aggregating raw information into a summary. Downsampling allows you to store insights on long-term trends without retaining the exact values forever.

What are the key characteristics of a time series database?

A time series database has a range of core characteristics that set it apart from traditional databases.

High write throughput

TSDBs are specifically engineered for continuous data ingestion, using their append-only model and batching techniques to handle enormous volumes of incoming data.

Time-based indexing for enhanced query performance

Time is the primary indexing system and how you query data within a TSDB. By storing data in time blocks, you can quickly retrieve data from a specific date or time period.

Built-in aggregation functions for time window analysis

A time series database offers features that let you aggregate data across specific time windows. For example, you could aggregate the average, minimum, and maximum values, percentiles, and rates of change over a specific period. Downsampling aggregates is a specific technique in TSDBs.

Visualizing time series data across seasonality, trend, and residual with Amazon SageMaker Data Wrangler

Automatic data expiry and tiered storage

You can configure time series databases to automatically tier or delete data to efficiently store data. Set a retention period that will archive data once it reaches an expiration age to ensure operational efficiency, and delete it at a later age.

Columnar or chunk-based storage layout

To best store time series data, these databases will use columnar or chunk-based storage formats, or a combination of the two. Both of these approaches work well when indexing time-series data, help improve compression, and enhance query efficiency when searching for a specific time period.

What are the types of time series databases?

Below are the main types of TSDBs you can use.

Purpose-built TSDBs

A purpose-built TSDB is built specifically for temporal data. They are specialized databases that offer optimized data ingestion at scale, compression strategies, and techniques to improve query performance with minimal configuration needed.

InfluxDB and OpenTSDB are examples of purpose-built TSDBs.

Relational databases with time series extensions

You can extend some relational databases with additional capabilities to be able to support time series workloads. Doing this is useful if you want to leverage any existing relational databases you have.

For example, you can extend PostgreSQL with TimescaleDB to add time-series data querying, partitioning, and compression.

Column-oriented databases adapted for time series data analysis

A column-oriented database can also be adapted for time series data. A suitable column-oriented database with partition keys, modified, can handle time-series workloads with high write throughput and efficient storage.

For example, you can add time series workload capabilities to Apache Cassandra.

Cloud-native managed TSDBs

Cloud providers can offer fully managed TSDBs that can scale to meet performance and ingestion demands. You can also integrate them with other cloud-native tools to improve maintenance, data observability, and real-time analytics and visualization.

Amazon Timestream is an example of a cloud-native TSDB.

What are common time series database use cases?

There are a number of use cases for time series databases.

For example, you can use a TSDB to:

  • Capture and process time-stamped data from IoT sensors
  • Monitor application performance in real time
  • Track financial data, such as stock prices and trades within financial markets
  • Track energy consumption and performance with timestamped data across a facility
  • Analyze user behavior over time

Time series database vs. relational database

Time series databases and relational databases both store and query data, but have key differences. Here are a few differences that set them apart.

Write patterns

TSDBs use an append-only model to optimize for the continual ingestion of time-series data. Relational databases instead perform regular updates and deletes, which would introduce too much overhead for high-volume streaming workloads.

Query patterns

Time series databases are specifically built to query over specific time intervals, retrieving data from a certain period or aggregating metrics over a window of time. Relational databases are built for complex queries that may involve joins, multi-dimensional filtering, and transactions.

Schema design

Time series databases often use denormalized schemas with wide tables (multiple columns) to store related data together for efficient reads. Relational databases use normalized schemas to reduce redundancy and enforce strict relationships within their systems.

Compression

TSDBs use a wide range of compression techniques to store incremental data efficiently. Relational databases instead use a range of general-purpose compression strategies, which are less appropriate for time-ordered data.

Data lifecycle management

TSDBs use retention policies to guide how long they should keep their data and what they should do once it expires. Relational databases use similar retention policies but require either manual implementation or a specific archival strategy in place.

How can AWS help with time series databases?

AWS offers a range of managed services for storing and processing time series data at scale across monitoring, IoT, and analytics workloads. Here are two options for TSDBs on AWS:

  • Amazon Keyspaces (for Apache Cassandra) is a serverless workload model for your Cassandra time series databases.
  • Amazon Managed Service for Prometheus allows you to use Prometheus query language (PromQL) to filter, aggregate, ingest, and query millions of unique time series metrics from your self-managed Kubernetes clusters.
  • Amazon Timestream offers fully managed, purpose-built time-series database engines for workloads from low-latency queries to large-scale data ingestion.

Get started with time series databases on AWS by creating a free account today.

Browse all cloud computing concepts

Browse all cloud computing concepts content here:

Loading
Loading
Loading
Loading
Loading

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages