NOTE

1.1 Time-Series Database

Time-series databases, time-series data, data models, and storage implementation.

ObservabilityCreated Updated 2 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is a Time-Series Database

A database optimized specifically for timestamp or time-series data. Compared with many traditional business tables that only keep the current state, a time-series database keeps historical data that changes over time. Time-series queries also often use time as a filter condition.

2. What Is Time-Series Data

Time-series data is a series of data based on time.

3. Why a Time-Series Database Is Needed

3.1. Recording with a Relational Model

  • If a database is used to record time-series data, it looks like this:
    • Disadvantages:
      • Writes: large amount of data,
      • Queries: low efficiency when aggregating by time
      • Cost: high cost

3.2. Recording with a Time-Series Model

  • metric: the metric name and identifier of the current data. It is equivalent to a table in a relational database.

  • data point: a data point, equivalent to a row in a relational database.

  • timestamp: the timestamp representing when the data point was generated.

  • field: different fields under a metric. In general, these store attribute information that changes as the timestamp changes. For example, a location metric can have longitude and latitude as two fields.

  • tag: a label or additional information. In general, these store attribute information that does not change as the timestamp changes. A timestamp plus all tags can be regarded as the table’s primary key.

  • In the example above:

    • The metric is Wind, and every data point has a timestamp.
    • There are two fields: direction and speed.
    • There are two tags: sensor and city. The first and third rows both store the device whose sensor number is 95D8-7913, and whose city attribute is Shanghai. As time changes, both wind direction and wind speed change: the direction changes from 23.4 to 23.2, while the speed changes from 3.4 to 3.3.

4. Time-Series Database Implementation

4.1. LSM Tree

LSM.md

4.2. Distributed Storage

Time-series databases face massive data writing, storage, and reading workloads that a single machine cannot handle, so multiple machines are needed for storage, that is, distributed storage.

4.2.1. Distribution Algorithm

Distributed system partitioning

4.2.2. Shard Key

Shard by metric + tags. Because queries are often performed over a time range, data with the same metric and tags can be assigned to one machine and stored continuously. Sequential disk reads are fast.

5. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub