NOTE

1.8 Distributed-System Cluster Metadata Management

1. What Is Cluster Metadata - Data in any file system can be divided into actual data and metadata. - Data refers to the data we store in files. - Metadata refers to characteristics of files, such as access permissions and data-block distribution. 2. Why Cluster Metadata Is Needed - Only through cluster metadata can we know the mapping relationship between data and partitions.

Distributed SystemsCreated Updated 1 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is Cluster Metadata?

  • Data in any file system can be divided into actual data and metadata.
    • Data refers to the data we store in files.
    • Metadata refers to characteristics of files, such as access permissions and data-block distribution.

2. Why Is Cluster Metadata Needed?

  • Only through cluster metadata can we know the mapping relationship between data and partitions.

3. Cluster Metadata Maintenance Methods

There are generally two ways to maintain metadata in a cluster.

3.1. Centralized

  • A separate node is responsible for tracking cluster metadata. This node is called a coordination service. For example, ZooKeeper.

3.2. Distributed (P2P)

  • Each node maintains its own information and exchanges information with the others through ping and pong. For example, Redis Cluster.

3.2.1. Gossip Protocol

4. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub