1. 1.1 Distributed Systemshistorical

    1. What Is a Distributed System - A system composed of multiple subsystems. The subsystems communicate through a network, and each subsystem consists of multiple machines. 2. Why Distributed Systems Are Needed - A single machine has limited read/write capability and is not safe. 3. How to Design a Distributed System 3.1. Replication. 3.2. Partitioning. 4. Theoretical Foundations of Distributed Systems - CAP, BASE.

  2. 1.2 How to Implement Distributed Lockshistorical

    1. What Is a Distributed Lock - A lock in a distributed environment (across processes or machines). It satisfies the following conditions: atomicity, mutual exclusion, no deadlock, and locking and unlocking must be performed by the same client. 2. Why Distributed Locks Are Needed - Built-in locks in various languages, such as Java synchronized and Go mutex, can only guarantee lock properties within a single process and cannot span processes or machines.

  3. 1.3 How to Implement Distributed IDshistorical

    1. What Is a Distributed ID - A unique ID in a distributed environment (multiple machines). Characteristics of distributed IDs: globally unique; roughly increasing; high concurrency; high availability. 2. How to Implement Distributed IDs - Database auto-increment, UUID, Redis-generated IDs, Snowflake IDs.

  4. 1.4 How to Implement Distributed Sessionshistorical

    1. What Is a Distributed Session - A session is server-side memory. - A session shared in a distributed environment (multiple machines). 2. Why Distributed Sessions Are Needed - In a distributed environment, if the traditional Tomcat session mechanism is still used, the following can happen: after a user logs in to system A, the user needs to jump to system B to perform some operations, but the user's login information is stored in system A and system B has no related session.

  5. 1.5 How to Implement Distributed Storagehistorical

    1. What Is Distributed Storage - File storage in a distributed environment (across multiple machines). 2. Why Distributed Storage Is Needed - A single machine has limited performance and capacity. - A single machine does not have high availability. 3. How to Implement Distributed Storage - Distributed-System Replication - Distributed-System Partitioning - Distributed Consistency - Distributed-System Cluster Metadata Management.

  6. 1.6 BASEhistorical

    1. What Is BASE - An extension of AP theory. - It guarantees eventual consistency rather than strong consistency. That is, because failures are unavoidable, I allow data to be different during this period, but after this period the data needs to be consistent. - Availability is obtained by sacrificing strong consistency. When a failure occurs, partial unavailability is allowed, but core functions must remain available.

  7. 1.7 CAPhistorical

    1. Why CAP Exists - A distributed system has multiple nodes, and state needs to be synchronized among the nodes. This requires support from CAP theory. 2. What Is CAP - Only two of the three can be chosen. 2.1. C (Consistency) - Consistency. - A read after a write must return that value. When data is distributed across multiple nodes, the data read from any node must be that value.

  8. 1.8 Distributed-System Cluster Metadata Managementhistorical

    1. What Is Cluster Metadata - Data in any file system can be divided into actual data and metadata. - Data refers to the data we store in files. - Metadata refers to characteristics of files, such as access permissions and data-block distribution. 2. Why Cluster Metadata Is Needed - Only through cluster metadata can we know the mapping relationship between data and partitions.

  9. 1.9 Distributed Consistencyhistorical

    1. What Is Distributed Consistency - Data remains consistent across multiple replicas, that is, data consistency. 2. Why Distributed Consistency Is Needed - In a distributed environment, the same data needs to be replicated to multiple nodes for fault tolerance. Because replication is delayed by network problems, the same data may differ across multiple nodes at the same moment. Distributed consistency exists to solve this problem.

  10. 1.10 Distributed Computinghistorical

    1. What Is Distributed Computing 2. Why Distributed Computing Is Needed 3. Categories of Distributed Computing 4. Batch Processing 5. Stream Processing - Stream processing - Stream: data that gradually increases over time - An event is the smallest unit of stream processing - Each event contains a timestamp indicating its creation time - Events are produced by producers and correspond to multiple consumers.