NOTE

Network Tuning

1. Network Performance Metrics 1.1. Network Layer It is mainly responsible for network packet encapsulation, addressing, routing, sending, and receiving. Performance metric: Number of network packets that can be processed per second, PPS. The kernel'

Operating Systems / LinuxCreated Updated 3 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. Network Performance Metrics

1.1. Network Layer

  • It is mainly responsible for network-packet encapsulation, addressing, routing, sending, and receiving.
  • Performance metric:
    • Number of network packets that can be processed per second, PPS.
  • The kernel’s built-in packet-generation tool pktgen can be used for testing.

1.2. Transport Layer

  • It is mainly responsible for network transmission.
  • Performance metrics:
    • Throughput (BPS)
    • Number of connections and latency
  • iperf or netperf can be used to test transport-layer performance.

1.3. Application Layer

  • Performance metrics:
    • Throughput (BPS)
    • Requests per second
    • Latency
  • Tools such as wrk and ab can be used to test application performance.

2. Network Performance Tools

2.1. sar

sar.md

2.2. nethogs

nethogs.md

2.3. iftop

iftop.md

2.4. netstat

netstat.md

2.5. ping

2.6. traceroute

traceroute.md

2.7. nslookup

nslookup.md

2.8. dig

dig.md

2.9. tcpdump

tcpdump.md

2.10. wireshark

3. How to Analyze Network Performance Bottlenecks

3.1. Check Whether It Is an I/O Bottleneck

  • Use top to check CPU load (load average) and CPU utilization (%Cpu). If the former is high and the latter is low, then the bottleneck is network I/O or disk I/O.

3.2. Check Whether It Is a Network I/O Bottleneck

Use top to check whether iowait is relatively high. If so, it may be a disk I/O bottleneck; use iostat -xdk 1 10 to check whether %util is relatively high. If so, it is indeed a disk I/O bottleneck. If both are low, then it is a network I/O problem.

3.3. Find the Process with High I/O Usage

  • Use iftop to find the IP:Port consuming the most traffic.
  • Use netstat to find the process PID corresponding to the IP:Port.
  • ps -e -L h o state,cmd | awk '{if($1=="R"||$1=="D"){print $0}}' | sort | uniq -c | sort -k 1nr

3.4. Analyze the Process

  • Latency problem: use ping or traceroute to check whether latency between the two machines is relatively high, and solve it by deploying in the same region and similar methods.
  • Bandwidth problem: use tools such as pktgen to test bandwidth, and tools such as sar -n DEV 1 to check the currently used bandwidth. If it exceeds the limit, then it is a network I/O bottleneck.
    • Data-volume problem: use tcpdump -i interface-name -v -nn tcp port port-number -w test.cap to capture packets on the business server and storage server; use Wireshark to analyze whether packet fragmentation exists. If so, compress according to the application logic. Suppose QPS is 100K and the packet size is 4K, then 3.2 Gbit/s of bandwidth is required, so a 10 Gbit/s link is needed.
  • Duplicate-packet/packet-loss problem: use tcpdump -i interface-name -v -nn tcp port port-number -w test.cap to capture packets on the business server and storage server; use Wireshark to analyze whether duplicate packets / packet loss exist. If so, work with operations to handle it.
  • Connection-count problem: use netstat -antp |awk '/tcp/ {print $6}' |sort|uniq -c to check the number of ESTABLISHED connections. If it is too small, it may be a connection-pool problem.

4. Packet Loss

4.1. Check Whether There Is Packet Loss

dmesg | grep "TCP: drop open request from"
netstat -ant|grep SYN_RECV|wc -l

4.2. Analyze the Cause of Packet Loss

  1. The half-connection queue is full.
sysctl -w net.ipv4.tcp_max_syn_backlog=1024
// syncookie mechanism
sysctl -w net.ipv4.tcp_syncookies=1
  1. The full-connection queue is full.
ss -lnt
Recv-Q: current size of the full-connection queue, that is, TCP connections that have completed the three-way handshake and are waiting for the server to accept();
Send-Q: current maximum full-connection queue length. The output above indicates that the TCP service listening on port 8088 has a maximum full-connection length of 128;
cat /proc/sys/net/ipv4/tcp_abort_on_overflow
  1. Maximum number of connections
cat /proc/sys/net/netfilter/nf_conntrack_max

5. Timeout Case in Practice

I checked the timeout of <SERVICE_A>’s <API_A> API (claiming a gift pack); it was mainly because <SERVICE_B>’s <API_B> API was slow. It may be because it needs to call <DELIVERY_SYSTEM> for delivery on one hand, and the SQL is relatively complex on the other; In addition, it only provides single-item claiming and does not support batch claiming, while my side calls it concurrently. When volume becomes large, it is relatively easy to time out

6. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub