Research and Implementation of Distributed Reliable Learning Technologies for Unmanned Cluster Systems
摘要
This study focuses on developing robust distributed learning methodologies for unmanned cluster systems that must operate under severe resource limitations and rapidly changing conditions. We tackle fundamental issues related to system fault tolerance, inter-device knowledge sharing, and coordinated computation among diverse unmanned platforms using a novel distributed anti-destructive learning approach. When tested across multiple machine learning architectures—including CNN, MobileNet, EfficientNet, YOLO, and FastRCNN—our system maintains performance losses below 5% even during catastrophic scenarios where more than 30% of network nodes fail simultaneously. Through comprehensive benchmarking against state-of-the-art fault-tolerant methods including FedAvg and Byzantine-resilient approaches, our framework demonstrates superior resilience. The core innovation lies in combining hybrid synchronous-asynchronous parallel processing with multi-center dynamic networking and a veteran-rookie knowledge transfer mechanism specifically designed for unmanned system limitations. We deployed and evaluated our approach using heterogeneous hardware configurations, specifically JetRover autonomous vehicles and RosMaster X3 Plus robotic units, connected through a managed 100 Mbps network infrastructure. The implementation integrates intelligent device clustering, sophisticated knowledge sharing protocols based on distance correlation metrics, and self-healing recovery systems utilizing BFT-SMaRt consensus protocol.