Session: 12-11: Software-Defined Vehicles, Digital Twins, and Connected AI Platforms
Paper Number: 197354
197354 - From Pilot to Production: Operationalizing Federated Learning for Fleet-Scale Edge Ai
Abstract:
Connected vehicle fleets and industrial transportation assets generate telemetry at a scale that makes centralized model training untenable for bandwidth, cost, and privacy reasons, particularly when data must cross OEM, supplier, or operator boundaries. Federated Learning (FL) is widely cited as a solution, and recent surveys in connected and automated vehicles have mapped its applicability across perception, planning, predictive maintenance, and traffic-flow tasks [1]. Yet most FL efforts in transportation remain in pilot stages and rarely reach production at fleet scale.
While algorithmic progress in FL has accelerated, the primary barriers to fleet-scale deployment are increasingly operational. Production FL across heterogeneous edge fleets requires operational capabilities that off-the-shelf FL frameworks were never designed to provide. We organize these into five requirements any fleet-scale FL platform must satisfy: (1) data security and governance across organizational boundaries, including secure communication and distributed training architectures; (2) orchestration and lifecycle management for large client populations with intermittent connectivity, covering scheduling, rollout, and rollback; (3) observability across training health, client participation, convergence behavior, and communication performance; (4) compute and deployment feasibility on constrained edge hardware under unreliable networks; and (5) algorithmic robustness for non-IID data, partial participation, and personalization.
Open-source FL frameworks have accelerated experimentation but were not designed for the reliability, recoverability, and governance demands of fleet-scale deployment. We contrast the capabilities these frameworks provide against the operational requirements outlined above, and identify the infrastructure and tooling patterns that distinguish a viable production system from a successful pilot.
References
Chellapandi, V. P., Yuan, L., Brinton, C. G., Zak, S. H., & Wang, Z. (2023). Federated Learning for Connected and Automated Vehicles: A Survey of Existing Approaches and Challenges. arXiv. https://doi.org/10.48550/arXiv.2308.10407
Presenting Author: David Solooki Rhino Federated Computing
Presenting Author Biography: David Solooki is a Forward Deployed Engineer for AI/ML at Rhino Federated Computing, where he leads the technical deployment of federated learning systems into production environments. His work bridges privacy-preserving machine learning research and the operational realities of training models across sensitive, distributed data silos - spanning secure compute infrastructure, multi-site data harmonization, and model orchestration with frameworks like NVIDIA FLARE. David partners closely with customers on technical implementations, taking federated learning from proof-of-concept to durable, real-world deployments.
Authors:
David Solooki Rhino Federated ComputingAdrish Sanyassi Rhino Federated Computing
From Pilot to Production: Operationalizing Federated Learning for Fleet-Scale Edge Ai
Paper Type
Technical Presentation Only