Beyond Client Clustering: Fine-Grained Preference Alignment in Federated RLHF via Self-Evolving Routing
Beyond Client Clustering: Fine-Grained Preference Alignment in Federated RLHF via Self-Evolving Routing
Ke Wang, Shaojing Fu, Yuchuan Luo, Guilin Deng, Silong Chen, Zheng Yuan, Lin Liu
Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Main Track. Pages 4958-4966.
https://doi.org/10.24963/ijcai.2026/552
Federated Reinforcement Learning from Human Feedback (RLHF) enables the collaborative alignment of Large Language Models (LLMs) while preserving privacy, yet it faces critical bottlenecks arising from data heterogeneity. Existing approaches typically rely on rigid client-level clustering, which overlooks intra-client heterogeneity and fails to adapt to the multifaceted needs of individual users. To address this, we propose FedPrism, a novel framework that shifts alignment granularity from the coarse client level to the precise instance level. Similar to an optical prism dispersing mixed light into distinct spectral components, FedPrism decomposes complex, intra-client heterogeneous data streams by dynamically routing individual samples to specialized experts based on semantic features. Crucially, we devise a Posterior-Guided Performance Distillation (PGPD) mechanism that leverages experts' actual training loss as self-supervision to autonomously refine routing policies without explicit labels. Extensive experiments show that FedPrism not only establishes new state-of-the-art results but also effectively mitigates the negative transfer prevalent in non-IID settings, ensuring superior alignment fidelity.
Keywords:
Machine Learning: Federated learning
Multidisciplinary Topics and Applications: Security and privacy
