A journal of IEEE and CAA , publishes high-quality papers in English on original theoretical/experimental research and development in all areas of automation
Volume 13 Issue 8
Aug.  2026

IEEE/CAA Journal of Automatica Sinica

  • JCR Impact Factor: 18.3, Top 1 (SCI Q1)
    CiteScore: 28.2, Top 1% (Q1)
    Google Scholar h5-index: 95, TOP 5
Turn off MathJax
Article Contents
H. Chen, T. Xu, X. Wu, and J. Kittler, “A knowledge-imparting generative modelling framework for heterogeneous federated learning,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 8, pp. 1938–1951, Aug. 2026. doi: 10.1109/JAS.2026.125840
Citation: H. Chen, T. Xu, X. Wu, and J. Kittler, “A knowledge-imparting generative modelling framework for heterogeneous federated learning,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 8, pp. 1938–1951, Aug. 2026. doi: 10.1109/JAS.2026.125840

A Knowledge-Imparting Generative Modelling Framework for Heterogeneous Federated Learning

doi: 10.1109/JAS.2026.125840
Funds:  This work was supported in part by the National Natural Science Foundation of China (62576152, 62332008, 62336004), the Basic Research Program of Jiangsu (BK20250104), the Fundamental Research Funds for the Central Universities (JUSRP202504007), and the Leverhulme Trust Emeritus Fellowship (EM-2025-06-09)
More Information
  • Federated learning aims to provide security for client data privacy in practical machine learning applications. In principle, a global server aggregates the models produced by local clients to obtain a global model. However, the server is challenged when collaborating with local clients handling non-identically distributed data without authorisation to access it. Therefore, advanced solutions advocate the use of generative modules to deliver surrogate data to local clients during a server-agent interaction, without revealing private particulars. We argue that such a unidirectional transfer of surrogate patterns cannot fully represent and harmonise knowledge during the server-client interactions. To this end, we propose a knowledge-imparting generative modelling framework (FedKIG) based on adversarial feature learning and bidirectional knowledge distillation, to explore the potential of interactive generative modelling. In particular, FedKIG trains a feature discriminator for each local client to identify the surrogate patterns extracted by the global model. Under the supervision of the local feature discriminators, the server learns a global generator to generate pseudo samples that convey its global perspective. In this manner, local models are enabled to absorb global knowledge, thereby mitigating the training data divergence caused by data heterogeneity. In addition, we develop a bidirectional knowledge distillation strategy to support the entire learning process. This strategy breaks the rigidity of federated distillation by updating knowledge transfer between the server and the clients iteratively, thus overcoming the learning-forgetting issue. The proposed privacy-protected server-client interaction solution supports explicit knowledge generation for exploitation in federated learning. Extensive experimental results indicate that FedKIG significantly improves the generalisation performance and the stability of the model in heterogeneous federated learning scenarios.

     

  • loading
  • [1]
    B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Proc. 20th Int. Conf. Artificial Intelligence and Statistics, Fort Lauderdale, USA, 2017, pp. 1273−1282.
    [2]
    M. Wei, W. Yu, D. Chen, M. Kang, and G. Cheng, “Privacy distributed constrained optimization over time-varying unbalanced networks and its application in federated learning,” IEEE/CAA J. Autom. Sinica, vol. 12, no. 2, pp. 335–346, Feb. 2025. doi: 10.1109/JAS.2024.124869
    [3]
    H. Yuan, M. Zhou, Q. Liu, and A. Abusorrah, “Fine-grained resource provisioning and task scheduling for heterogeneous applications in distributed green clouds,” IEEE/CAA J. Autom. Sinica, vol. 7, no. 5, pp. 1380–1393, Sep. 2020. doi: 10.1109/jas.2020.1003177
    [4]
    M. J. Sheller, G. A. Reina, B. Edwards, J. Martin, and S. Bakas, “Multi-institutional deep learning modeling without sharing patient data: A feasibility study on brain tumor segmentation,” in Proc. 4th Int. Workshop, Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries, Granada, Spain, 2019, pp. 92−104.
    [5]
    W. Li, F. Milletarì, D. Xu, N. Rieke, J. Hancox, W. Zhu, et al, “Privacy-preserving federated brain tumour segmentation,” in Proc. 10th Int. Workshop Machine Learning in Medical Imaging, Shenzhen, China, 2019, pp. 133−141.
    [6]
    L. Zong, Q. Xie, J. Zhou, P. Wu, X. Zhang, and B. Xu, “FedCMR: Federated cross-modal retrieval,” in Proc. 44th Int. ACM SIGIR Conf. Research and Development in Information Retrieval, Canada, 2021, pp. 1672−1676.
    [7]
    F. Pinelli, G. Tolomei, and G. Trappolini, “FLIRT: Federated learning for information retrieval,” in Proc. 46th Int. ACM SIGIR Conf. Research and Development in Information Retrieval, Taipei, China, 2023, pp. 3472−3475.
    [8]
    W. Yuan, Q. V. H. Nguyen, T. He, L. Chen, and H. Yin, “Manipulating federated recommender systems: Poisoning with synthetic users and its countermeasures,” in Proc. 46th Int. ACM SIGIR Conf. Research and Development in Information Retrieval, Taipei, China, 2023, pp. 1690−1699.
    [9]
    C. Chen, X. Feng, J. Zhou, J. Yin, and X. Zheng, “Federated large language model: A position paper,” arXiv preprint arXiv: 2307.08925, 2023.
    [10]
    H. Wu, X. Xu, D. Zhang, X. Li, J. Wu, and Z. Liu, “CG-FedLLM: How to compress gradients in federated fune-tuning for large language models,” arXiv preprint arXiv: 2405.13746, 2024.
    [11]
    R. Ye, W. Wang, J. Chai, D. Li, Z. Li, Y. Xu, Y. Du, Y. Wang, and S. Chen, “OpenFedLLM: Training large language models on decentralized private data via federated learning,” in Proc. 30th ACM SIGKDD Conf. Knowledge Discovery and Data Mining, Barcelona, Spain, 2024.
    [12]
    M. V. Luzón, N. Rodríguez-Barroso, A. Argente-Garrido, D. Jiménez-López, J. M. Moyano, J. Del Ser, W. Ding, and F. Herrera, “A tutorial on federated learning from theory to practice: Foundations, software frameworks, exemplary use cases, and selected trends,” IEEE/CAA J. Autom. Sinica, vol. 11, no. 4, pp. 824–850, Apr. 2024. doi: 10.1109/JAS.2024.124215
    [13]
    Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated machine learning: Concept and applications,” ACM Trans. Intell. Syst. Technol. (TIST), vol. 10, no. 2, Art. no. 12, Mar. 2019.
    [14]
    H. Chen, T. Xu, X. Wu, and J. Kittler, “Hybrid batch normalisation: Resolving the dilemma of batch normalisation in federated learning,” in Proc. 42nd Int. Conf. Machine Learning, Vancouver, Canada, 2025.
    [15]
    T.-M. H. Hsu, H. Qi, and M. Brown, “Measuring the effects of non-identical data distribution for federated visual classification,” arXiv preprint arXiv: 1909.06335, 2019.
    [16]
    S. P. Karimireddy, S. Kale, M. Mohri, S. J. Reddi, S. U. Stich, and A. T. Suresh, “SCAFFOLD: Stochastic controlled averaging for federated learning,” in Proc. 37th Int. Conf. Machine Learning, 2020, pp. 5132−5143.
    [17]
    T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” in Proc. 3rd Conf. Machine Learning and Systems, Austin, USA, 2020, pp. 429−450.
    [18]
    D. A. E. Acar, Y. Zhao, R. M. Navarro, M. Mattina, P. N. Whatmough, and V. Saligrama, “Federated learning based on dynamic regularization,” in Proc. 9th Int. Conf. Learning Representations, Austria, 2021.
    [19]
    R. Ye, Y. Du, Z. Ni, Y. Wang, and S. Chen, “Fake it till make it: Federated learning with consensus-oriented generation,” in Proc. 12th Int. Conf. Learning Representations, Vienna, Austria, 2024.
    [20]
    T. Lin, L. Kong, S. U. Stich, and M. Jaggi, “Ensemble distillation for robust model fusion in federated learning,” in Proc. 34th Int. Conf. Neural Information Processing Systems, Vancouver, Canada, 2020. Art. no. 198.
    [21]
    Z. Wu, S. Sun, Y. Wang, M. Liu, X. Jiang, and R. Li, “Survey of knowledge distillation in federated edge learning,” arXiv preprint arXiv: 2301.05849v1, 2023.
    [22]
    D. Li and J. Wang, “FedMD: Heterogenous federated learning via model distillation,” arXiv preprint arXiv: 1910.03581, 2019.
    [23]
    J. Shao, F. Wu, and J. Zhang, “Selective knowledge sharing for privacy-preserving federated distillation without a good teacher,” Nature Communications, vol. 15, no. 1, Art. no. 349, 2024.
    [24]
    H. Q. Le, M. N. H. Nguyen, S. R. Pandey, C. Zhang, and C. S. Hong, “CDKT-FL: Cross-device knowledge transfer using proxy dataset in federated learning,” Eng. Appl. Artif. Intell., vol. 133, Art. no. 108093, Jul. 2024. doi: 10.1016/j.engappai.2024.108093
    [25]
    Z. Zhu, J. Hong, and J. Zhou, “Data-free knowledge distillation for heterogeneous federated learning,” in Proc. 38th Int. Conf. Machine Learning, 2021, pp. 12878−12889.
    [26]
    L. Zhang, L. Shen, L. Ding, D. Tao, and L.-Y. Duan, “Fine-tuning global model via data-free knowledge distillation for non-IID federated learning,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, New Orleans, USA, 2022, pp. 10174−10183.
    [27]
    H. Wang, Y. Li, W. Xu, R. Li, Y. Zhan, and Z. Zeng, “DaFKD: Domain-aware federated knowledge distillation,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Vancouver, Canada, 2023, pp. 20412−20421.
    [28]
    J. Wang, Q. Liu, H. Liang, G. Joshi, and H. V. Poor, “Tackling the objective inconsistency problem in heterogeneous federated optimization,” in Proc. 34th Int. Conf. Neural Information Processing Systems, Vancouver, Canada, 2020, Art. no. 638.
    [29]
    Q. Li, B. He, and D. Song, “Model-contrastive federated learning,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, Nashville, USA, 2021, pp. 10713−10722.
    [30]
    G. Lee, M. Jeong, Y. Shin, S. Bae, and S.-Y. Yun, “Preservation of the global knowledge by not-true distillation in federated learning,” in Proc. 36th Int. Conf. Neural Information Processing Systems, New Orleans, USA, 2022, Art. no. 2787.
    [31]
    F. Sattler, T. Korjakow, R. Rischke, and W. Samek, “FedAUX: Leveraging unlabeled auxiliary data in federated learning,” IEEE Trans. Neural Netw. Learn. Syst., vol. 34, no. 9, pp. 5531–5543, Sep. 2023. doi: 10.1109/TNNLS.2021.3129371
    [32]
    Z. Yang, Y. Zhang, Y. Zheng, X. Tian, H. Peng, T. Liu, and B. Han, “FedFed: Feature distillation against data heterogeneity in federated learning,” in Proc. 37th Int. Conf. Neural Information Processing Systems, New Orleans, USA, 2024, Art. no. 2639.
    [33]
    J. Lü, G. Wen, R. Lu, Y. Wang, and S. Zhang, “Networked knowledge and complex networks: An engineering view,” IEEE/CAA J. Autom. Sinica, vol. 9, no. 8, pp. 1366–1383, Aug. 2022. doi: 10.1109/JAS.2022.105737
    [34]
    G. Patel, K. R. Mopuri, and Q. Qiu, “Learning to retain while acquiring: Combating distribution-shift in adversarial data-free knowledge distillation,” in Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023, pp. 7786−7794.
    [35]
    I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Proc. 28th Int. Conf. Neural Information Processing Systems, Montreal, Canada, 2014, pp. 2672−2680.
    [36]
    L. Mescheder, S. Nowozin, and A. Geiger, “The numerics of GANs,” in Proc. 31st Int. Conf. Neural Information Processing Systems, Long Beach, USA, 2017, pp. 1823−1833
    [37]
    C. Wu, F. Wu, L. Lyu, Y. Huang, and X. Xie, “Communication-efficient federated learning via knowledge distillation,” Nat. Commun., vol. 13, no. 1, Art. no. 2032, Apr. 2022. doi: 10.1038/s41467-022-29763-x
    [38]
    X. Li, K. Huang, W. Yang, S. Wang, and Z. Zhang, “On the convergence of FedAvg on non-IID data,” in Proc. 8th Int. Conf. Learning Representations, Addis Ababa, Ethiopia, 2020.
    [39]
    G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv: 1503.02531, 2015.
    [40]
    H. Chen, Y. Wang, C. Xu, Z. Yang, C. Liu, B. Shi, C. Xu, C. Xu, and Q. Tian, “Data-free learning of student networks,” in Proc. IEEE/CVF Int. Conf. Computer Vision, Seoul, Korea (South), 2019, pp. 3514−3522.
    [41]
    G. Fang, J. Song, X. Wang, C. Shen, X. Wang, and M. Song, “Contrastive model inversion for data-free knowledge distillation,” arXiv preprint arXiv: 2105.08584, 2021.
    [42]
    J. Oh, S. Kim, and S.-Y. Yun, “FedBABU: Toward enhanced representation for federated image classification,” in Proc. 10th Int. Conf. Learning Representations, 2022.
    [43]
    H. B. McMahan, E. Moore, D. Ramage, and B. A. y Arcas, “Federated learning of deep networks using model averaging,” arXiv preprint arXiv: 1602.05629, 2016.
    [44]
    M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in Proc. 34th Int. Conf. Machine Learning, Sydney, Australia, 2017, pp. 214−223.
    [45]
    S. Jadon, “A survey of loss functions for semantic segmentation,” in Proc. IEEE Conf. Computational Intelligence in Bioinformatics and Computational Biology, Viña del Mar, Chile, 2020, pp. 1−7.
    [46]
    M. Mirza and S. Osindero, “Conditional generative adversarial nets,” arXiv preprint arXiv: 1411.1784, 2014.
    [47]
    Y. Kossale, M. Airaj, and A. Darouichi, “Mode collapse in generative adversarial networks: An overview,” in Proc. 8th Int. Conf. Optimization and Applications, Genoa, Italy, 2022, pp. 1−6.
    [48]
    H. Xiao, K. Rasul, and R. Vollgraf, “Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms,” arXiv preprint arXiv: 1708.07747, 2017.
    [49]
    Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” in Proc. NIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011.
    [50]
    G. Cohen, S. Afshar, J. Tapson, and A. van Schaik, “EMNIST: Extending MNIST to handwritten letters,” in Proc. Int. Joint Conf. Neural Networks, Anchorage, USA, 2017, pp. 2921−2926.
    [51]
    A. Krizhevsky, “Learning multiple layers of features from tiny images,” M.S. thesis, University of Toronto, Toronto, Canada, 2009.
    [52]
    X. Ma, J. Zhu, Z. Lin, S. Chen, and Y. Qin, “A state-of-the-art survey on solving non-IID data in federated learning,” Future Generation Computer Systems, vol. 135, pp. 244–258, 2022.
    [53]
    Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proc. IEEE, vol. 86, no. 11, pp. 2278–2324, Nov. 1998. doi: 10.1109/5.726791
    [54]
    X. Li, Z. Song, and J. Yang, “Federated adversarial learning: A framework with convergence analysis,” in Proc. 40th Int. Conf. Machine Learning, Honolulu, USA, 2023, Art. no. 823.
    [55]
    D. Yao, W. Pan, Y. Dai, Y. Wan, X. Ding, H. Jin, Z. Xu, and L. Sun, “Local-global knowledge distillation in heterogeneous federated learning with non-IID data,” arXiv preprint arXiv: 2107.00051, 2021.
    [56]
    Y. Yu, W. Zhang, and Y. Deng, “Frechet inception distance (fid) for evaluating GANs,” China University of Mining Technology Beijing Graduate School, 2021.

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(8)  / Tables(10)

    Article Metrics

    Article views (12) PDF downloads(0) Cited by()

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return