A journal of IEEE and CAA , publishes high-quality papers in English on original theoretical/experimental research and development in all areas of automation
Volume 13 Issue 8
Aug.  2026

IEEE/CAA Journal of Automatica Sinica

  • JCR Impact Factor: 18.3, Top 1 (SCI Q1)
    CiteScore: 28.2, Top 1% (Q1)
    Google Scholar h5-index: 95, TOP 5
Turn off MathJax
Article Contents
T. Chen, Y. Wang, H. Chang, G. Liu, D. Li, L. Rob, X. Xu, and E. Q. Wu, “Instructing the learning of language model with the token interpretation to improve language understanding,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 8, pp. 1915–1925, Aug. 2026. doi: 10.1109/JAS.2026.125900
Citation: T. Chen, Y. Wang, H. Chang, G. Liu, D. Li, L. Rob, X. Xu, and E. Q. Wu, “Instructing the learning of language model with the token interpretation to improve language understanding,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 8, pp. 1915–1925, Aug. 2026. doi: 10.1109/JAS.2026.125900

Instructing the Learning of Language Model With the Token Interpretation to Improve Language Understanding

doi: 10.1109/JAS.2026.125900
Funds:  This work was supported in part by the National Natural Science Foundation of China (62303117, T2325018), the Fundamental and Interdisciplinary Disciplines Breakthrough Plan of the Ministry of Education of China (JYB2025XDXM605), Key Laboratory Independent Research Foundation (C2620409-5), and the Fujian Provincial Natural Science Foundation (2024J01278, 2025J09022)
More Information
  • Pretrained language models (PLMs) have established state-of-the-art performance across diverse natural language understanding (NLU) tasks. This study reveals that semantic-rich explanations of lexical units can effectively guide PLM learning processes. We propose a novel language understanding enhancement method with token interpretation (LUETI) that addresses two critical limitations in conventional PLMs: Incomplete token semantics caused by isolated contextual learning and insufficient semantic encoding in embedding matrices. LUETI operates through dual mechanisms, augmenting token representations by integrating hidden states with corresponding token interpretations and refining embedding spaces using interpretation-derived semantic vectors for token prediction. LUETI, which is implemented as a plug-in module for standard architectures, demonstrates significant improvements on BERT and GLM, achieving average performance gains of 3.36% and 4.87% respectively on the SuperGLUE benchmark with equivalent parameters and training data. Note that LUETI-equipped models attain comparable performance to baseline PLMs using only 60% of pretraining data. Findings establish token interpretation as a computationally efficient but semantically powerful enhancement strategy for language model pretraining.

     

  • loading
  • [1]
    X.-C. Wen, K.-H. Liu, Y. Luo, J. Ye, and L. Chen, “TWACapsNet: A capsule network with two-way attention mechanism for speech emotion recognition,” Soft Comput., vol. 28, no. 15−16, pp. 8701–8713, Aug. 2024. doi: 10.1007/s00500-023-08957-5
    [2]
    M. Potocar and M. Kvet, “Impact of preprocessing using substitution on the performance of selected NER models-methodology,” in Good Practices and New Perspectives in Information Systems and Technologies, Á Rocha, H. Adeli, G. Dzemyda, F. Moreira, and A. Poniszewska-Marańda, eds. Cham, Switzerland: Springer, 2024, pp. 141−150.
    [3]
    M. Namazifar, A. Papangelis, G. Tur, and D. Hakkani-Tür, “Language model is all you need: Natural language understanding as question answering,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing, Toronto, Canada, 2021, pp. 7803−7807.
    [4]
    M. Namazifar, A. Papangelis, G. Tur, and D. Hakkani-Tür, “Zero-shot learners for natural language understanding via a unified multiple-choice perspective,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 31, pp. 1728−1739, 2023.
    [5]
    R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Yu Du, et al., “LaMDA: Language models for dialog applications,” arXiv preprint arXiv: 2201.08239, 2022.
    [6]
    Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “GLM: General language model pretraining with autoregressive blank infilling,” in Proc. 60th Annu. Meeting of the Association for Computational Linguistics, Dublin, Ireland, 2022, pp. 320−335.
    [7]
    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al., “Training language models to follow instructions with human feedback,” in Proc. 36th Int. Conf. Neural Inform. Processing Systems, New Orleans, USA, 2022, Art. no. 2011.
    [8]
    E. Dunbar, N. Hamilakis, and E. Dupoux, “Self-supervised language learning from raw audio: Lessons from the zero resource speech challenge,” IEEE J. Sel. Top. Signal Process., vol. 16, no. 6, pp. 1211–1226, Oct. 2022. doi: 10.1109/JSTSP.2022.3206084
    [9]
    D. Li, B. Zhang, Y. Xiu, H. Deng, M. Zhang, W. Tong, R. Law, G. Zhu, E. Q. Wu, and L. Zhu, “Snake robots play an important role in social services and military needs,” Innovation, vol. 3, no. 6, Art. no. 100333, Nov. 2022. doi: 10.1016/j.xinn.2022.100333
    [10]
    Y. Xiu, D. Li, M. Zhang, H. Deng, R. Law, Y. Huang, E. Q. Wu, and X. Xu, “Finite-time sideslip differentiator-based LOS guidance for robust path following of snake robots,” IEEE/CAA J. Autom. Sinica, vol. 10, no. 1, pp. 239–253, Jan. 2023. doi: 10.1109/JAS.2022.106052
    [11]
    C. Eang and S. Lee, “Improving the accuracy and effectiveness of text classification based on the integration of the BERT model and a recurrent neural network (RNN_Bert_Based),” Appl. Sci., vol. 14, no. 18, Art. no. 8388, Sep. 2024. doi: 10.3390/app14188388
    [12]
    X. Yang, F. Lv, F. Liu, and G. Lin, “Self-training vision language BERTs with a unified conditional model,” IEEE Trans. Circ. Syst. Video Technol., vol. 33, no. 8, pp. 3560–3569, Aug. 2023. doi: 10.1109/TCSVT.2023.3235704
    [13]
    R. K. Singh, M. K. Sachan, and R. B. Patel, “Cross-domain sentiment classification using decoding-enhanced bidirectional encoder representations from transformers with disentangled attention,” Concurr. Comput. Pract. Exp., vol. 35, no. 6, Art. no. 1, Mar. 2023. doi: 10.1002/cpe.7589
    [14]
    B. Liu, “Comparative analysis of encoder-only, decoder-only, and encoder-decoder language models,” in Proc. 1st Int. Conf. Data Science and Engineering, Singapore, 2024, pp. 524−530.
    [15]
    Y. Qian, X. Bian, Y. Shi, N. Kanda, L. Shen, Z. Xiao, and M. Zeng, “Speech-language pre-training for end-to-end spoken language understanding,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing, Toronto, Canada, 2021, pp. 7458−7462.
    [16]
    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res., vol. 21, no. 140, pp. 1–67, Jun. 2020.
    [17]
    H. Bao, L. Dong, F. Wei, W. Wang, N. Yang, X. Liu, Y. Wang, S. Piao, J. Gao, M. Zhou, et al., “UniLMv2: Pseudo-masked language models for unified language model pre-training,” in Proc. 37th Int. Conf. Machine Learning, 2020, pp. 642−652.
    [18]
    R. Zhang, H. Du, Y. Liu, D. Niyato, J. Kang, S. Sun, X. Shen, and H. V. Poor, “Interactive AI with retrieval-augmented generation for next generation networking,” IEEE Network, vol. 38, no. 6, pp. 414–424, Nov. 2024. doi: 10.1109/MNET.2024.3401159
    [19]
    J. Liu, M. Xiao, J. Wen, J. Kang, R. Zhang, T. Zhang, D. Niyato, W. Zhang, and Y. Liu, “Optimizing resource allocation for multi-modal semantic communication in mobile AIGC networks: A diffusion-based game approach,” IEEE Trans. Cognit. Commun. Netw., vol. 11, no. 5, pp. 3346–3360, Oct. 2025. doi: 10.1109/TCCN.2025.3529747
    [20]
    R. Zhang, Z. Xu, and X. Gou, “An integrated method for multi-criteria decision-making based on the best-worst method and Dempster-Shafer evidence theory under double hierarchy hesitant fuzzy linguistic environment,” Appl. Intell., vol. 51, no. 2, pp. 713–735, Feb. 2021. doi: 10.1007/s10489-020-01777-2
    [21]
    D. Li, L. Zeng, Y. Xiu, Z. Pan, D. Zhang, and H. Deng, “Sideslip elimination and coefficient approximation-based trajectory tracking control for snake robots,” IEEE Trans. Ind. Inf., vol. 19, no. 8, pp. 8754–8764, Aug. 2023. doi: 10.1109/TII.2022.3220846
    [22]
    D. Li, B. Zhang, R. Law, E. Q. Wu, and X. Xu, “Error constrained-formation path-following method with disturbance elimination for multisnake robots,” IEEE Trans. Ind. Electron., vol. 71, no. 5, pp. 4987–4998, May 2024. doi: 10.1109/TIE.2023.3288202
    [23]
    Y. Xiu, H. Deng, D. Li, R. Law, E. Q. Wu, and L. Zhu, “Collision avoidance regulation-compliant orientation guidance and maneuvering control approach for snake robots,” IEEE Trans. Ind. Electron., vol. 71, no. 9, pp. 10955–10965, Sep. 2024. doi: 10.1109/TIE.2023.3342304
    [24]
    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. 31st Int. Conf. Neural Inform. Processing Systems, Long Beach, USA, 2017, pp. 6000−6010.
    [25]
    Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, “Aligning books and movies: Towards story-like visual explanations by watching movies and reading books,” in Proc. IEEE Int. Conf. Computer Vision, Santiago, Chile, 2015, pp. 19−27.
    [26]
    A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “SuperGLUE: A stickier benchmark for general-purpose language understanding systems,” in Proc. 33rd Int. Conf. Neural Inform. Processing Systems, Vancouver, Canada, 2019, Art. no. 294.
    [27]
    T. Schick and H. Schütze, “Exploiting cloze-questions for few-shot text classification and natural language inference,” in Proc. 16th Conf. European Chapter of the Association for Computational Linguistics: Main Volume, 2020, pp. 255−269.
    [28]
    R. Nallapati, B. Zhou, C. dos Santos, Ç. Gülçehre, and B. Xiang, Abstractive text summarization using sequence-to-sequence RNNs and beyond,” in Proc. 20th SIGNLL Conf. Computational Natural Language Learning, Berlin, Germany, 2016, pp. 280−290.
    [29]
    S. Narayan, S. B. Cohen, and M. Lapata, “Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization,” in Proc. Conf. Empirical Methods in Natural Language Processing, Brussels, Belgium, 2018, pp. 1797−1807.
    [30]
    P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100, 000+ questions for machine comprehension of text,” in Proc. Conf. Empirical Methods in Natural Language Processing, Austin, USA, 2016, pp. 2383−2392.

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(11)  / Tables(6)

    Article Metrics

    Article views (19) PDF downloads(0) Cited by()

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return