研究者詳細

顔写真

リ ウンモウ
李 云蒙
Yunmeng Li
所属
大学院情報科学研究科 システム情報科学専攻 知能情報科学講座(自然言語処理学分野)
職名
特任助教(研究)
学位
  • 博士(情報科学)(東北大学)

  • 修士(情報科学)(東北大学)

e-Rad 研究者番号
71019750

論文 4

  1. MQM-Chat: Multidimensional Quality Metrics for Chat Translation

    Yunmeng Li, Jun Suzuki, Makoto Morishita, Kaori Abe, Kentaro Inui

    2024年8月29日

    詳細を見る 詳細を閉じる

    The complexities of chats pose significant challenges for machine translation models. Recognizing the need for a precise evaluation metric to address the issues of chat translation, this study introduces Multidimensional Quality Metrics for Chat Translation (MQM-Chat). Through the experiments of five models using MQM-Chat, we observed that all models generated certain fundamental errors, while each of them has different shortcomings, such as omission, overly correcting ambiguous source content, and buzzword issues, resulting in the loss of stylized information. Our findings underscore the effectiveness of MQM-Chat in evaluating chat translation, emphasizing the importance of stylized content and dialogue consistency for future studies.

  2. Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset

    Diana Galvan-Sosa, Gabrielle Gaudeau, Pride Kavumba, Yunmeng Li, Hongyi gu, Zheng Yuan, Keisuke Sakaguchi, Paula Buttery

    2025年3月31日

    詳細を見る 詳細を閉じる

    The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for users to distinguish good from bad explanations. To address this issue, we present Rubrik's CUBE, an education-inspired rubric and a dataset of 26k explanations, written and later quality-annotated using the rubric by both humans and six open- and closed-source LLMs. The CUBE dataset focuses on two reasoning and two language tasks, providing the necessary diversity for us to effectively test our proposed rubric. Using Rubrik, we find that explanations are influenced by both task and perceived difficulty. Low quality stems primarily from a lack of conciseness in LLM-generated explanations, rather than cohesion and word choice. The full dataset, rubric, and code are available at https://github.com/RubriksCube/rubriks_cube.

  3. An Investigation of Warning Erroneous Chat Translations in Cross-lingual Communication

    Yunmeng Li, Jun Suzuki, Makoto Morishita, Kaori Abe, Kentaro Inui

    2024年8月28日

    DOI: 10.18653/v1/2023.ijcnlp-srw.2  

    詳細を見る 詳細を閉じる

    Machine translation models are still inappropriate for translating chats, despite the popularity of translation software and plug-in applications. The complexity of dialogues poses significant challenges and can hinder crosslingual communication. Instead of pursuing a flawless translation system, a more practical approach would be to issue warning messages about potential mistranslations to reduce confusion. However, it is still unclear how individuals perceive these warning messages and whether they benefit the crowd. This paper tackles to investigate this question and demonstrates the warning messages' contribution to making chat translation systems effective.

  4. Chat Translation Error Detection for Assisting Cross-lingual Communications

    Yunmeng Li, Jun Suzuki, Makoto Morishita, Kaori Abe, Ryoko Tokuhisa, Ana Brassard, Kentaro Inui

    2023年8月2日

    DOI: 10.18653/v1/2022.eval4nlp-1.9  

    詳細を見る 詳細を閉じる

    In this paper, we describe the development of a communication support system that detects erroneous translations to facilitate crosslingual communications due to the limitations of current machine chat translation methods. We trained an error detector as the baseline of the system and constructed a new Japanese-English bilingual chat corpus, BPersona-chat, which comprises multiturn colloquial chats augmented with crowdsourced quality ratings. The error detector can serve as an encouraging foundation for more advanced erroneous translation detection systems.