Details of the Researcher

PHOTO

Heinzerling Benjamin Tobias
Section
Center for Language AI Research
Job title
Specially Appointed Associate Professor(Research)

Awards 5

  1. スポンサー賞(日立製作所賞)

    2026/03 言語処理学会第32回年次大会(NLP2026) 言語モデルにおける既知性判断のメカニズム

  2. 委員特別賞

    2026/03 言語処理学会第32回年次大会(NLP2026) TopK Language Models

  3. 委員特別賞

    2026/03 言語処理学会第32回年次大会(NLP2026) 大規模言語モデルの潜在言語は一貫しているべきか?

  4. 委員特別賞

    2026/03 言語処理学会第32回年次大会(NLP2026) Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

  5. 奨励賞

    2025/09 第20回YANSシンポジウム(YANS2025) マルチモーダルLLMのモダリティ間共有表現の解明

Papers 39

  1. Cell-Based Representation of Relational Binding in Language Models.

    Qin Dai, Benjamin Heinzerling, Kentaro Inui

    CoRR abs/2604.19052 2026/04

    DOI: 10.48550/arXiv.2604.19052  

  2. Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation.

    Joe Stacey, Hadas Orgad, Kentaro Inui, Benjamin Heinzerling, Nafise Sadat Moosavi

    CoRR abs/2604.11662 2026/04

    DOI: 10.48550/arXiv.2604.11662  

  3. Linear Representations of Hierarchical Concepts in Language Models.

    Masaki Sakata, Benjamin Heinzerling, Takumi Ito, Sho Yokoi, Kentaro Inui

    CoRR abs/2604.07886 2026/04

    DOI: 10.48550/arXiv.2604.07886  

  4. Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki.

    Mutsumi Sasaki, Go Kamoda, Ryosuke Takahashi, Kosuke Sato, Kentaro Inui, Keisuke Sakaguchi, Benjamin Heinzerling

    IJCNLP-AACL (Short Papers) 444-463 2025/12

    DOI: 10.18653/v1/2025.ijcnlp-short.36  

  5. How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders

    Tatsuro Inaba, Go Kamoda, Kentaro Inui, Masaru Isonuma, Yusuke Miyao, Yohei Oseki, Yu Takagi, Benjamin Heinzerling

    Findings of the Association for Computational Linguistics: EMNLP 2025 13458-13470 2025/11

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2025.findings-emnlp.725  

  6. On Entity Identification in Language Models

    Masaki Sakata, Benjamin Heinzerling, Sho Yokoi, Takumi Ito, Kentaro Inui

    Findings of the Association for Computational Linguistics: ACL 2025 16717-16741 2025/07

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2025.findings-acl.858  

  7. Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles

    Munachiso S Nwadike, Zangir Iklassov, Toluwani Aremu, Tatsuya Hiraoka, Benjamin Heinzerling, Velibor Bojkovic, Hilal AlQuabeh, Martin Takáč, Kentaro Inui

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) abs/2501.13491 25365-25377 2025/07

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2025.acl-long.1232  

  8. TopK Language Models.

    Ryosuke Takahashi, Tatsuro Inaba, Kentaro Inui, Benjamin Heinzerling

    CoRR abs/2506.21468 2025/06

    DOI: 10.48550/arXiv.2506.21468  

  9. Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance.

    Shintaro Ozaki, Tatsuya Hiraoka, Hiroto Otake, Hiroki Ouchi, Masaru Isonuma, Benjamin Heinzerling, Kentaro Inui, Taro Watanabe, Yusuke Miyao, Yohei Oseki, Yu Takagi

    CoRR abs/2505.21458 2025/05

    DOI: 10.48550/arXiv.2505.21458  

  10. Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge.

    Ying Zhang, Benjamin Heinzerling, Dongyuan Li, Ryoma Ishigaki, Yuta Hitomi, Kentaro Inui

    CoRR abs/2505.16178 2025/05

    DOI: 10.48550/arXiv.2505.16178  

  11. Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference

    Go Kamoda, Benjamin Heinzerling, Tatsuro Inaba, Keito Kudo, Keisuke Sakaguchi, Kentaro Inui

    Findings of the Association for Computational Linguistics: NAACL 2025 6324-6343 2025/04

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2025.findings-naacl.355  

  12. Understanding the Side Effects of Rank-One Knowledge Editing

    Ryosuke Takahashi, Go Kamoda, Benjamin Heinzerling, Keisuke Sakaguchi, Kentaro Inui

    Proceedings of the 8th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP 189-205 2025/04

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2025.blackboxnlp-1.11  

  13. The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces

    Ahmed Oumar El-Shangiti, Tatsuya Hiraoka, Hilal AlQuabeh, Benjamin Heinzerling, Kentaro Inui

    NAACL (Short Papers) 550-561 2025/04

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2025.naacl-short.47  

  14. The Curse of Popularity: Popular Entities have Catastrophic Side Effects when Deleting Knowledge from Language Models

    Ryosuke Takahashi, Go Kamoda, Benjamin Heinzerling, Keisuke Sakaguchi, Kentaro Inui

    2024/06/10

    More details Close

    Language models (LMs) encode world knowledge in their internal parameters through training. However, LMs may learn personal and confidential information from the training data, leading to privacy concerns such as data leakage. Therefore, research on knowledge deletion from LMs is essential. This study focuses on the knowledge stored in LMs and analyzes the relationship between the side effects of knowledge deletion and the entities related to the knowledge. Our findings reveal that deleting knowledge related to popular entities can have catastrophic side effects. Furthermore, this research is the first to analyze knowledge deletion in models trained on synthetic knowledge graphs, indicating a new direction for controlled experiments.

  15. ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation

    Ana Brassard, Benjamin Heinzerling, Keito Kudo, Keisuke Sakaguchi, Kentaro Inui

    2024/05/08

    More details Close

    Evaluating free-text explanations is a multifaceted, subjective, and labor-intensive task. Large language models (LLMs) present an appealing alternative due to their potential for consistency, scalability, and cost-efficiency. In this work, we present ACORN, a new dataset of 3,500 free-text explanations and aspect-wise quality ratings, and use it to gain insights into how LLMs evaluate explanations. We observed that replacing one of the human ratings sometimes maintained, but more often lowered the inter-annotator agreement across different settings and quality aspects, suggesting that their judgments are not always consistent with human raters. We further quantified this difference by comparing the correlation between LLM-generated ratings with majority-voted human ratings across different quality aspects. With the best system, Spearman's rank correlation ranged between 0.53 to 0.95, averaging 0.72 across aspects, indicating moderately high but imperfect alignment. Finally, we considered the alternative of using an LLM as an additional rater when human raters are scarce, and measured the correlation between majority-voted labels with a limited human pool and LLMs as an additional rater, compared to the original gold labels. While GPT-4 improved the outcome when there were only two human raters, in all other observed cases, LLMs were neutral to detrimental when there were three or more human raters. We publicly release the dataset to support future improvements in LLM-in-the-loop evaluation here: https://github.com/a-brassard/ACORN.

  16. Monotonic Representation of Numeric Attributes in Language Models.

    Benjamin Heinzerling, Kentaro Inui

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics 175-195 2024

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2024.acl-short.18  

  17. Representational Analysis of Binding in Language Models

    Qin Dai, Benjamin Heinzerling, Kentaro Inui

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing 31st 17468-17493 2024

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2024.emnlp-main.967  

    ISSN: 2188-4420

  18. Test-time Augmentation for Factual Probing

    Go Kamoda, Benjamin Heinzerling, Keisuke Sakaguchi, Kentaro Inui

    EMNLP (Findings) 29th 3650-3661 2023/10/26

    DOI: 10.18653/v1/2023.findings-emnlp.236  

    ISSN: 2188-4420

    More details Close

    Factual probing is a method that uses prompts to test if a language model "knows" certain world knowledge facts. A problem in factual probing is that small changes to the prompt can lead to large changes in model output. Previous work aimed to alleviate this problem by optimizing prompts via text mining or fine-tuning. However, such approaches are relation-specific and do not generalize to unseen relation types. Here, we propose to use test-time augmentation (TTA) as a relation-agnostic method for reducing sensitivity to prompt variations by automatically augmenting and ensembling prompts at test time. Experiments show improved model calibration, i.e., with TTA, model confidence better reflects prediction accuracy. Improvements in prediction accuracy are observed for some models, but for other models, TTA leads to degradation. Error analysis identifies the difficulty of producing high-quality prompt variations as the main challenge for TTA.

  19. Examining the effect of whitening on static and contextualized word embeddings.

    Shota Sasaki, Benjamin Heinzerling, Jun Suzuki, Kentaro Inui

    Information Processing and Management 60 (3) 103272-103272 2023/05

    DOI: 10.1016/j.ipm.2023.103272  

    ISSN: 0306-4573

  20. Tracing and Manipulating Intermediate Values in Neural Math Problem Solvers

    Yuta Matsumoto, Benjamin Heinzerling, Masashi Yoshikawa, Kentaro Inui

    CoRR abs/2301.06758 (4) 2023/01/17

    DOI: 10.48550/arXiv.2301.06758  

    ISSN: 2185-8314

    More details Close

    How language models process complex input that requires multiple steps of inference is not well understood. Previous research has shown that information about intermediate values of these inputs can be extracted from the activations of the models, but it is unclear where that information is encoded and whether that information is indeed used during inference. We introduce a method for analyzing how a Transformer model processes these inputs by focusing on simple arithmetic problems and their intermediate values. To trace where information about intermediate values is encoded, we measure the correlation between intermediate values and the activations of the model using principal component analysis (PCA). Then, we perform a causal intervention by manipulating model weights. This intervention shows that the weights identified via tracing are not merely correlated with intermediate values, but causally related to model predictions. Our findings show that the model has a locality to certain intermediate values, and this is useful for enhancing the interpretability of the models.

  21. Prompting for explanations improves Adversarial NLI. Is this true? Yes it is true because it weakens superficial cues.

    Pride Kavumba, Ana Brassard, Benjamin Heinzerling, Kentaro Inui

    Findings of the Association for Computational Linguistics: EACL 2023 2120-2135 2023

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2023.findings-eacl.162  

  22. Can LMs Store and Retrieve 1-to-N Relational Knowledge?

    Haruki Nagasawa, Benjamin Heinzerling, Kazuma Kokuta, Kentaro Inui

    ACL (student) 130-138 2023

    DOI: 10.18653/v1/2023.acl-srw.22  

  23. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

    BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, Dragomir Radev, Eduardo González Ponferrada, Efrat Levkovizh, Ethan Kim, Eyal Bar Natan, Francesco De Toni, Gérard Dupont, Germán Kruszewski, Giada Pistilli, Hady Elsahar, Hamza Benyamina, Hieu Tran, Ian Yu, Idris Abdulmumin, Isaac Johnson, Itziar Gonzalez-Dios, Javier de la Rosa, Jenny Chim, Jesse Dodge, Jian Zhu, Jonathan Chang, Jörg Frohberg, Joseph Tobing, Joydeep Bhattacharjee, Khalid Almubarak, Kimbo Chen, Kyle Lo, Leandro Von Werra, Leon Weber, Long Phan, Loubna Ben allal, Ludovic Tanguy, Manan Dey, Manuel Romero Muñoz, Maraim Masoud, María Grandury, Mario Šaško, Max Huang, Maximin Coavoux, Mayank Singh, Mike Tian-Jian Jiang, Minh Chien Vu, Mohammad A. Jauhar, Mustafa Ghaleb, Nishant Subramani, Nora Kassner, Nurulaqilla Khamis, Olivier Nguyen, Omar Espejel, Ona de Gibert, Paulo Villegas, Peter Henderson, Pierre Colombo, Priscilla Amuok, Quentin Lhoest, Rheza Harliman, Rishi Bommasani, Roberto Luis López, Rui Ribeiro, Salomey Osei, Sampo Pyysalo, Sebastian Nagel, Shamik Bose, Shamsuddeen Hassan Muhammad, Shanya Sharma, Shayne Longpre, Somaieh Nikpoor, Stanislav Silberberg, Suhas Pai, Sydney Zink, Tiago Timponi Torrent, Timo Schick, Tristan Thrush, Valentin Danchev, Vassilina Nikoulina, Veronika Laippala, Violette Lepercq, Vrinda Prabhu, Zaid Alyafeai, Zeerak Talat, Arun Raja, Benjamin Heinzerling, Chenglei Si, Davut Emre Taşar, Elizabeth Salesky, Sabrina J. Mielke, Wilson Y. Lee, Abheesht Sharma, Andrea Santilli, Antoine Chaffin, Arnaud Stiegler, Debajyoti Datta, Eliza Szczechla, Gunjan Chhablani, Han Wang, Harshit Pandey, Hendrik Strobelt, Jason Alan Fries, Jos Rozen, Leo Gao, Lintang Sutawika, M Saiful Bari, Maged S. Al-shaibani, Matteo Manica, Nihal Nayak, Ryan Teehan, Samuel Albanie, Sheng Shen, Srulik Ben-David, Stephen H. Bach, Taewoon Kim, Tali Bers, Thibault Fevry, Trishala Neeraj, Urmish Thakker, Vikas Raunak, Xiangru Tang, Zheng-Xin Yong, Zhiqing Sun, Shaked Brody, Yallow Uri, Hadar Tojarieh, Adam Roberts, Hyung Won Chung, Jaesung Tae, Jason Phang, Ofir Press, Conglong Li, Deepak Narayanan, Hatim Bourfoune, Jared Casper, Jeff Rasley, Max Ryabinin, Mayank Mishra, Minjia Zhang, Mohammad Shoeybi, Myriam Peyrounette, Nicolas Patry, Nouamane Tazi, Omar Sanseviero, Patrick von Platen, Pierre Cornette, Pierre François Lavallée, Rémi Lacroix, Samyam Rajbhandari, Sanchit Gandhi, Shaden Smith, Stéphane Requena, Suraj Patil, Tim Dettmers, Ahmed Baruwa, Amanpreet Singh, Anastasia Cheveleva, Anne-Laure Ligozat, Arjun Subramonian, Aurélie Névéol, Charles Lovering, Dan Garrette, Deepak Tunuguntla, Ehud Reiter, Ekaterina Taktasheva, Ekaterina Voloshina, Eli Bogdanov, Genta Indra Winata, Hailey Schoelkopf, Jan-Christoph Kalo, Jekaterina Novikova, Jessica Zosa Forde, Jordan Clive, Jungo Kasai, Ken Kawamura, Liam Hazan, Marine Carpuat, Miruna Clinciu, Najoung Kim, Newton Cheng, Oleg Serikov, Omer Antverg, Oskar van der Wal, Rui Zhang, Ruochen Zhang, Sebastian Gehrmann, Shachar Mirkin, Shani Pais, Tatiana Shavrina, Thomas Scialom, Tian Yun, Tomasz Limisiewicz, Verena Rieser, Vitaly Protasov, Vladislav Mikhailov, Yada Pruksachatkun, Yonatan Belinkov, Zachary Bamberger, Zdeněk Kasner, Alice Rueda, Amanda Pestana, Amir Feizpour, Ammar Khan, Amy Faranak, Ana Santos, Anthony Hevia, Antigona Unldreaj, Arash Aghagol, Arezoo Abdollahi, Aycha Tammour, Azadeh HajiHosseini, Bahareh Behroozi, Benjamin Ajibade, Bharat Saxena, Carlos Muñoz Ferrandis, Daniel McDuff, Danish Contractor, David Lansky, Davis David, Douwe Kiela, Duong A. Nguyen, Edward Tan, Emi Baylor, Ezinwanne Ozoani, Fatima Mirza, Frankline Ononiwu, Habib Rezanejad, Hessie Jones, Indrani Bhattacharya, Irene Solaiman, Irina Sedenko, Isar Nejadgholi, Jesse Passmore, Josh Seltzer, Julio Bonis Sanz, Livia Dutra, Mairon Samagaio, Maraim Elbadri, Margot Mieskes, Marissa Gerchick, Martha Akinlolu, Michael McKenna, Mike Qiu, Muhammed Ghauri, Mykola Burynok, Nafis Abrar, Nazneen Rajani, Nour Elkott, Nour Fahmy, Olanrewaju Samuel, Ran An, Rasmus Kromann, Ryan Hao, Samira Alizadeh, Sarmad Shubber, Silas Wang, Sourav Roy, Sylvain Viguier, Thanh Le, Tobi Oyebade, Trieu Le, Yoyo Yang, Zach Nguyen, Abhinav Ramesh Kashyap, Alfredo Palasciano, Alison Callahan, Anima Shukla, Antonio Miranda-Escalada, Ayush Singh, Benjamin Beilharz, Bo Wang, Caio Brito, Chenxi Zhou, Chirag Jain, Chuxin Xu, Clémentine Fourrier, Daniel León Periñán, Daniel Molano, Dian Yu, Enrique Manjavacas, Fabio Barth, Florian Fuhrimann, Gabriel Altay, Giyaseddin Bayrak, Gully Burns, Helena U. Vrabec, Imane Bello, Ishani Dash, Jihyun Kang, John Giorgi, Jonas Golde, Jose David Posada, Karthik Rangasai Sivaraman, Lokesh Bulchandani, Lu Liu, Luisa Shinzato, Madeleine Hahn de Bykhovetz, Maiko Takeuchi, Marc Pàmies, Maria A Castillo, Marianna Nezhurina, Mario Sänger, Matthias Samwald, Michael Cullan, Michael Weinberg, Michiel De Wolf, Mina Mihaljcic, Minna Liu, Moritz Freidank, Myungsun Kang, Natasha Seelam, Nathan Dahlberg, Nicholas Michio Broad, Nikolaus Muellner, Pascale Fung, Patrick Haller, Ramya Chandrasekhar, Renata Eisenberg, Robert Martin, Rodrigo Canalli, Rosaline Su, Ruisi Su, Samuel Cahyawijaya, Samuele Garda, Shlok S Deshmukh, Shubhanshu Mishra, Sid Kiblawi, Simon Ott, Sinee Sang-aroonsiri, Srishti Kumar, Stefan Schweter, Sushil Bharati, Tanmay Laud, Théo Gigant, Tomoya Kainuma, Wojciech Kusa, Yanis Labrak, Yash Shailesh Bajaj, Yash Venkatraman, Yifan Xu, Yingxin Xu, Yu Xu, Zhe Tan, Zhongli Xie, Zifan Ye, Mathilde Bras, Younes Belkada, Thomas Wolf

    2022/11/09

    More details Close

    Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed by resource-rich organizations and are frequently kept from the public. As a step towards democratizing this powerful technology, we present BLOOM, a 176B-parameter open-access language model designed and built thanks to a collaboration of hundreds of researchers. BLOOM is a decoder-only Transformer language model that was trained on the ROOTS corpus, a dataset comprising hundreds of sources in 46 natural and 13 programming languages (59 in total). We find that BLOOM achieves competitive performance on a wide variety of benchmarks, with stronger results after undergoing multitask prompted finetuning. To facilitate future research and applications using LLMs, we publicly release our models and code under the Responsible AI License.

  24. Cross-stitching Text and Knowledge Graph Encoders for Distantly Supervised Relation Extraction

    Qin Dai, Benjamin Heinzerling, Kentaro Inui

    Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing(EMNLP) 29th 6947-6958 2022/11/02

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2022.emnlp-main.467  

    ISSN: 2188-4420

    More details Close

    Bi-encoder architectures for distantly-supervised relation extraction are designed to make use of the complementary information found in text and knowledge graphs (KG). However, current architectures suffer from two drawbacks. They either do not allow any sharing between the text encoder and the KG encoder at all, or, in case of models with KG-to-text attention, only share information in one direction. Here, we introduce cross-stitch bi-encoders, which allow full interaction between the text encoder and the KG encoder via a cross-stitch mechanism. The cross-stitch mechanism allows sharing and updating representations between the two encoders at any layer, with the amount of sharing being dynamically controlled via cross-attention-based gates. Experimental results on two relation extraction benchmarks from two different domains show that enabling full interaction between the two encoders yields strong improvements.

  25. COPA-SSE: Semi-structured Explanations for Commonsense Reasoning

    Ana Brassard, Benjamin Heinzerling, Pride Kavumba, Kentaro Inui

    Proceedings of the Thirteenth Language Resources and Evaluation Conference(LREC) 3994-4000 2022/01/18

    Publisher: European Language Resources Association

    More details Close

    We present Semi-Structured Explanations for COPA (COPA-SSE), a new crowdsourced dataset of 9,747 semi-structured, English common sense explanations for Choice of Plausible Alternatives (COPA) questions. The explanations are formatted as a set of triple-like common sense statements with ConceptNet relations but freely written concepts. This semi-structured format strikes a balance between the high quality but low coverage of structured data and the lower quality but high coverage of free-form crowdsourcing. Each explanation also includes a set of human-given quality ratings. With their familiar format, the explanations are geared towards commonsense reasoners operating on knowledge graphs and serve as a starting point for ongoing work on improving such systems. The dataset is available at https://github.com/a-brassard/copa-sse.

  26. Learning to Learn to be Right for the Right Reasons

    Pride Kavumba, Benjamin Heinzerling, Ana Brassard, Kentaro Inui

    Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies(NAACL-HLT) 3890-3898 2021/04/23

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2021.naacl-main.304  

    More details Close

    Improving model generalization on held-out data is one of the core objectives in commonsense reasoning. Recent work has shown that models trained on the dataset with superficial cues tend to perform well on the easy test set with superficial cues but perform poorly on the hard test set without superficial cues. Previous approaches have resorted to manual methods of encouraging models not to overfit to superficial cues. While some of the methods have improved performance on hard instances, they also lead to degraded performance on easy instances. Here, we propose to explicitly learn a model that does well on both the easy test set with superficial cues and hard test set without superficial cues. Using a meta-learning objective, we learn such a model that improves performance on both the easy test set and the hard test set. By evaluating our models on Choice of Plausible Alternatives (COPA) and Commonsense Explanation, we show that our proposed method leads to improved performance on both the easy test set and the hard test set upon which we observe up to 16.5 percentage points improvement over the baseline.

  27. Language Models as Knowledge Bases: On Entity Representations, Storage Capacity, and Paraphrased Queries

    Benjamin Heinzerling, Kentaro Inui

    Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume(EACL) 1772-1791 2020/08/20

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/2021.eacl-main.153  

    More details Close

    Pretrained language models have been suggested as a possible alternative or complement to structured knowledge bases. However, this emerging LM-as-KB paradigm has so far only been considered in a very limited setting, which only allows handling 21k entities whose single-token name is found in common LM vocabularies. Furthermore, the main benefit of this paradigm, namely querying the KB using a variety of natural language paraphrases, is underexplored so far. Here, we formulate two basic requirements for treating LMs as KBs: (i) the ability to store a large number facts involving a large number of entities and (ii) the ability to query stored facts. We explore three entity representations that allow LMs to represent millions of entities and present a detailed case study on paraphrased querying of world knowledge in LMs, thereby providing a proof-of-concept that language models can indeed serve as knowledge bases.

  28. NLP's clever hans moment has arrived

    Heinzerling, B.

    Journal of Cognitive Science 21 (1) 2020

    ISSN: 1976-6939 1598-2327

  29. When Choosing Plausible Alternatives, Clever Hans can be Clever

    Pride Kavumba, Naoya Inoue, Benjamin Heinzerling, Keshav Singh, Paul Reisert, Kentaro Inui

    2019/11/01

    More details Close

    Pretrained language models, such as BERT and RoBERTa, have shown large improvements in the commonsense reasoning benchmark COPA. However, recent work found that many improvements in benchmarks of natural language understanding are not due to models learning the task, but due to their increasing ability to exploit superficial cues, such as tokens that occur more often in the correct answer than the wrong one. Are BERT's and RoBERTa's good performance on COPA also caused by this? We find superficial cues in COPA, as well as evidence that BERT exploits these cues. To remedy this problem, we introduce Balanced COPA, an extension of COPA that does not suffer from easy-to-exploit single token cues. We analyze BERT's and RoBERTa's performance on original and Balanced COPA, finding that BERT relies on superficial cues when they are present, but still achieves comparable performance once they are made ineffective, suggesting that BERT learns the task to a certain degree when forced to. In contrast, RoBERTa does not appear to rely on superficial cues.

  30. Riposte! A Large Corpus of Counter-Arguments

    Paul Reisert, Benjamin Heinzerling, Naoya Inoue, Shun Kiyono, Kentaro Inui

    2019/10/08

    More details Close

    Constructive feedback is an effective method for improving critical thinking skills. Counter-arguments (CAs), one form of constructive feedback, have been proven to be useful for critical thinking skills. However, little work has been done for constructing a large-scale corpus of them which can drive research on automatic generation of CAs for fallacious micro-level arguments (i.e. a single claim and premise pair). In this work, we cast providing constructive feedback as a natural language processing task and create Riposte!, a corpus of CAs, towards this goal. Produced by crowdworkers, Riposte! contains over 18k CAs. We instruct workers to first identify common fallacy types and produce a CA which identifies the fallacy. We analyze how workers create CAs and construct a baseline model based on our analysis.

  31. On the Importance of Subword Information for Morphological Tasks in Truly Low-Resource Languages

    Yi Zhu, Benjamin Heinzerling, Ivan Vulić, Michael Strube, Roi Reichart, Anna Korhonen

    Proceedings of the 23rd Conference on Computational Natural Language Learning(CoNLL) 216-226 2019/09/26

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/K19-1021  

    More details Close

    Recent work has validated the importance of subword information for word representation learning. Since subwords increase parameter sharing ability in neural models, their value should be even more pronounced in low-data regimes. In this work, we therefore provide a comprehensive analysis focused on the usefulness of subwords for word representation learning in truly low-resource scenarios and for three representative morphological tasks: fine-grained entity typing, morphological tagging, and named entity recognition. We conduct a systematic study that spans several dimensions of comparison: 1) type of data scarcity which can stem from the lack of task-specific training data, or even from the lack of unannotated data required to train word embeddings, or both; 2) language type by working with a sample of 16 typologically diverse languages including some truly low-resource ones (e.g. Rusyn, Buryat, and Zulu); 3) the choice of the subword-informed word representation method. Our main results show that subword-informed models are universally useful across all language types, with large gains over subword-agnostic embeddings. They also suggest that the effective use of subwords largely depends on the language (type) and the task at hand, as well as on the amount of available data for training the embeddings and task-based models, where having sufficient in-task data is a more critical requirement.

  32. Fine-Grained Entity Typing in Hyperbolic Space

    Federico López, Benjamin Heinzerling, Michael Strube

    2019/06/06

    More details Close

    How can we represent hierarchical information present in large type inventories for entity typing? We study the ability of hyperbolic embeddings to capture hierarchical relations between mentions in context and their target types in a shared vector space. We evaluate on two datasets and investigate two different techniques for creating a large hierarchical entity type inventory: from an expert-generated ontology and by automatically mining type co-occurrences. We find that the hyperbolic model yields improvements over its Euclidean counterpart in some, but not all cases. Our analysis suggests that the adequacy of this geometry depends on the granularity of the type inventory and the way hierarchical relations are inferred.

  33. Sequence Tagging with Contextual and Non-Contextual Subword Representations: A Multilingual Evaluation

    Benjamin Heinzerling, Michael Strube

    Proceedings of the 57th Conference of the Association for Computational Linguistics 273-291 2019/06/04

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/p19-1027  

    More details Close

    Pretrained contextual and non-contextual subword embeddings have become available in over 250 languages, allowing massively multilingual NLP. However, while there is no dearth of pretrained embeddings, the distinct lack of systematic evaluations makes it difficult for practitioners to choose between them. In this work, we conduct an extensive evaluation comparing non-contextual subword embeddings, namely FastText and BPEmb, and a contextual representation method, namely BERT, on multilingual named entity recognition and part-of-speech tagging. We find that overall, a combination of BERT, BPEmb, and character representations works best across languages and tasks. A more detailed analysis reveals different strengths and weaknesses: Multilingual BERT performs well in medium- to high-resource languages, but is outperformed by non-contextual subword embeddings in a low-resource setting.

  34. Aspects of Coherence for Entity Analysis

    Benjamin Heinzerling

    2019

    DOI: 10.11588/heidok.00026117  

  35. What's Important in a Text? An Extensive Evaluation of Linguistic Annotations for Summarization.

    Markus Zopf, Teresa Botschen, Tobias Falke, Benjamin Heinzerling, Ana Marasovic, Todor Mihaylov, Avinesh P. V. S., Eneldo Loza Mencía, Johannes Fürnkranz, Anette Frank

    Fifth International Conference on Social Networks Analysis, Management and Security(SNAMS) 272-277 2018

    Publisher: IEEE

    DOI: 10.1109/SNAMS.2018.8554853  

  36. BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages

    Benjamin Heinzerling, Michael Strube

    2017/10/05

    More details Close

    We present BPEmb, a collection of pre-trained subword unit embeddings in 275 languages, based on Byte-Pair Encoding (BPE). In an evaluation using fine-grained entity typing as testbed, BPEmb performs competitively, and for some languages bet- ter than alternative subword approaches, while requiring vastly fewer resources and no tokenization. BPEmb is available at https://github.com/bheinzerling/bpemb

  37. Revisiting Selectional Preferences for Coreference Resolution

    Benjamin Heinzerling, Nafise Sadat Moosavi, Michael Strube

    2017/07/20

    More details Close

    Selectional preferences have long been claimed to be essential for coreference resolution. However, they are mainly modeled only implicitly by current coreference resolvers. We propose a dependency-based embedding model of selectional preferences which allows fine-grained compatibility judgments with high coverage. We show that the incorporation of our model improves coreference resolution performance on the CoNLL dataset, matching the state-of-the-art results of a more complex system. However, it comes with a cost that makes it debatable how worthwhile such improvements are.

  38. Trust, but Verify! Better Entity Linking through Automatic Verification.

    Benjamin Heinzerling, Michael Strube 0001, Chin-Yew Lin

    Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics 828-838 2017

    Publisher: Association for Computational Linguistics

    DOI: 10.18653/v1/e17-1078  

  39. Visual Error Analysis for Entity Linking.

    Benjamin Heinzerling, Michael Strube 0001

    Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing 37-42 2015

    Publisher: The Association for Computer Linguistics

    DOI: 10.3115/v1/p15-4007  

Show all ︎Show first 5

Misc. 27

  1. 似た単語の知識ニューロンは似た形成過程を経る

    有山知希, 有山知希, HEINZERLING Benjamin, HEINZERLING Benjamin, 穀田一真, 穀田一真, 乾健太郎, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 31st 2025

    ISSN: 2188-4420

  2. スパースオートエンコーダーを用いた大規模言語モデルのチェックポイント横断分析

    稲葉達郎, 稲葉達郎, 乾健太郎, 乾健太郎, 乾健太郎, 宮尾祐介, 宮尾祐介, 大関洋平, HEINZERLING Benjamin, HEINZERLING Benjamin, 高木優

    言語処理学会年次大会発表論文集(Web) 31st 2025

    ISSN: 2188-4420

  3. LMは日本の時系列構造をどうエンコードするか

    佐々木睦史, 鴨田豪, 高橋良允, HEINZERLING Benjamin, HEINZERLING Benjamin, 坂口慶祐, 坂口慶祐

    言語処理学会年次大会発表論文集(Web) 31st 2025

    ISSN: 2188-4420

  4. 言語モデルの内部表現における文法情報の局所性について

    佐藤宏亮, 鴨田豪, HEINZERLING Benjamin, HEINZERLING Benjamin, 坂口慶祐, 坂口慶祐

    言語処理学会年次大会発表論文集(Web) 31st 2025

    ISSN: 2188-4420

  5. Anchoringを行う生成的関係抽出

    広田航, 高橋洸丞, HEINZERLING Benjamin, HEINZERLING Benjamin, DAI Qin, 近江崇宏, 乾健太郎, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 31st 2025

    ISSN: 2188-4420

  6. 継続事前学習によるLLMの知識獲得

    高橋洸丞, 近江崇宏, 有馬幸介, HEINZERLING Benjamin, HEINZERLING Benjamin, DAI Qin, 乾健太郎, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 31st 2025

    ISSN: 2188-4420

  7. Internal Representations of Familiarity Judgments in Language Models

    佐藤魁, 高橋良允, HEINZERLING Benjamin, HEINZERLING Benjamin, 田中健史朗, ZHAO Yufeng, 坂井吉弘, 井之上直也, 井之上直也, 乾健太郎, 乾健太郎, 乾健太郎

    人工知能学会全国大会論文集(Web) 39th 2025

    ISSN: 2758-7347

  8. Analysis of Personality Trait Directions and Intervention Feasibility in Large Language Models

    石垣龍馬, 石垣龍馬, HEINZERLING Benjamin, HEINZERLING Benjamin, ZHANG Ying, 人見雄太, 乾健太郎, 乾健太郎, 乾健太郎

    情報処理学会研究報告(Web) 2025 (NL-264) 2025

  9. Analysis of Internal Representations of Knowledge with Expressions of Familiarity

    田中健史朗, 坂井吉弘, ZHAO Yufeng, 井之上直也, 井之上直也, 佐藤魁, 高橋良允, HEINZERLING Benjamin, HEINZERLING Benjamin, 乾健太郎, 乾健太郎, 乾健太郎

    人工知能学会全国大会論文集(Web) 39th 2025

    ISSN: 2758-7347

  10. 言語モデルのパラメータから探るDetokenizationメカニズム

    鴨田豪, HEINZERLING Benjamin, HEINZERLING Benjamin, 稲葉達郎, 工藤慧音, 工藤慧音, 坂口慶祐, 坂口慶祐, 乾健太郎, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 31st 2025

    ISSN: 2188-4420

  11. 言語モデルからの知識削除:頻出実体の知識は副作用が破滅的

    高橋良允, 鴨田豪, HEINZERLING Benjamin, HEINZERLING Benjamin, 坂口慶祐, 坂口慶祐, 乾健太郎, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 30th 2024

    ISSN: 2188-4420

  12. 事前学習済み言語モデルによるエンティティの概念化

    坂田将樹, 坂田将樹, 横井祥, 横井祥, HEINZERLING Benjamin, HEINZERLING Benjamin, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 29th 2023

    ISSN: 2188-4420

  13. Sequence-to-sequenceモデルを用いた一対多関係知識の記憶とその取り出し

    長澤春希, HEINZERLING Benjamin, HEINZERLING Benjamin, 穀田一真, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 29th 2023

    ISSN: 2188-4420

  14. 事前学習済み言語モデルの知識に基づく演繹推論能力の調査

    穀田一真, 長澤春希, HEINZERLING Benjamin, HEINZERLING Benjamin, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 29th 2023

    ISSN: 2188-4420

  15. ニューラル数式ソルバーにおける途中結果の追跡と操作

    松本悠太, HEINZERLING Benjamin, HEINZERLING Benjamin, 吉川将司, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 29th 2023

    ISSN: 2188-4420

  16. 言語モデルの学習における知識ニューロンの形成過程について

    有山知希, HEINZERLING Benjamin, HEINZERLING Benjamin, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 29th 2023

    ISSN: 2188-4420

  17. 因果的プロンプトによるNLIの敵対的ロバスト性の強化

    KAVUMBA Pride, KAVUMBA Pride, BRASSARD Ana, BRASSARD Ana, HEINZERLING Benjamin, HEINZERLING Benjamin, 坂口慶祐, 坂口慶祐, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 29th 2023

    ISSN: 2188-4420

  18. 白色化が単語埋め込みに及ぼす効果の検証

    佐々木翔大, 佐々木翔大, HEINZERLING Benjamin, HEINZERLING Benjamin, 鈴木潤, 鈴木潤, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 29th 2023

    ISSN: 2188-4420

  19. End-to-End学習可能な記号処理層の検討と数量推論への応用における課題の分析

    吉川将司, 吉川将司, HEINZERLING Benjamin, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 28th 2022

    ISSN: 2188-4420

  20. 四則演算を用いたTransformerの再帰的構造把握能力の調査

    松本悠太, 吉川将司, 吉川将司, HEINZERLING Benjamin, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 28th 2022

    ISSN: 2188-4420

  21. Transformerモデルのニューロンには局所的に概念についての知識がエンコードされている

    有山知希, HEINZERLING Benjamin, HEINZERLING Benjamin, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 28th 2022

    ISSN: 2188-4420

  22. ニューラル言語モデルによる一対多関係知識の記憶と操作

    長澤春希, HEINZERLING Benjamin, HEINZERLING Benjamin, 乾健太郎, 乾健太郎

    言語処理学会年次大会発表論文集(Web) 28th 2022

    ISSN: 2188-4420

  23. Universal Graph based Relation Extraction

    DAI Qin, HEINZERLING Benjamin, INUI Kentaro, INUI Kentaro

    言語処理学会年次大会発表論文集(Web) 28th 2022

    ISSN: 2188-4420

  24. Universal Graph based Distantly Supervised Relation Extraction

    Dai Qin, Heinzerling Benjamin, Heinzerling Benjamin, Inoue Naoya, Inui Kentaro, Inui Kentaro

    自然言語処理(Web) 29 (4) 2022

    ISSN: 2185-8314

  25. None the wiser? Adding “None” Mitigates Superficial Cues in Multiple-Choice Benchmarks

    KAVUMBA Pride, KAVUMBA Pride, BRASSARD Ana, BRASSARD Ana, HEINZERLING Benjamin, HEINZERLING Benjamin, INOUE Naoya, INOUE Naoya, INUI Kentaro, INUI Kentaro

    言語処理学会年次大会発表論文集(Web) 27th 2021

    ISSN: 2188-4420

  26. Balanced COPA: Countering Superficial Cues in Causal Reasoning

    KAVUMBA Pride, INOUE Naoya, INOUE Naoya, HEINZERLING Benjamin, HEINZERLING Benjamin, SINGH Keshav, REISERT Paul, REISERT Paul, INUI Kentaro, INUI Kentaro

    言語処理学会年次大会発表論文集(Web) 26th 2020

    ISSN: 2188-4420

  27. Constructing a Large Corpus of Counter-Arguments

    REISERT Paul, REISERT Paul, HEINZERLING Benjamin, HEINZERLING Benjamin, INOUE Naoya, INOUE Naoya, KIYONO Shun, KIYONO Shun, INUI Kentaro, INUI Kentaro

    計測自動制御学会システム・情報部門学術講演会講演論文集(CD-ROM) 2019 2019

Show all ︎Show first 5

Research Projects 2

  1. Computational Modeling of Argumentation Understanding

    Offer Organization: Japan Society for the Promotion of Science

    System: Grants-in-Aid for Scientific Research

    Category: Grant-in-Aid for Scientific Research (A)

    Institution: Tohoku University

    2022/04/01 - 2027/03/31

  2. Knowledge-Base-Grounded Language Models

    HEINZERLING BENJAMIN

    Offer Organization: 日本学術振興会

    System: 科学研究費助成事業

    Category: 若手研究

    Institution: 国立研究開発法人理化学研究所

    2021/04/01 - 2024/03/31

    More details Close

    In the second year of the grant period we devised, implemented, and evaluated a neural network model architecture for combining symbolic information from a knowledge base with non-symbolic representations of textual information. <BR> The model architecture consists of two multi-layered encoder stacks, one for symbolic information and one for textual information. The two encoder stacks interact at arbitrary layers via cross-attention and gates that determine how much one encoder use the information of the encoder to update its internal representations. <BR> Evaluation on distantly-supervised relation extraction benchmarks demonstrated state-of-the-art performance. The work was published at EMNLP 2022, as well as domestically at NLP 2023 where it was recognized as an outstanding paper.