Semantic Shifts in Low-Resource Domains: Computational Humanities and Natural Language Processing in Dialogue

Semantic Shifts in Low-Resource Domains: Computational Humanities and Natural Language Processing in Dialogue

Recognizing and describing lexical semantic change is a fundamental task for the humanities, but modern computational methods encounter significant obstacles in this regard. Large language models (LLMs) are powerful, but they require huge training corpora that are not available in the low-resource domains of historical and specialized humanities research. This limits the practical applicability and scientific value of the computational methods for analyzing language change that have been researched in recent years.

The interdisciplinary research group SILD (Semantic shifts In Low-resource Domains) was conceived to overcome these problems. By bringing together researchers from different disciplines such as Natural Language Processing (NLP), Visual Analytics, and Computational Humanities, SILD will develop novel methods for detecting semantic change that are data-efficient, interpretable, and tailored to real humanities research questions. Our focus extends beyond individual words to the analysis of multi-word expressions (MWEs), which are crucial carriers of historical and cultural meaning. A key technical innovation is the close integration of LLMs with structured domain knowledge (e.g., from historical dictionaries) represented in knowledge graphs (KGs) to improve language representations in data-poor environments.

SILD’s joint work program is structured in a cyclical workflow of three interacting project groups. Three projects from the computational humanities (CH) form the basis and provide corpora from low-resource domains, evaluation data, and subject expertise from the diachronic study of Hong Kong English (linguistics), early modern German drama (literary studies), and medieval Arabic-Latin scientific translations (philosophy). Three NLP projects use this data and KGs to develop integrated, sample-efficient language models and robust methods for identifying single-word and multi-word expressions. Two transfer projects bridge the gap between NLP and CH by creating visual analytics tools for exploring and interpreting the models and developing a metamodel of semantic change to guide the analysis of meaning and the evaluation.

The project will create valuable, freely accessible resources, including annotated corpora, linguistic KGs, evaluation datasets, and domain-adapted LLMs, as well as new methods for creating, analyzing, and visualizing LLMs for powerful semantic analysis in low-resource domains. The project is thus pioneering new computer-assisted methods and state-of-the-art models for a new, application-oriented paradigm in computational humanities.

Funding

  • German Research Foundation
German Research Foundation

German Research Foundation (DFG)

Official Channels

Zum Inhaltsanfang

© Universität Konstanz 2026