Skip to content
Show Almanac

Talk

LING Circle: Dr. Erhard Hinrichs | The Added Value of LLMs for Computational Lexicography

When

· 3:30 PM MDT

Shown in the venue’s time zone (America/Denver), not yours.

Where

Tickets

No on-sale date recorded

This event came from a public calendar feed. Calendar feeds publish when a show happens, not when its tickets go on sale — so this is genuinely unknown rather than zero, and rather than “already on sale”.

Official ticket page ↗

What the source said

The Added Value of LLMs for Computational Lexicography Due to the success of large language models (LLMs), first introduced by (Vaswani et al. 2017; Devlin et al. 2018), a wide range of scientific disciplines explore their use for tasks which previously used human resources or more traditional technologies. LLMs have also been explored in lexicography to support experts in constructing and maintaining dictionaries (McKean and Fitzgerald 2024; Marcondes et al. 2024). In this presentation, I report on joint work with Reinhild Barkey, Marie Hinrichs, and Claus Zinn on using three prominent representatives of LLMs, namely, ChatGPT (Brown et al. 2020), Deepseek (DeepSeek-AI et al. 2024) and Claude (Enis and Hopkins 2024), in the context of GermaNet (Hamp and Feldweg 1997; Henrich and Hinrichs 2010b). More specifically, we study the quality of LLM-generated content on a single but crucial task: the generation of example sentences for word senses to illustrate their use in context. This study shows that while LLMs can support lexicographers, they cannot replace them. Despite large-scale training data, LLMs lack the linguistic competence and world knowledge of trained experts, making human validation of generated examples essential for high-quality resources such as GermaNet. The two best-performing models frequently produce sentences that meet established dictionary-quality criteria (Atkins and Rundell 2008), particularly for monosemous lemmas, thereby substantially reducing manual effort. Although performance is lower for polysemous nouns, generating multiple candidates enables efficient selection and editing. Overall, two thirds of the examples for polysemous lemmas are suitable for inclusion, demonstrating that LLMs can meaningfully reduce, though not eliminate, lexicographic workload. This mitigated usage of LLMs is now being made available in GernEdit (Henrich and Hinrichs 2010a), our lexicographer’s workbench for GermaNet. The results of our study show that LLMs do not mark the end of lexicography, as has been claimed by, among others, de Schryver (2023), but rather suggest a fruitful synergy between lexicographers and LLM-generated output.