Eurach Research

UniTermGPT

University terminology in German in the age of ChatGPT

Specialized translation and terminology work are impacted by the acceleration in AI breakthroughs, such as GPTs (generative pre-trained transformers). However, in the field of specialized translation, GPTs like ChatGPT currently do not take terminology variation sufficiently into account. This project focuses on higher education, as a domain characterized by terminological variation. Despite the introduction of the Bologna system, which aimed at the harmonization of EU higher education systems, there are still major terminological differences between countries and regions. This includes countries that have the same language in common, like the German-speaking countries. The overall objective of this project is to analyze how ChatGPT treats university terminology in different German language varieties (Austria, Germany and South Tyrol).

Objective: The objective of UniTermGPT is to analyze how ChatGPT treats university terminology of selected German language varieties in Europe (depending on the prompt). 

The specific objectives of UniTermGPT are: 1) to assess the language-variety specificity of terminology ‘translation’ when using ChatGPT on a subcorpus representative for the analyzed language varieties and text types; 2) to examine the influence of prompt engineering (incl. the integration of terminological resources through e.g. an API) on the ‘translation’ of language-variety-specific university terminology from English into German, and 3) to give recommendations for the use of ChatGPT when translating texts containing language-variety-specific terminology.

Innovative potential: UniTermGPT can set the stage for prompt engineering in the field of domain-specific, language-variety-specific terminology injection in translations generated with large language models. 

Research question: The research question addressed by UniTermGPT is: How can ChatGPT solve, if prompted accordingly, the issue of terminology injection in the translation of specialized language taking German language varieties into account? 

Method: The project combines corpus creation, terminological resource development and prompt engineering to study how large language models (LLMs) handle terminological variation in German university contexts. First, a corpus of university texts from Austria, Germany, Switzerland and South Tyrol is compiled and annotated with the help of domain experts, ensuring that regional and institutional varieties are systematically represented. From this corpus, terminology is extracted, harmonized with existing databases and compiled into a bilingual term list, including English and the different German varieties. Parallel to this, prompts for ChatGPT are designed, tested and refined to evaluate how effectively the model translates and manages specialized terminology. Expert annotators assess the quality of ChatGPT’s outputs, allowing systematic comparison between LLM-generated translations, curated resources and human expertise. The annotated datasets, terminological resources and prompt strategies are then analyzed and openly published to enable replication, reuse and further exploration by both researchers and practitioners. 


Discover

See related experts, publications, and research topics.

Search for " UniTermGPT University term…"
Explore " UniTermGPT" in Discover

Science Shots Eurac Research Newsletter

Get your monthly dose of our best science stories and upcoming events.

Choose language
Eurac Research logo

Eurac Research is a private research center based in Bolzano (South Tyrol) with researchers from a wide variety of scientific fields who come from all over the globe. Together, through scientific knowledge and research, they share the goal of shaping the future.

No Woman No Panel

What we do

Where we are

Eurac Research Headquarter

Viale Druso 1

39100 Bolzano

Italy

Get directions

Eurac Research NOI Techpark

Via A.-Volta

39100 Bolzano

Italy

Get directions
WORK WITH US

Except where otherwise noted, content on this site is licensed under a Creative Commons Attribution 4.0 International license.

Follow us
  • Facebook icon
  • X icon
  • Instagram icon
  • Youtube icon
  • linkedin icon