Benchmark zur Evaluierung von KI-generierten Rechtstexten im Südtiroler Deutschen
Das Projekt ist ein interdisziplinäres Dissertationsvorhaben, das verschiedene Forschungsbereiche zusammenführt: (deutsche) Sprachvariation, Terminologie, maschinelle Übersetzung und generative KI für bestimmte Sprachvarietäten. Ziel des Projekt ist es zu erforschen, inwieweit die neuesten Fortschritte in der KI die Redaktion und das Übersetzen in verschiedenen Sprachvarietäten und Sprachkombinationen, in denen Englisch fehlt, unterstützen können. Die Fallstudie befasst sich mit der in Südtirol verwendeten Standardvarietät des Deutschen in Kombination mit dem Italienischen.
Background
Natural Language Generation evaluation assesses GenAI performance using benchmarks that assess model capabilities via predefined tasks and agreed-upon metrics. However, currently used benchmarks have major design flaws: lack of real-world utility, data contamination, preference for fluency of text over accuracy of information and insufficient validation.
For South Tyrol, this inadequacy is compounded by the neglect around local needs of terminological adaptation, including the system-bound nature of legal language. Consequently, the development of any AI text generation application targeted to South Tyrolean users is hindered by the absence of a fair, representative benchmark and proper evaluation metrics. Currently, developers face the lack of good-quality datasets as the major challenge, as documented for other under-resourced languages and varieties.
Research gaps
The PhD addresses different research gaps:
- the scarce insights into how minor varieties of (European) languages risk being overshadowed by major varieties in AI tools due to lack of data for training and customisation;
- the well-known performance gap of generative AI facing text production/editing in non-English languages and of machine translation systems in non-English language combinations, while both are greatly needed by (multilingual) public and private organisations;
- the specific challenges of applying AI to the legal domain, which has system-bound terminology that varies within the same language (e.g. German in Austria, Germany, Switzerland and South Tyrol) and where terminology errors may lead to dramatic consequences.
The research question can be summed up as: To what extent can we exploit terminological data to develop an evaluation benchmark that assesses the adequacy of mono- and multilingual AI-generated texts into a minority language variety?
Objectives
The objective of the PhD is to create:
- A high-quality, highly specialized dataset of South Tyrolean legal language, achieved by using terminological data published in bistro.
- An evaluation benchmark for multilingual (e.g. translation) and monolingual (e.g. legal text drafting) assessment of AI-generated texts in legal South Tyrolean German.
- An automated evaluation metric, achieved by fine-tuning a pre-trained multilingual model on sentences with South Tyrolean terminology.
Methods
bistro contains more than 13,000 terminological entries related to various subdomains of law in Italian and German. In the German part, it considers the language varieties of South Tyrol, Austria, Germany, Switzerland, EU and international law. For each Italian term bistro does not provide a single, generic equivalent but all equivalent terms in the German, Austrian, Swiss, etc. legal systems. It also contains the terminology in German approved by the South Tyrolean Terminology Commission. For both Italian and the South Tyrolean German variant, it features higher-level information such as concept IDs and terms (designations), and lower level fields such as context of use (a usage example featuring the term).
The methodology for dataset creation starts from the extraction of the contexts of use from the XML-formatted bistro data to the ends of their later augmentation with semi-synthetic methods. This is achieved by the creation of:
- Semi-synthetic contexts: substituting deterministically South Tyrolean terms and phraseologies with equivalents from other German varieties, attained with the use of LLMs or grammar checkers. This will create authentic contexts of use as positive exemplars, and term-replaced contexts of use as negative exemplars;
- Parallel segments: because contexts of use in the two languages come from different sources, parallel segments would be created by machine translating positive exemplars into Italian with the constrained use of the correct equivalent Italian term, ensuring the South Tyrolean German term is always present;
- Synthetic contexts of use: additional sentences would be created by having an LLM generate instances with the constraint of including the term of interest.
As for the fine-tuning of metrics, the proposal aims to explore two approaches relative to the types of automatic machine translation evaluation metrics: the training of an embedding –based metric and of a learned regression metric.
The first approach – applicable to mono- and bi-lingual evaluation – would rely on fine-tuning pre-trained embeddings of a multilingual model. It would employ the SIMCSE technique (Tang et al. 2024) to fine-tune a metric on a low-resource language. There is a contrastive learning objective: given a positive exemplar, a negative exemplar and an anchor, this technique reduces the latent vector distance between the representation of the anchor and the positive exemplar, while maximizing the distance with the negative exemplar.
While the Italian parallel data previously created allow to gain the boost of multilinguality during training, this approach has the advantage of being viable also without parallel data by using unsupervised learning, making it a feasible option for purely monolingual text evaluation too. This solution would measure the semantic similarity between the reference and the candidate translation, or between the source and the candidate.
The second approach, exclusively applicable to bilingual data, would consist of further fine-tuning existing learned metrics onto parallel segments associated with a translation score. An open-source platform for facilitated training will be used.
A few rules of thumb should be followed when training a regression model. Ideally, the score range represented should cut across the whole score scale to avoid training the regression model to cluster the output within very narrow score intervals. For example, if only good translations are provided, the model would learn to output high scores regardless of the translation quality. In this regard, the creation of dataset sentences with erroneous terminology is extremely valuable. Secondly, because Comet models have already been fine-tuned on most common types of error, the advantage is to rely on this basis to inject task-specific errors – such as terminology variation – at a reduced computational expense.
Results
The PhD will achieve innovation by developing: 1) an automatic evaluation model trained on terminological data; 2) neural evaluation metrics with non-English and minority language variety combinations; 3) evaluation metrics specialised in terminology, one of the most critical aspects in legal language; 4) a pipeline for repurposing terminology databases to train AI applications. It contributes to scientific advancement by providing a novel dataset specialised in terminology variation across legal German, a dataset benchmark for AI applications and an evaluation metric for mono-/multilingual South Tyrolean legal texts.
The results will benefit all public and private actors involved in legal drafting and translation (public administration, courts, law firms, companies) as well as the Terminology Commission whose work might see a boosted dissemination through the tools. The project would also greatly contribute to a planned international ISO standard on “Terminology management and artificial intelligence”.
Potential follow-up products are a real-time AI writing assistant for South Tyrolean legal texts (from laws to contracts) and a customised MT tool for South Tyrolean German.
- Project duration: -
- Project status:
- Funding: Internal funding EURAC (Project)
- Institute: Institut für Angewandte Sprachforschung
Discover
Verwandte Expertinnen und Experten, Publikationen und Forschungsthemen ansehen.




