Eurach Research

ELEXAI

Europäische Lexikografie-Infrastruktur für Künstliche Intelligenz

Das Projekt ELEXAI sieht sowohl eine wesentliche Weiterentwicklung bestehender KI-Basismodelle als auch einen umfassenden Ausbau der europäischen Lexikografie-Infrastruktur „ELEXIS“ vor, die aus dem gleichnamigen Horizon-2020-Projekt (2018–2022) hervorgegangen ist.

Dies stellt eine grundlegende Neugestaltung der bestehenden Infrastruktur dar, basierend auf den jüngsten Entwicklungen im Bereich der Künstlichen Intelligenz (KI), insbesondere im Zusammenhang mit dem Aufkommen großer Sprachmodelle (Large Language Models, LLMs). Die neue Infrastruktur wird lexikografische Bemühungen auf nationaler, regionaler und institutioneller Ebene zusammenführen und aktualisierte Technologien, Sprachdaten, Tools und Dienste bereitstellen. Letztere sind die für die Anreicherung transformerbasierter Sprachmodelle mit mehrsprachigem Wissen aus hochwertigen lexikografischen Ressourcen von entscheidender Bedeutung.

Maschinenlesbare Wissensrepräsentationen (z. B. Datenbanken und Graphen), die sprachwissenschaftlich fundiert sind und in LLMs integriert werden können, ermöglichen es, einen positiven Kreislauf anzukurbeln: Die Relevanz des integrierten sprachlichen Wissens wird anhand verbesserter Ergebnisse von LLM-Anwendungen überprüft, die ihrerseits zur Weiterentwicklung der Wissensrepräsentationen und Sprachdaten genutzt werden können. Um die Auswirkungen der Einbindung linguistischer Kenntnisse in LLMs zu überprüfen, ist als Ergebnis die Erstellung zuverlässiger Benchmarks und anderer Mittel zur Bewertung maschinengenerierter Ergebnisse vorgesehen. Somit wird die Gestaltung der Infrastruktur einen Beitrag zur Bewältigung noch ungelöster Aufgaben im Bereich des Verständnisses natürlicher Sprache leisten.

Der Aufbau einer neuen virtuellen lexikografischen Infrastruktur wird von einem breit aufgestellten und vielfältigen Konsortium durchgeführt, dem Partner aus allen relevanten Bereichen angehören: Lexikografie, Computerlinguistik und KI. Im Hinblick auf die langfristige Nachhaltigkeit stützt sich die Infrastruktur auf mehrere bedeutende Infrastrukturinitiativen: CLARIN und DARIAH, zwei ESFRI-Landmark-Infrastrukturen, sowie ALT-EDIC, die neue paneuropäische Initiative zur Entwicklung offener, massenhaft mehrsprachiger Sprachmodelle in Europa.

Background

Large pre-trained language models have transformed natural language processing, but they continue to suffer from a number of well-recognised limitations: hallucination, poor reasoning and weak interpretability. For Europe – home to a rich and diverse linguistic landscape – these weaknesses pose particular risks. Without access to accurate semantic information, LLMs often struggle with under-resourced languages and domain-specific terminology, amplifying digital inequalities and reducing people’s trust in AI outputs.

Lexicographic resources offer a way forward. Carefully curated resources contain the explicit, structured semantic knowledge that LLMs lack. Injecting this knowledge into LLMs has the potential to dramatically improve their reasoning, factual reliability, and cross-lingual performance. However, the interaction between LLMs, lexicographic databases and heterogeneous user communities (i.e. researchers, industry, policy makers, general public) remains fragmented: multiple LLM architectures, dozens of language formats, and diverse user needs prevent seamless integration.

Recent advances in agentic AI frameworks provide the missing connective tissue. Agentic systems have already demonstrated their ability to manage unpredictable, complex environments. By applying them to the European lexicographic ecosystem, ELEXAI can enable smooth interactions between LLMs, lexical resources, and human users.

Aims

Existing infrastructures must be technologically and organisationally upgraded if they are to remain relevant in the era of LLMs and agentic AI. The ELEXAI infrastructure will deliver integration of lexicographic data and LLMs through:

  1.     Lexicography-enhanced LLMs, specialised for European languages and lexicographic tasks.
  2.     LLM-enhanced lexical resources, where machine learning and lexicographic expertise mutually strengthen each other.
  3.     Accessible interfaces tailored to both expert lexicographers and general users.
  4.     Lexicography-based evaluation benchmarks, ensuring that progress is measurable and trustworthy.


Method
Addressing the LLM blackbox requires three key advances:

  1.     Deeper semantic description: an improved understanding of higher levels of linguistic meaning, as provided by lexicography, especially in contexts where humans must interpret machine-generated text (and vice versa).
  2.     Machine-readable knowledge resources: the development of knowledge bases and knowledge graphs that can be systematically injected into LLMs.
  3.     Robust evaluation frameworks: the creation of benchmarks and testing protocols that can assess machine-generated output in realistic NLU scenarios, ensuring that results are valid and scientifically grounded.


Therefore, ELEXAI will:

  1. Expand existing ELEXIS services through AI-enhanced language models. 
  2. Provide lexicographically informed knowledge representations. 
  3. Deliver European benchmarks and evaluation datasets.
  4. Build an agentic AI-powered ELEXAI platform. 
  5. Ensure access, standards, and interoperability. 
  6. Promote Open Science and responsible innovation in lexicography. 
  7. Secure the long-term sustainability and growth of ELEXAI. 


Eurac Research
Eurac Research's role within ELEXAI is to share existing terminological and corpus data focusing on the specific South Tyrolean language variety and on domain-related data, test LLM performance on multilingual terminological data (e,g. for definitionwriting) and support the development of sound benchmarks.

Discover

Verwandte Expertinnen und Experten, Publikationen und Forschungsthemen ansehen.

Science Shots Eurac Research Newsletter

Die besten Wissenschafts-Stories und Veranstaltungen des Monats.

Sprache wählen
Eurac Research logo

Eurac Research ist ein privates Forschungszentrum mit Sitz in Bozen, Südtirol. Unsere Forscherinnen und Forscher kommen aus allen Teilen der Welt und arbeiten in vielen verschiedenen Disziplinen. Gemeinsam widmen sie sich dem, was ihr Beruf und ihre Berufung ist – Zukunft zu gestalten.

No Woman No Panel

Was wir tun

Wo wir sind

Eurac Research Headquarter

Viale Druso 1

39100 Bolzano

Italy

Get directions

Eurac Research NOI Techpark

Via A.-Volta

39100 Bolzano

Italy

Get directions
ARBEITE MIT UNS

Except where otherwise noted, content on this site is licensed under a Creative Commons Attribution 4.0 International license.

Folge uns
  • Facebook icon
  • X icon
  • Instagram icon
  • Youtube icon
  • linkedin icon