
Language is messy, subjective, and influenced by the intentions of those who produce it. Yet it also captures how environmental change is experienced, discussed, and interpreted by society. By combining advances in artificial intelligence with rigorous evaluation and domain expertise, applied linguistics can contribute to transforming these observations into structured knowledge.
When people think about environmental data, they typically imagine weather stations, satellite images or labs measurements from highly technological sensors. However, traditional measurements often struggle to capture the complex and increasing interplay between social and environmental systems. This limitation becomes particularly evident during extreme events such as droughts, where climate and human pressures interact and lead to severe impacts. Within this context, natural language (the language people use when speaking or writing) can provide a valuable source of unconventional data. Local newspapers report crop losses or when and how municipalities announce water restrictions, or tourism operators describe falling lake levels, and environmental organizations warn about impacts on ecosystems. Together, these texts contain valuable information about how drought affects society and the environment.
But how can we access this kind of information for environmental research? Clearly, reading and making sense of thousands and thousands of articles mentioning drought in more or less explicit ways, is not efficient to analyse and correlate traditional data on droughts with the impacts mentioned in the texts. Also, language can be messy, subjective, and influenced by the intentions of those who produce it, it captures how environmental change is experienced, discussed, and interpreted by society, but it is rarely a neat source of information. Newspaper articles do not simply state that a drought caused a specific impact at a specific location and time. They discuss concerns about future water shortages, debate political responses, quote different stakeholders, or report impacts that may have multiple causes. A lake's low water level, for example, might be linked to drought, a concurrent heatwave, or long-term water management issues. Distinguishing between real impacts, potential future conditions, opinions, and speculation is therefore challenging not only for AI systems, but sometimes even for human readers.
For this reason, the combination of modern language technologies and sectorial drought assessment can help to automatically identify and extract relevant information in the text such as when a drought occurred, what sectors it impacted and where these impacts were felt. Thanks to their ability to interpret meaning and context, these systems can identify even complex relationships that would have been difficult to capture using earlier techniques or without extensive amounts of manually annotated training data.
While modern language models with their advanced performance in semantic reasoning offer unprecedented opportunities for analysing large collections of text, researchers must carefully assess how reliably these systems extract information and whether the resulting datasets are suitable for scientific use. At the same time, the growing computational costs of AI make it important to develop approaches that are not only accurate, but also efficient and sustainable.
In the DryAlps project,we investigate how far language can become environmental data. By analysing a large collection of Italian newspaper articles, we use language technologies to identify where drought impacts occurred, which sectors were affected, and when these impacts took place, with the aim of creating an openly available and regularly updated dataset that can help researchers study and monitor drought impacts across locations and time, complementing traditional environmental observations with information derived from human reporting.

Jennifer-Carmen Frey
Computational linguist with a background in language education, technology-enhanced learning and German linguistics. Fascinated by language in all its forms and uses and always on the search for the right amount of words needed.

Alessandra Pomella
Data scientist working with text data on climate impacts, especially from droughts. She combines expert knowledge with critical approaches to AI and digital technologies

Stefano Terzi
Environmental engineer expert in drought risk assessments. He works at the interplay between hydrology and social factors evaluating critical conditions and related impacts.
Discover
Making research visible.
Find the connected experts, projects, publications, and research areas behind this blog post.
This content is licensed under a Creative Commons Attribution 4.0 International license except for third-party materials or where otherwise noted.

