Facts and Language models

Welcome to the dedicated page for the RCF-funded Facts and Language models project, or FLaM for short!

The FLaM project is funded by an RCF fellowship grand awarded to Timothee Mickus, who serves as the PI for this project. The project will run between Sep 1st, 2026 and Aug 30, 2030. The project is part of broader research activities of the Language Technology group at the University of Helsinki, Helsinki-NLP.

Factuality is an area of concerns for large language models — they suffer from hallucinations, often fail to capture human ambiguity, and underestimate the variation inherent to natural language data. Yet, we also need to reconcile this with the observation that LLMs can produce factually correct information in many instances, and that they can and do memorize lengthy pieces of texts. The Facts and Language Models (FLaM) project will answer this focusing on the type of data they are exposed to: can we delineate which facts will not be properly portrayed by a model, given its training data? can we identify what data are responsible for specific LLM behaviors?

The work comprised in FLaM will converge onto practical, concrete use-cases; namely (i) generating dictionary definitions across languages and (ii) predicting when hallucinations will occur, rather than detecting that one occurred. While the former has practical relevance for minoritized language communities for which lexicographic work is usually deemed financially impractical, the latter has a broader impact in terms of providing a clearer scope of what topics a model can and cannot be trusted on.

FLaM builds upon a series of projects, some of which are still ongoing:

More broadly, the University of Helsinki fosters a number of initiatives that dovetail with the objectives of FLaM

News

We’re expecting to hire PhD students within the scope of this project. We’re explicitly avoiding quantitative indicators in assessing applications or research success. Stay tuned for more information.