Follow

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Subscribe

Award for scientific paper on training multilingual LLMs

image

A scientific paper on “Mitigating Catastrophic Forgetting in Multilingual Continual Pretraining: Lessons from EU Institutional LLMs” has been awarded Best Paper at the EuroHPC User Days (organised by the European High Performance Computing Joint Undertaking).  

The paper was drafted by a team of colleagues from DG Translation, the Directorate-General for Communications Networks, Content and Technology and the Joint Research Centre of the European Commission. It was evaluated by a peer review committee and selected for its scientific excellence and best practice in using supercomputer resources.  

Research findings 

Third-party large language models (LLMs) tend to lack the language coverage required by the EU institutions, which is why the European Commission built its own EU Institutional LLM.  

When an LLM acquires new languages, it loses fluency in some of the ones it has previously learned. To mitigate this effect, the research team behind the award-winning paper ran various practical tests.   

They measured the LLM’s general reasoning and language understanding, using standard benchmarks and sentence completion tasks for EU formal language. While it was not possible to fully avoid catastrophic forgetting, the researchers were able to reduce it through a combination of learning rate, data and training process. 

The model was also tested for its resistance to profanity – or toxicity – in all 24 EU official languages. It was concluded that training the model on EU data didn’t affect toxicity resistance.  In summary, the paper offers practical lessons for building high-quality multilingual LLMs better suited to European public-sector needs.  

The EU Institutional LLM is available to EU public administrations, small businesses, academia and non-governmental organisations, and is used to power some of the AI translation and language tools available under the Digital Europe programme. It is enhanced with high-quality multilingual text from the Commission’s own translation database, and is trained and pretrained using EuroHPC supercomputers.

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
EIC CoC Label