Multilingual Information Retrieval for Semantic Search, Retrieval-Augmented Generation, and Machine Translation of Tibetan

Date:

I co-chaired the panel “Advances in Tibetan NLP” with Marieke Meelen at the 17th Seminar of the International Association for Tibetan Studies (IATS 2026) and presented this paper in it. The panel also featured Elie Roux and Eric Werner (Buddhist Digital Resource Center) on Tibetan OCR development and Tashi Tsering and Tenzin Kaldan on LLM-assisted generation of Tibetan Wikipedia articles.

With the ongoing digitization of Tibetan textual material, ever-larger datasets are becoming available, which present a great opportunity for scholars when it comes to searching for relevant information. Traditional keyword-based search methods can fall short when it comes to lexical ambiguity, orthographic variations, and the inability to retrieve conceptually related passages within the same language and across language boundaries. While general-purpose LLMs have been advancing quickly in recent years, their application to specialized scholarly domains still requires robust mechanisms to ground them in reliable, curated, textual source data. This presentation addresses how information retrieval (IR) systems can be tailored for the specific needs of the Tibetan Studies community. I show how sparse IR systems such as BM25, and dense methods such as deep neural embeddings based on the Gemma2 MITRA semantic embedding model, can be used on Tibetan material, how different retrieval methods can be combined with reranking to achieve optimal results, and how downstream applications such as machine translation, semantic search for philological use, and retrieval-augmented generation can be realized with advanced IR methods. A special focus is on multilingual retrieval settings, where queries in English or other modern languages are used to retrieve results from Classical Tibetan texts, and on IR between classical languages, i.e., from Sanskrit queries to Tibetan results, or between Buddhist Chinese and Tibetan.