A Buddhist Chinese Parallel Corpus and AI Infrastructure from the Kokuyaku Issaikyō
Date:
As part of the SPReAD 1000 program, I gave a lightning talk and presented a poster at the DiHuCo Hub Symposium organized by the ROIS-DS Center for Open Data in the Humanities (CODH), National Institute of Informatics, at Hitotsubashi Hall in Tokyo. The project turns the Kokuyaku Issaikyō (『国訳一切経』) into AI training data, building a high-precision parallel corpus of more than 300,000 Buddhist Chinese–Japanese sentence pairs to strengthen Dharmamitra’s machine translation, search, and retrieval-augmented generation.
