The main goal of this proposed project is to develop a novel bilingual topic model, which explicitly models the word co-occurrence cross-lingual in document-aligned comparable data using a novel merging and shuffling strategy, called CL-BTM. Given a document-aligned multilingual corpus, CL-BTM can be employed to extract latent cross-lingual topics that optimally describe the observed data […]
Read More