Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
M. K. Arabov, Svetlana Sergeevna Khaybullina, Dinara M. Naumetova
Kazan Federal University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
This paper addresses the end-to-end task of automatic geospatial question answering over multilingual toponymic data. An original bilingual (Russian–Tatar) dataset of toponyms of the Republic of Tatarstan is introduced, comprising 9,688 structured records with detailed linguistic, etymological, administrative, and coordinate information (93.1% of objects are georeferenced). Based on this dataset, a specialized question-answering corpus of approximately 39,000 “question–context–extractable answer” triples is constructed, with guaranteed answer localization within the text. To solve the task, an architecture combining two key components is proposed: a hybrid retriever that integrates dense semantic indexing using multilingual-e5-large with a geospatial filter and ranking (KD-trees, haversine distance), and an extractive reader based on fine-tuned transformer models. On 500 test queries, the hybrid search achieves Recall@1 = 0.988, Recall@5 = 1.000, and MRR = 0.994, statistically significantly outperforming both BM25 and the purely spatial method. Among the tested reader architectures (RuBERT, XLM-RoBERTa-large, T5-RUS), the best answer extraction quality is attained by the multilingual XLM-RoBERTa-large model: EM = 0.992, F1 = 0.994. A contrasting effect is observed: on raw outputs, RuBERT-based models fail to answer coordinate-related questions (F1 = 0), whereas XLM-RoBERTa-large achieves F1 = 0.984; however, simple post-processing completely eliminates the gaps in numerical values and restores RuBERT accuracy to 100%. This discrepancy is attributed to tokenization specifics and the composition of pre-training corpora. The created resources (dataset, QA corpus, trained model weights, and web demonstrator) are openly published on the Hugging Face platform. The obtained results can be directly applied in the development of geospatial question-answering services, geocoding systems, and digital humanities projects that benefit from etymological and multilingual place-name data.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
社会科学Geographic Information Systems Studies
Advanced Image and Video Retrieval Techniques · Human Mobility and Location-Based Analysis