A TECHNIQUE FOR BRIDGING LANGUAGE BARRIERS TO DEVELOP AN URDU-ENGLISH INFORMATION RETRIEVAL SYSTEM
Keywords:
Cross-Lingual Information Retrieval, Urdu–English Retrieval, Machine Translation, Query Expansion, Cross-Lingual Embeddings, Natural Language Processing, Multilingual Information Access, Semantic Search.Abstract
Cross-Lingual Information Retrieval (CLIR) aims to enable users to access information written in a language different from that of their search query. In a CLIR environment, a query submitted in one language can retrieve relevant documents written in another language, thereby facilitating multilingual information access. This is achieved by developing retrieval systems capable of matching queries and documents across different languages. In recent years, CLIR has emerged as one of the most challenging research areas in Information Retrieval due to the linguistic diversity, semantic variations, and translation complexities involved. Urdu, the national language of Pakistan, is spoken by approximately 163 million people worldwide. Its rich morphological structure, complex syntax, and large speaker population make Urdu Information Retrieval an important and specialized area of research. Despite the growing demand for multilingual information access, efficient CLIR solutions for the Urdu–English language pair remain limited. To address this gap, this study proposes a hybrid query model designed to improve retrieval effectiveness and optimize performance in a cross-lingual retrieval environment. The CLIR framework proposed above allows retrieving relevant documents both in the cases when the search queries are posed either in Urdu or in English in both monolingual and cross-lingual environments such as Urdu to English and English to Urdu information retrieval tasks. The assessment of the system was carried out using the Text Retrieval Conference (TREC) approach, where standardized methodology exists for testing the performance of IR systems via use of document collections, query topics and relevance assessments. For experimental purposes, the UIR-21 dataset was chosen, which is a collection of Urdu news articles in the style of TREC. In total, 37,000 documents were used in the research. For the purpose of cross-lingual IR, Urdu documents were translated to English using the Google Translate API. The performance of the model was investigated by comparing the effectiveness of the process of retrieving the documents when the query-document matching takes place in both Urdu-URDU and translated English-English languages. The criteria for assessment included precision, recall, and F1-score. It should be noted that, according to the experimental results, English IRS was characterized by the highest precision, while Urdu IRS showed low precision.












