Cross Language Information Retrieval System Using Machine Learning Algorithms

Introduction :

Cross Language Information Retrieval (CLIR), is the process of retrieving relevant documents, where in the language of the given query is different from the language of the retrieved documents. CLIR allowed the users to search and access documents in the language different from the language of the search query. The process of CLIR system modeling consists of three basic approaches such as Query Translation, Document Translation and Hybrid approach. The basic techniques of CLIR in Query translation can be divided as; Dictionary based CLIR, Corpus based CLIR and Machine Translation based CLIR.

Significance of the Study :

CLIR provides a new paradigm for searching across a variety of languages. It plays a very vital role in a country like India having a wide variety of regional languages. In an increasingly globalized economy, the ability to find information in regional languages is becoming a necessity of time. The proposed CLIR system facilitates users to pose query in both Tamil and Malayalam languages. It plays a vital role in future because large volume of information stored in the web is in English. The cross-search engine can be used for multilingual countries; so that the people belonging to various languages community can retrieve documents in their native languages.

Aim of the study :

  • The aim of the research work is to investigate and design new framework to improve the performance of Cross Language Information Retrieval.
  • The work aims to investigate, design, and implement Cross Language Information Retrieval model using Machine Learning algorithm for Tamil- Malayalam.
  • The Study focus on developing effective mechanisms for translating the information needed from the query language to the documents language. The work also attempts to investigate all the relevant and related activity in this area of research.