Vector space based Information Retrieval System for Text Processing - Information retrieval

Last update: Jan 01, 2022

Related tags

Text Processing BITS-IR-PROJECT

Overview

Information Retrieval: Text Processing

Group 13

Sequence of operations

Install Requirements
Add given wikipedia files to the corpus directory.
Download glove.6B.100d.txt dataset (Ignore if already present) and place it in the project root directory.
Run construct_index.py
Run construct_index.py --zoned_index True
Run trim_embeddings.py
Run test_queries.py
Run test_queries.py --score_title True
Run test_queries.py --expand_query True

Installing Requirements:

   pip install -r requirements.txt

corpus

Contains the files to be indexed. Add files directly to this directory. Do not create subdirectories.
For this assignment, we have used the following files present in the AA folder of wikipedia files.
wiki_00
wiki_01
wiki_05
wiki_06
wiki_10
wiki_11
wiki_15
wiki_16
wiki_20
wiki_21
wiki_25
wiki_26
wiki_30
wiki_31

index_files

Contains the inverted indices constructed using construct_index.py.

construct_index.py

Constructs the inverted indices used for query evaluation.
Command Line Arguments:
--zoned_index: True if zoned indexing must be used. Set to False by default.

trim_embeddings.py

Trims the GloVe embeddings to contain terms only from corpus. Download the glove.6B.100d.txt dataset before running this file.

test_queries.py

Evaluates queries and displays retrieved documents.
Command Line Arguments:
--score_title: True if zoned index considered for evaluation. Set to False by default.
--expand_query: True if query expansion must be used. Set to False by default.

helper_module.py

Contains helper functions used by other files. Do not run this file.

document_list.txt

Contains the document ids and names used for evaluation.

Vector space based Information Retrieval System for Text Processing - Information retrieval

Related tags

Overview

Information Retrieval: Text Processing

Group 13

Sequence of operations

Installing Requirements:

corpus

index_files

construct_index.py

trim_embeddings.py

test_queries.py

helper_module.py

document_list.txt

Owner

🚩 A simple and clean python banner generator - Banners

This project aims to test check if your RegExp are being matched by grep.

Utility for Text Normalisation or Inverse Normalisation

Format Covid values to ASCII-Table (Only for Germany and Austria)

Migrates translations to the REDCap native Multi-Language Management system

Fuzz a language by mixing up only few words.

A Python library that provides an easy way to identify devices like mobile phones, tablets and their capabilities by parsing (browser) user agent strings.

Build a translation program similar to Google Translate with Python programming language and QT library

Umamusume story patcher with python

JSON and CSV data for Swahili dictionary with over 16600+ words

Hamming code generation, error detection & correction.

Chilean Digital Vaccination Pass Parser (CDVPP) parses digital vaccination passes from PDF files

Auto translate Localizable.strings for multiple languages in Xcode

Convert English text to IPA using the toPhonetic

Phone Number formatting for PlaySMS Platform - BulkSMS Platform

Getting git-style versioning working on RDFlib

AnnIE - Annotation Platform, tool for open information extraction annotations using text files.

A production-ready pipeline for text mining and subject indexing

Translate .sbv subtitle files

Answer some questions and get your brawler csvs ready!