Text Classification Using LSTM

Last update: Jan 03, 2023

Overview

Text-Classification-Using-LSTM

Ontology Classification-Using-LSTM

Introduction

Text classification is the task of assigning a set of predefined categories to free text. Text classifiers can be used to organize, structure, and categorize pretty much anything. For example, new articles can be organized by topics, support tickets can be organized by urgency, chat conversations can be organized by language, brand mentions can be organized by sentiment, and so on.

Technologies Used

1. IDE - Pycharm
2. LSTM - As a classification Deep learning Model
3. GPU - P-4000
4. Google Colab - Text Analysis
5. Flas- Fast API
6. Postman - API Tester
7. Gensim - Word2Vec embeddings

🔑 Prerequisites All the dependencies and required libraries are included in the file requirements.txt

  Python 3.6

Dataset

The DBpedia ontology classification dataset is constructed by picking 14 non-overlapping classes from DBpedia 2014. They are listed in classes.txt. From each of thse 14 ontology classes, we randomly choose 40,000 training samples and 5,000 testing samples. Therefore, the total size of the training dataset is 560,000 and testing dataset 70,000. The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3 columns in them, corresponding to class index (1 to 14), title and content. The title and content are escaped using double quotes ("), and any internal double quote is escaped by 2 double quotes (""). There are no new lines in title or content.

For Dataset Please click here

Process - Flow of This project

🚀 Installation of Text-Classification-Using-LSTM

Clone the repo

git clone https://github.com/KrishArul26/Text-Classification-DBpedia-ontology-classes-Using-LSTM.git

Change your directory to the cloned repo

cd Text-Classification-DBpedia-ontology-classes-Using-LSTM

Create a Python 3.6 version of virtual environment name 'lstm' and activate it

pip install virtualenv

virtualenv bert

lstm\Scripts\activate

Now, run the following command in your Terminal/Command Prompt to install the libraries required!!!

pip install -r requirements.txt

💡 Working

Type the following command:

python app.py

After that You will see the running IP adress just copy and paste into you browser and import or upload your speech then closk the predict button.

Implementations

In this section, contains the project directory, explanation of each python file presents in the directory.

1. Project Directory

Below picture illustrate the complete folder structure of this project.

2. preprocess.py

Below picture illustrate the preprocess.py file, It does the necessary text cleaning process such as removing punctuation, numbers, lemmatization. And it will create train_preprocessed, validation_preprocessed and test_preprocessed pickle files for the further analysis.

3. word_embedder_gensim.py

Below picture illustrate the word_embedder_gensim.py, After done with text pre-processing, this file will take those cleaned text as input and will be creating the Word2vec embedding for each word.

4. rnn_w2v.py

Below picture illustrate the rnn_w2v.py, After done with creating Word2vec for each word then those vectors will use as input for creating the LSTM model and Train the LSTM (RNN) model with body and Classes.

5. index.htmml

Below picture illustrate the index.html file, these files use to create the web frame for us.

6. main.py

Below picture illustrate the main.py, After evaluating the LSTM model, This files will create the Rest -API, To that It will use FLASK frameworks and get the request from the customer or client then It will Post into the prediction files and Answer will be deliver over the web browser.

Text Classification Using LSTM

Related tags

Overview

Text-Classification-Using-LSTM

Ontology Classification-Using-LSTM

Introduction

Technologies Used

Dataset

Process - Flow of This project

🚀 Installation of Text-Classification-Using-LSTM

💡 Working

Implementations

In this section, contains the project directory, explanation of each python file presents in the directory.

1. Project Directory

Below picture illustrate the complete folder structure of this project.

2. preprocess.py

3. word_embedder_gensim.py

4. rnn_w2v.py

5. index.htmml

6. main.py

7. Testing Rest-API

Owner

KrishArul26

Fuzzy String Matching in Python

A natural language modeling framework based on PyTorch

:mag: Transformers at scale for question answering & neural search. Using NLP via a modular Retriever-Reader-Pipeline. Supporting DPR, Elasticsearch, HuggingFace's Modelhub...

Indobenchmark are collections of Natural Language Understanding (IndoNLU) and Natural Language Generation (IndoNLG)

A CRM department in a local bank works on classify their lost customers with their past datas. So they want predict with these method that average loss balance and passive duration for future.

Text-Summarization-using-NLP - Text Summarization using NLP to fetch BBC News Article and summarize its text and also it includes custom article Summarization

Takes a string and puts it through different languages in Google Translate a requested amount of times, returning nonsense.

Train 🤗-transformers model with Poutyne.

This library is testing the ethics of language models by using natural adversarial texts.

天池中药说明书实体识别挑战冠军方案；中文命名实体识别；NER; BERT-CRF & BERT-SPAN & BERT-MRC；Pytorch

PyTorch Implementation of the paper Single Image Texture Translation for Data Augmentation

SpikeX - SpaCy Pipes for Knowledge Extraction

ProtFeat is protein feature extraction tool that utilizes POSSUM and iFeature.

GPT-2 Model for Leetcode Questions in python

Official source for spanish Language Models and resources made @ BSC-TEMU within the "Plan de las Tecnologías del Lenguaje" (Plan-TL).

Partially offline multi-language translator built upon Huggingface transformers.

AI_Assistant - This is a Python based Voice Assistant.

Installation, test and evaluation of Scribosermo speech-to-text engine

OCR을 이용하여 인원수를 인식 후 줌을 Kill 해줍니다

This is the writeup of all the challenges from Advent-of-cyber-2019 of TryHackMe