Quantifiers-and-Negations-in-RE-Documents

This project was part of my work for a seminar at the Technical University of Munich (TUM) during my bachelor studies in 2019. The python project can be used to find quantifiers and negations in documents. It searches for problematic findings. Problematic findings are i.e. sentences that use specific combinations of quantifiers and negations that are ambiguous. This means there are multiple valid interpretations of the sentence. It can extract those and report them.

Motivation:

You want to avoid ambiguous sentences as they can cause problems that are hard to find and possibly hard to fix. This is especially the case for technical specifications and similar use cases. In this project we compare two different approaches to finding ambiguous sentences:

String based search
NLP based search

We want to find out if the computational overhead of using NLP gives better results than standard string based search methods.

Features:

Detect quantifiers and negations in .xml or .txt documents
Search either by a string based search or by NLP based search (using Stanfords CoreNLP library [1])
Extract possibly ambiguous sentences
Compare string search results with NLP search results

Prerequisites:

Java 8 or higher
Python 3.6 or higher as project interpreter
Stanford Corenlp library: https://stanfordnlp.github.io/CoreNLP/download.html
Environment variable "CORENLP_HOME" set to where the CoreNLP library is stored

References:

[1] Christopher D.Manning, MihaiSurdeanu, JohnBauer, JennyFinkel, StevenJ.Bethard, and David McClosky. The Stanford CoreNLP natural language processing toolkit. In Association for Computational Linguistics (ACL) System Demonstrations, pages 55–60, 2014.

Quantifiers and Negations in RE Documents

Related tags

Overview

Quantifiers-and-Negations-in-RE-Documents

Owner

Nicolas Ruscher

Original implementation of the pooling method introduced in "Speaker embeddings by modeling channel-wise correlations"

基于“Seq2Seq+前缀树”的知识图谱问答

Neural Lexicon Reader: Reduce Pronunciation Errors in End-to-end TTS by Leveraging External Textual Knowledge

Autoregressive Entity Retrieval

BERT-based Financial Question Answering System

A natural language modeling framework based on PyTorch

New Modeling The Background CodeBase

Extract Keywords from sentence or Replace keywords in sentences.

Trained T5 and T5-large model for creating keywords from text

source code for paper: WhiteningBERT: An Easy Unsupervised Sentence Embedding Approach.

A Paper List for Speech Translation

Collection of scripts to pinpoint obfuscated code

🌐 Translation microservice powered by AI

Multiple implementations for abstractive text summurization , using google colab

Ptorch NLU, a Chinese text classification and sequence annotation toolkit, supports multi class and multi label classification tasks of Chinese long text and short text, and supports sequence annotation tasks such as Chinese named entity recognition, part of speech tagging and word segmentation.

Chatbot with Pytorch, Python & Nextjs

The (extremely) naive sentiment classification function based on NBSVM trained on wisesight_sentiment

A python package for deep multilingual punctuation prediction.

NLP, Machine learning

Gold standard corpus annotated with verb-preverb connections for Hungarian.