Package for controllable summarization

Last update: Dec 07, 2022

Related tags

Overview

summarizers

summarizers is package for controllable summarization based CTRLsum.
currently, we only supports English. It doesn't work in other languages.

Installation

pip install summarizers

Usage

1. Create Summarizers

First at all, create summarizers obejct to summarize your own article.

>>> from summarizers import Summarizers
>>> summ = Summarizers()

You can select type of source article between [normal, paper, patent].
If you don't input any parameter, default type is normal.

>>> from summarizers import Summarizers
>>> summ = Summarizers('normal')  # <-- default.
>>> summ = Summarizers('paper')
>>> summ = Summarizers('patent')

If you want GPU acceleration, set param device='cuda'.

>>> from summarizers import Summarizers
>>> summ = Summarizers('normal', device='cuda')

2. Basic Summarization

If you inputted source article, basic summariztion is conducted.

>>> contents = """
Tunip is the Octonauts' head cook and gardener. 
He is a Vegimal, a half-animal, half-vegetable creature capable of breathing on land as well as underwater. 
Tunip is very childish and innocent, always wanting to help the Octonauts in any way he can. 
He is the smallest main character in the Octonauts crew.
"""

>>> summ(contents)
'Tunip is a Vegimal, a half-animal, half-vegetable creature'

3. Query focused Summarization

If you want to input query together, Query focused summarization conducted.

>>> summ(contents, query="main character of Octonauts")
'Tunip is the smallest main character in the Octonauts crew.'

3. Abstractive QA (Auto Question Detection)

If you inputted question as query, Abstractive QA is conducted.

>>> summ(contents, query="What is Vegimal?")
'Half-animal, half-vegetable'

You can turn off this feature by setting param question_detection=False.

>>> summ(contents, query="SOME_QUERY", question_detection=False)

4. Prompt based Summarization

You can generate summary that begins with some sequence using param prompt.
It works like GPT-3's Prompt based generation. (but It doesn't work very well.)

>>> summ(contents, prompt="Q:Who is Tunip? A:")
"Q:Who is Tunip? A: Tunip is the Octonauts' head"

5. Query focused Summarization with Prompt

You can also input both query and prompt.
In this case, a query focus summary is generated that starts with a prompt.

>>> summ(contents, query="personality of Tunip", prompt="Tunip is very")
"Tunip is very childish and innocent, always wanting to help the Octonauts."

6. Options for Decoding Strategy

For generative models, decoding strategy is very important.
summarizers support variety of options for decoding strategy.

>>> summ(
...     contents=contents,
...     num_beams=10,
...     top_k=30,
...     top_p=0.85,
...     no_repeat_ngram_size=3,                  
... )

License

Copyright 2021 Hyunwoong Ko.

Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

Package for controllable summarization

Related tags

Overview

summarizers

Installation

Usage

1. Create Summarizers

2. Basic Summarization

3. Query focused Summarization

3. Abstractive QA (Auto Question Detection)

4. Prompt based Summarization

5. Query focused Summarization with Prompt

6. Options for Decoding Strategy

License

Owner

Hyunwoong Ko

Minimal GUI for accessing the Watson Text to Speech service.

A machine learning model for analyzing text for user sentiment and determine whether its a positive, neutral, or negative review.

GPT-2 Model for Leetcode Questions in python

Learning Spatio-Temporal Transformer for Visual Tracking

Active learning for text classification in Python

Yes it's true :broken_heart:

A collection of GNN-based fake news detection models.

Learn meanings behind words is a key element in NLP. This project concentrates on the disambiguation of preposition senses. Therefore, we train a bert-transformer model and surpass the state-of-the-art.

Implementation of some unbalanced loss like focal_loss, dice_loss, DSC Loss, GHM Loss et.al

Silero Models: pre-trained speech-to-text, text-to-speech models and benchmarks made embarrassingly simple

Adversarial Examples for Extreme Multilabel Text Classification

Natural Language Processing Specialization

KoBERT - Korean BERT pre-trained cased (KoBERT)

A CRM department in a local bank works on classify their lost customers with their past datas. So they want predict with these method that average loss balance and passive duration for future.

Implementation of ProteinBERT in Pytorch

AEC_DeepModel - Deep learning based acoustic echo cancellation baseline code

Generating new names based on trends in data using GPT2 (Transformer network)

Ecco is a python library for exploring and explaining Natural Language Processing models using interactive visualizations.

This repository contains the code for running the character-level Sandwich Transformers from our ACL 2020 paper on Improving Transformer Models by Reordering their Sublayers.

Residual2Vec: Debiasing graph embedding using random graphs