General tricks that may help you find bad, or noisy, labels in your dataset

Last update: Dec 26, 2022

Related tags

Overview

doubtlab

A lab for bad labels.

Warning still in progress.

This repository contains general tricks that may help you find bad, or noisy, labels in your dataset. The hope is that this repository makes it easier for folks to quickly check their own datasets before they invest too much time and compute on gridsearch.

Install

You can install the tool via pip.

python -m pip install doubtlab

Quickstart

Doubtlab allows you to define "reasons" for a row of data to deserve another look. These reasons can form a pipeline which can be used to retreive a sorted list of examples worth checking again.

from doubtlab import DoubtLab
from doubtlab.reasons import ProbaReason, WrongPredictionReason

# Let's say we have some model already
model.fit(X, y)

# Next we can the reasons for doubt. In this case we're saying
# that examples deserve another look if the associated proba values
# are low or if the model output doesn't match the associated label.
reasons = {
    'proba': ProbaReason(model=model),
    'wrong_pred': WrongPredictionReason(model=model)
}

# Pass these reasons to a doubtlab instance.
doubt = DoubtLab(**reasons)

# Get the predicates, or reasoning, behind the order
predicates = doubt.get_predicates(X, y)
# Get the ordered indices of examples worth checking again
indices = doubt.get_indices(X, y)
# Get the (X, y) candidates worth checking again
X_check, y_check = doubt.get_candidates(X, y)

Features

The library implemented many "reaons" for doubt.

ProbaReason: assign doubt when a models' confidence-values are low
RandomReason: assign doubt randomly, just for sure
LongConfidenceReason: assign doubt when a wrong class gains too much confidence
ShortConfidenceReason: assign doubt when the correct class gains too little confidence
DisagreeReason: assign doubt when two models disagree on a prediction
CleanLabReason: assign doubt according to cleanlab

Related Projects

The cleanlab project was an inspiration for this one. They have a great heuristic for bad label detection but I wanted to have a library that implements many. Be sure to check out their work on the labelerrors.com project.
My employer, Rasa, has always had a focus on data quality. Some of that attitude is bound to have seeped in here. Be sure to check out Rasa X if you're working on virtual assistants.

Comments

`QuantileDifferenceReason` and `StandardDeviationReason`
Hey! I was thinking if it would make sense to add two more reasons for regressions tasks, namely something like HighLeveragePointReason and HighStudentizedResidualReason.

Citing Wikipedia:

Leverage is a measure of how far away the independent variable values of an observation are from those of the other observations. High-leverage points, if any, are outliers with respect to the independent variables (link)

A studentized residual is the quotient resulting from the division of a residual by an estimate of its standard deviation. [...] This is an important technique in the detection of outliers. (link)
opened by FBruzzesi 31
Doubt Reason Based on Entropy

If a machine learning model is very "confident" then the proba scores will have low entropy. The most uncertain outcome is a uniform distribution which would contain high entropy. Therefore, it could be sensible to add entropy as a reason for doubt.

opened by koaning 10
Add staticmethods to reasons to prevent re-compute.
I really like the current design with reasons just being function calls.

However, when working with large datasets or in use cases where you already have the predictions of a model, I wonder if you have thought about letting users to pass either a sklearn model or the pre-computed probas (for those Reasons where it make sense). For threshold-based reasons and large datasets this could save some time and compute, allow for faster iteration, and it would open up the possibility of using other models beyond sklearn.

I understand that the design wouldn't be as clean as it is right now, might cause miss-alignments if users don't send the correct shapes/positions, but I wonder if you have considered this (or any other way to pass pre-computed predictions).

Just to illustrate what I mean (sorry about the dirty-pseudo code):

class ProbaReason: def __init__(self, model=None, probas=None, max_proba=0.55): if not model or probas: print("You should at least pass a model or probas") self.model = model self.probas = probas self.max_proba = max_proba def __call__(self, X, y=None): probas = probas if self.probas else self.model.predict_proba(X) result = probas.max(axis=1) <= self.max_proba return result.astype(np.float16)
opened by dvsrepo 9
"Fair" Sorting

Suppose there are 5 reasons for doubt, 4 of which overlap a lot. Then we may end up in a situation where we ignore a reason. That could be bad ... maybe it's worth exploring voting systems a bit to figure out alternative sorting methods.

opened by koaning 7
Add example to docs that shows lambda X, y: y.isna()

Hey! First of all: this is a very cool project ;) I have been thinking about potential new "reasons" to doubt and I personally often look into predictions generated by a model whenever the data instance had missing values (and part of the model-pipeline imputes them)... So I wonder if it would be useful to have a FillNaNReason (or something similar) based, for example in the MissingIndicator transformer.

opened by juanitorduz 4
added conda-install-option and badges to readme
This closes #14: doubtlab can now be installed with conda from conda-forge channel.

[x] Created conda-forge/doubtlab-feedstock to make doubtlab available on conda-forge channel.

[x] Added conda install option to readme.

[x] Added the following badges to readme.
opened by sugatoray 4
Added a LICENSE
Hi @koaning,

I am assuming MIT License is okay for this repository. If you think otherwise, please feel free to make changes in the PR accordingly.

[x] Added an MIT License

[x] ~~Added a Citation file~~ Removed the citation file and updated the name of the PR. - ~~If you have an orcid, please consider adding it to the citation.cff file.~~
opened by sugatoray 4
Add a conda installation option using conda-forge channel

I have already started this one. Will push a PR once the conda installation option is available.

See: Adding doubtlab from PyPI to conda-forge channel.

@koaning As the primary maintainer of this repo, would you like to be listed as one of the maintainers of doubtlab on conda-forge channel? Please let me know, I will add your name as another maintainer of conda-forge/doubtlab-feedstock, once it is accepted.

opened by sugatoray 3
Doubt about MarginConfidenceReason :-)

Hi Vincent,

Nice library! As mentioned a while ago on Twitter I'm doing a review to understand and compare different approaches to find label errors.

I'm playing with the AG News dataset, which we know it contains a lot of errors from our own previous experiments with Rubrix (using the training loss and using cleanlab).

While playing with the different reasons, I'm having difficulties to understand the reasoning behind the MarginConfidenceReason. As far as I can tell, if the model is doubting the margin between the top two predicted labels should be small, and that could point to an ambiguous example and/or a label error. If I read the code and description correctly, MarginConfidenceReason is doing the opposite, so I'd love to know the reasoning behind this to make sure I'm not missing something.

For context, using the MarginConfidenceReason with the AG News training set yields almost the entire dataset (117788 examples for the default threshold of 0.2, and 112995 for threshold=0.5). I guess this could become useful when there's overlap with other reasons, but I want to make sure about the reasoning :-).

opened by dvsrepo 2
updated docs: installation and badges
Updated docs:

[x] updated installation (with conda)

[x] ~~added badges from readme~~

@koaning I am not sure if you would prefer to include the badges in the docs (website). If you don't, please feel free to remove them.

UPDATE: removed badges from the docs (docs/index.md).
opened by sugatoray 2
Issue with cleanlab upgrading to v2

Issue

Environment details

Temporary fix

pip install "doubtlab==1.0.0"

More permanent fix

Pin doubtlab dependency to "doubtlab<2.0.0"

More more permanent fix

They've made some changes to their API

Let me know if you'd like me to make a PR

Thanks for a great package @koaning 😄

opened by duarteocarmo 1
Consider a fairlearn demo.

When two models disagree something interesting might be happening. But that'll only happen if you have two models that are actually different.

What if you have one model that's better at accuracy and another one that's better at fairness.

Maybe these labels deserve more attention too.

opened by koaning 0
Assign Doubt for Dissimilarity from Labelled Set

Suppose that y can contain NaN values if they aren't labeled. In that case, we may want to favor a subset of these NaN values. In particular: if they differ substantially from the already labeled datapoints.

The idea here is that we may be able to sample more diverse datapoints.

opened by koaning 10
Does it make sense to add an ensemble for spaCy?

This seems to be a like-able method of dealing with text outside the realm of scikit-learn. But I prefer to delay this feature until I really understand the use-case. For anything related to entities we cannot use sklearn, but tags/classes should work fine as-is.

opened by koaning 1

Releases(0.2.4)

0.2.4(Nov 25, 2022)

Support added for false-positive, false-negative reason via WrongPredictionReason.
Source code(tar.gz)
Source code(zip)
0.2.3(Apr 22, 2022)

Supports CleanLab v2.
Source code(tar.gz)
Source code(zip)
0.2.2(Apr 14, 2022)

Sped up confidence-based reasons by vectorization. Added more priority on proba based methods on docs.
Source code(tar.gz)
Source code(zip)
0.1.5(Dec 21, 2021)

Added a StandardizedErrorReason https://github.com/koaning/doubtlab/pull/29. Thanks @FBruzzesi!
Source code(tar.gz)
Source code(zip)
0.1.4(Dec 15, 2021)

Added a reason based on entropy and from_proba-style methods.
Source code(tar.gz)
Source code(zip)
0.1.3(Dec 7, 2021)

Fixed a bug related to the margin reason. Details found here: https://github.com/koaning/doubtlab/pull/22
Source code(tar.gz)
Source code(zip)
0.1.2(Nov 23, 2021)

Added the MarginConfidenceReason.
Source code(tar.gz)
Source code(zip)
0.1.0(Nov 9, 2021)

Source code(tar.gz)
Source code(zip)

Owner

vincent d warmerdam

Solving problems involving data. Mostly NLP these days. AskMeAnything[tm].

GitHub Repository

Package pyVHR is a comprehensive framework for studying methods of pulse rate estimation relying on remote photoplethysmography (rPPG)

Package pyVHR (short for Python framework for Virtual Heart Rate) is a comprehensive framework for studying methods of pulse rate estimation relying on remote photoplethysmography (rPPG)

261 Jan 03, 2023

SQL centered, docker process running game

REQUIREMENTS Linux Docker Python/bash set up image "docker build -t game ." create db container "run my_whatever/game_docker/pdb create" # creating po

1 Jan 11, 2022

通过简单的卷积神经网络直接预测出验证码图片中滑块的位置

使用说明 1. 在本地测试运行python3 prdict_one.py即可，默认需要预测的图片路径位于testImg文件夹下的test1.png 运行python3 predict_folder.py预测testImg下的所有图片 2. 部署到服务器运行python3 run_a_server

12 Mar 08, 2022

Web3 Solidity Connector

With this project, you can compile your sol files and create new transactions including creating contract and calling the state changer functions. You can integrate integrate your sol files with Pyth

3 Oct 09, 2022

A collection of tips for using MISP.

MISP Tip of the Week A collection of tips for using MISP. Published via BelgoMISP (todo) and this repository. Available in MD and JSON. Do you want to

52 Jan 07, 2023

User management system (UMS), has the primary purpose of connecting to an Active Directory (AD)

💿 Sistema de Gerenciamento de Usuário (SGU) 📚 Sobre o projeto Sistema de gerenciamento de usuários (SGU), tem o objetivo primário de se conectar a u

2 Feb 25, 2022

Chess bot can play automatically as white or black on lichess.com, chess.com and any website using drag and drop to move pieces

Chessbot "Why create another chessbot ?" The explanation is simple : I did not find a free bot I liked online : all the bots I saw on internet are par

2 Nov 11, 2021

Types for the Rasterio package

types-rasterio Types for the rasterio package A work in progress Install Not yet published to PyPI pip install types-rasterio These type definitions

7 Sep 10, 2021

This code makes the logs provided by Fiddler proxy of the Google Analytics events coming from iOS more readable.

GA-beautifier-iOS This code makes the logs provided by Fiddler proxy of the Google Analytics events coming from iOS more readable. To run it, create a

3 Feb 02, 2022

Script to calculate the italian fiscal code of a person.

fiscal_code Hi! This is my first public repository, so please be kind if it is not well formatted or it contains errors. I started learning Python abo

1 Nov 20, 2021

IG Trading Algos and Scripts in Python

IG_Trading_Algo_Scripts_Python IG Trading Algos and Scripts in Python This project is a collection of my work over 2 years building IG Trading Algorit

191 Oct 11, 2022

Banking management project using Tkinter GUI in python.

Bank-Management Banking management project using Tkinter GUI in python. Packages required Tkinter - Tkinter is the standard GUI library for Python. sq

7 Jul 03, 2022

A set of tools for ripping music from Konami mobile games

Konami Mobile Ripping Toolset A set of tools for ripping music from Konami mobile games Contents nigger.py for niggering konami's website, ripping all

5 Oct 20, 2022

Tool for running a high throughput data ingestion/transformation workload with MongoDB

Mongo Mangler The mongo-mangler tool is a lightweight Python utility, which you can run from a low-powered machine to execute a high throughput data i

9 Jan 02, 2023

Python library for generating CycloneDX SBOMs

Python Library for generating CycloneDX This CycloneDX module for Python can generate valid CycloneDX bill-of-material document containing an aggregat

31 Dec 16, 2022

Checks for Vaccine Availability at your district and notifies you using E-mail, subscribe to our website.

Vaccine Availability Notifier Project Description Checks for Vaccine Availability at your district and notifies you using E-mail every 10 mins. Kindly

19 Jun 03, 2021

Fixes your Microphone Level to one specific value.

MicLeveler Fixes your Microphone Level to one specific value. Intention A friend of mine has the problem that some programs are setting his microphone

2 Oct 14, 2021

Python Classes Without Boilerplate

attrs is the Python package that will bring back the joy of writing classes by relieving you from the drudgery of implementing object protocols (aka d

4.6k Jan 02, 2023

dbt adapter for Firebolt

dbt-firebolt dbt adapter for Firebolt dbt-firebolt supports dbt 0.21 and newer Installation First, download the JDBC driver and place it wherever you'

23 Dec 14, 2022

An example of Connecting a MySQL Database with Python Code

An example of Connecting And Query Data a MySQL Database with Python Code And How to install Table of contents General info Technologies Setup General

1 Nov 23, 2021