Pandas-based utility to calculate weighted means, medians, distributions, standard deviations, and more.

Last update: Dec 31, 2022

Related tags

Overview

weightedcalcs

weightedcalcs is a pandas-based Python library for calculating weighted means, medians, standard deviations, and more.

Features

Plays well with pandas.
Support for weighted means, medians, quantiles, standard deviations, and distributions.
Support for grouped calculations, using DataFrameGroupBy objects.
Raises an error when your data contains null-values.
Full test coverage.

Installation

pip install weightedcalcs

Usage

Getting started

Every weighted calculation in weightedcalcs begins with an instance of the weightedcalcs.Calculator class. Calculator takes one argument: the name of your weighting variable. So if you're analyzing a survey where the weighting variable is called "resp_weight", you'd do this:

import weightedcalcs as wc
calc = wc.Calculator("resp_weight")

Types of calculations

Currently, weightedcalcs.Calculator supports the following calculations:

calc.mean(my_data, value_var): The weighted arithmetic average of value_var.
calc.quantile(my_data, value_var, q): The weighted quantile of value_var, where q is between 0 and 1.
calc.median(my_data, value_var): The weighted median of value_var, equivalent to .quantile(...) where q=0.5.
calc.std(my_data, value_var): The weighted standard deviation of value_var.
calc.distribution(my_data, value_var): The weighted proportions of value_var, interpreting value_var as categories.
calc.count(my_data): The weighted count of all observations, i.e., the total weight.
calc.sum(my_data, value_var): The weighted sum of value_var.

The obj parameter above should one of the following:

A pandas DataFrame object
A pandas DataFrame.groupby object
A plain Python dictionary where the keys are column names and the values are equal-length lists.

Basic example

Below is a basic example of using weightedcalcs to find what percentage of Wyoming residents are married, divorced, et cetera:

import pandas as pd
import weightedcalcs as wc

# Load the 2015 American Community Survey person-level responses for Wyoming
responses = pd.read_csv("examples/data/acs-2015-pums-wy-simple.csv")

# `PWGTP` is the weighting variable used in the ACS's person-level data
calc = wc.Calculator("PWGTP")

# Get the distribution of marriage-status responses
calc.distribution(responses, "marriage_status").round(3).sort_values(ascending=False)

# -- Output --
# marriage_status
# Married                                0.425
# Never married or under 15 years old    0.421
# Divorced                               0.097
# Widowed                                0.046
# Separated                              0.012
# Name: PWGTP, dtype: float64

More examples

See this notebook to see examples of other calculations, including grouped calculations.

Max Ghenis has created a version of the example notebook that can be run directly in your browser, via Google Colab.

Pandas-based utility to calculate weighted means, medians, distributions, standard deviations, and more.

Related tags

Overview

weightedcalcs

Features

Installation

Usage

Getting started

Types of calculations

Basic example

More examples

Weightedcalcs in the wild

Other Python weighted-calculation libraries

Owner

Jeremy Singer-Vine

International Space Station data with Python research 🌎

BIGDATA SIMULATION ONE PIECE WORLD CENSUS

A Python package for Bayesian forecasting with object-oriented design and probabilistic models under the hood.

Projects that implement various aspects of Data Engineering.

Elasticsearch tool for easily collecting and batch inserting Python data and pandas DataFrames

This repository contains some analysis of possible nerdle answers

MidTerm Project for the Data Analysis FT Bootcamp, Adam Tycner and Florent ZAHOUI

PyNHD is a part of HyRiver software stack that is designed to aid in watershed analysis through web services.

We're Team Arson and we're using the power of predictive modeling to combat wildfires.

Reading streams of Twitter data, save them to Kafka, then process with Kafka Stream API and Spark Streaming

Advanced Pandas Vault — Utilities, Functions and Snippets (by @firmai).

A real-time financial data streaming pipeline and visualization platform using Apache Kafka, Cassandra, and Bokeh.

Business Intelligence (BI) in Python, OLAP

Project under the certification "Data Analysis with Python" on FreeCodeCamp

ELFXtract is an automated analysis tool used for enumerating ELF binaries

The micro-framework to create dataframes from functions.

ETL pipeline on movie data using Python and postgreSQL

🧪 Panel-Chemistry - exploratory data analysis and build powerful data and viz tools within the domain of Chemistry using Python and HoloViz Panel.

Tuplex is a parallel big data processing framework that runs data science pipelines written in Python at the speed of compiled code

Get mutations in cluster by querying from LAPIS API