Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Last update: Dec 10, 2022

Overview

https://travis-ci.org/uchicago-cs/deepdish.svg?branch=master

https://img.shields.io/badge/license-BSD%203--Clause-blue.svg?style=flat

deepdish

Flexible HDF5 saving/loading and other data science tools from the University of Chicago. This repository also host a Deep Learning blog:

http://deepdish.io

Installation

pip install deepdish

Alternatively (if you have conda with the conda-forge channel):

conda install -c conda-forge deepdish

Main feature

The primary feature of deepdish is its ability to save and load all kinds of data as HDF5. It can save any Python data structure, offering the same ease of use as pickling or numpy.save. However, it improves by also offering:

Interoperability between languages (HDF5 is a popular standard)
Easy to inspect the content from the command line (using h5ls or our specialized tool ddls)
Highly compressed storage (thanks to a PyTables backend)
Native support for scipy sparse matrices and pandas DataFrame, Series and Panel
Ability to partially read files, even slices of arrays

An example:

import deepdish as dd

d = {
    'foo': np.ones((10, 20)),
    'sub': {
        'bar': 'a string',
        'baz': 1.23,
    },
}
dd.io.save('test.h5', d)

This can be reconstructed using dd.io.load('test.h5'), or inspected through the command line using either a standard tool:

$ h5ls test.h5
foo                      Dataset {10, 20}
sub                      Group

Or, better yet, our custom tool ddls (or python -m deepdish.io.ls):

$ ddls test.h5
/foo                       array (10, 20) [float64]
/sub                       dict
/sub/bar                   'a string' (8) [unicode]
/sub/baz                   1.23 [float64]

Documentation

http://deepdish.readthedocs.io/

Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Related tags

Overview

deepdish

Installation

Main feature

Documentation

Owner

UChicago - Department of Computer Science

Python Project on Pro Data Analysis Track

4CAT: Capture and Analysis Toolkit

Learn machine learning the fun way, with Oracle and RedBull Racing

My first Python project is a simple Mad Libs program.

Orchest is a browser based IDE for Data Science.

Hidden Markov Models in Python, with scikit-learn like API

Repositori untuk menyimpan material Long Course STMKGxHMGI tentang Geophysical Python for Seismic Data Analysis

💬 Python scripts to parse Messenger, Hangouts, WhatsApp and Telegram chat logs into DataFrames.

Implementation in Python of the reliability measures such as Omega.

Senator Trades Monitor

PipeChain is a utility library for creating functional pipelines.

Functional tensors for probabilistic programming

BioMASS - A Python Framework for Modeling and Analysis of Signaling Systems

Big Data & Cloud Computing for Oceanography

Karate Club: An API Oriented Open-source Python Framework for Unsupervised Learning on Graphs (CIKM 2020)

Feature engineering and machine learning: together at last

fds is a tool for Data Scientists made by DAGsHub to version control data and code at once.

Spaghetti: an open-source Python library for the analysis of network-based spatial data

CS50 pset9: Using flask API to create a web application to exchange stocks' shares.

A powerful data analysis package based on mathematical step functions. Strongly aligned with pandas.