Template for a Dataflow Flex Template in Python

Last update: Apr 28, 2022

Related tags

Overview

Dataflow Flex Template in Python

This repository contains a template for a Dataflow Flex Template written in Python that can easily be used to build Dataflow jobs to run in STOIX using Dataflow runner.

The code is based on the same example data as Google Cloud Python Quickstart, "King Lear" which is a tragedy written by William Shakespeare.

The Dataflow job reads the file content, count occurencies of each word and inserts it to a BigQuery table. The schedule date is also added to the table name producing a sharded table for the output.

Source data:

https://storage.cloud.google.com/dataflow-samples/shakespeare/kinglear.txt
gs://dataflow-samples/shakespeare/kinglear.txt

Template maintained by STOIX.

Configuration

The job is configured with the following pipeline options:

stoix_scheduled - Scheduled datetime as RFC3339
input_file - Text to read
output_dataset - BigQuery dataset for output table
output_table_prefix - BigQuery output table name prefix
project - Google Cloud project id

When using Dataflow runner, stoix_scheduled is automatically set and other pipeline options can be added as described in the Dataflow runner README.

Test the code

Tox is used to format, test and lint the code. Make sure to install it with pip install tox and then just run tox within the project folder.

Run pipeline

In order to work with the code locally, you can use Python virtual environments. Make sure to use Python version 3.7.10 as it is the version supported by Google Dataflow.

$ python3 -m venv venv
$ source venv/bin/activate
$ pip install -e .

Run on local machine

See quickstart python for further description of arguments.

python -m main \
    --region europe-north1 \
    --runner DirectRunner \
    --stoix_scheduled 2021-01-01T00:00:00Z \
    --input_file gs://dataflow-samples/shakespeare/kinglear.txt \
    --output_table_prefix kinglear \
    --output_dataset 
   
     \
    --project 
    
      \
    --temp_location gs://
     
      /tmp/

Build Docker image for STOIX

In order to run the pipeline the Flex Template needs to be packaged in a Docker image and pushed to a Docker image repository. In this example Docker Hub is used.

Set the tag to the name and version of your pipeline, e.g: stoix/count-words:1.0.0.

$ docker build --tag stoix/count-words:1.0.0 .

Then upload the image to the Docker image repository.

$ docker push stoix/count-words:1.0.0

Run Dataflow on STOIX

Now the Dataflow Flex Template job can be ran using Dataflow runner. Add a new job with the image stoix/dataflow-runner and the following environment variables:

GCP_PROJECT_ID:
GCP_REGION: europe-north1
GCP_SERVICE_ACCOUNT: BASE64 encoded service account JSON
JOB_IMAGE: stoix/count-words:1.0.0
JOB_NAME_PREFIX: count-words
JOB_PARAM_INPUT_FILE: gs://dataflow-samples/shakespeare/kinglear.txt
JOB_PARAM_OUTPUT_DATASET: dataflow
JOB_PARAM_OUTPUT_TABLE_PREFIX: kinglear
JOB_SDK_LANGUAGE: python

Note: When running this in production, set GCP_SERVICE_ACCOUNT as a secret instead of environment variable.

License

MIT

Template for a Dataflow Flex Template in Python

Related tags

Overview

Dataflow Flex Template in Python

Configuration

Test the code

Run pipeline

Build Docker image for STOIX

Run Dataflow on STOIX

License

Owner

STOIX

Binance Kline Data With Python

Hg002-qc-snakemake - HG002 QC Snakemake

A project consists in a set of assignements corresponding to a BI process: data integration, construction of an OLAP cube, qurying of a OPLAP cube and reporting.

A computer algebra system written in pure Python

Cleaning and analysing aggregated UK political polling data.

An Indexer that works out-of-the-box when you have less than 100K stored Documents

Nobel Data Analysis

Synthetic data need to preserve the statistical properties of real data in terms of their individual behavior and (inter-)dependences

MeSH2Matrix - A set of Python codes for the generation of biomedical ontologies from the MeSH keywords of the PubMed scholarly publications

Driver Analysis with Factors and Forests: An Automated Data Science Tool using Python

Python Kalman filtering and optimal estimation library. Implements Kalman filter, particle filter, Extended Kalman filter, Unscented Kalman filter, g-h (alpha-beta), least squares, H Infinity, smoothers, and more. Has companion book 'Kalman and Bayesian Filters in Python'.

Synthetic Data Generation for tabular, relational and time series data.

Minimal working example of data acquisition with nidaqmx python API

CPSPEC is an astrophysical data reduction software for timing

Improving your data science workflows with

Investigating EV charging data

signac-flow - manage workflows with signac

Create HTML profiling reports from pandas DataFrame objects

Udacity - Data Analyst Nanodegree - Project 4 - Wrangle and Analyze Data

This mini project showcase how to build and debug Apache Spark application using Python