Downstream pipeline examples

July 24, 2026 · View on GitHub

These scripts answer "why UniSchema?" — they consume normalized ConstituentEvent output and feed PhilanthroPy ML pipelines.

Primary guide: docs/philanthropy-integration.md

Install

# Basic analytics
pip install -r examples/downstream/requirements.txt

# PhilanthroPy ML bridge (optional)
pip install -r examples/downstream/requirements-philanthropy.txt

Local egress (pilot / Docker Compose)

After npm run demo:multi:

python3 examples/downstream/read_local_egress.py data/egress
python3 examples/downstream/philanthropy_crm_pipeline.py data/egress samples/crm-golden-record.csv

Or run the full chain:

npm run downstream-demo

Notebook: egress_report.ipynb — egress summary + optional PhilanthroPy histogram.

Feature column contract

philanthropy.ingest is the single source of truth: read_constituent_events walks the date-partitioned local egress tree (recursive; skips .manifest.json), and constituent_events_to_features aggregates it (leakage-safe, dedup by eventId). Owned and tested in the PhilanthroPy repo.

ColumnPhilanthroPy use
total_gift_amountDonorPropensityModel input
years_activeDonorPropensityModel input
event_attendance_countDonorPropensityModel input

Full contract → philanthropy-integration.md

Scripts

ScriptPurpose
philanthropy_crm_pipeline.pyRecommended ML path with CRM labels
philanthropy_pipeline.pyDemo with proxy labels
crm_join_example.pyCRM join (externalConstituentId or email)
read_local_egress.pyStakeholder text report
read_s3_ndjson_batch.pyProduction S3 batch reader

S3 NDJSON batches (production)

When EGRESS_TARGET=s3, UniSchema flushes micro-batches:

s3://{bucket}/{prefix}/batches/{YYYY}/{MM}/{DD}/{batchId}.ndjson
s3://{bucket}/{prefix}/batches/{YYYY}/{MM}/{DD}/{batchId}.manifest.json
pip install boto3
python3 examples/downstream/read_s3_ndjson_batch.py \
  s3://your-bucket/constituent-events/batches/2026/06/20/abc123.ndjson

dbt

See dbt/README.md — includes mart_constituent_rfm_features for PhilanthroPy batch scoring.