Data factory for the VIEWS conflict forecasting platform — harvesting, consolidation, viewpoint building, grid compilation, and query.
Part of the VIEWS Platform ecosystem.
- Overview
- Role in VIEWS Platform
- Architecture
- Packages
- Project Structure
- Installation
- Quick Start
- Strategic Documents
- Where to Find What
- Contributing
- License
views-datafactory provides the data foundation for conflict forecasting in the VIEWS platform. It handles the full lifecycle of spatiotemporal conflict data — from raw ingestion through to compiled grid arrays ready for model consumption.
The system is designed as a graph of independent nodes, not a linear pipeline. Source nodes (harvesters) produce data independently. Compilation edges transform source data into consumer-specific formats. Consumer nodes read compiled outputs. This architecture ensures that adding a new data source or output format never requires changing existing components.
- Auditable data harvesting — UCDP (annual, candidate, .9), ACLED, GHS-POP, GHS-BUILT-S, V-Dem, SHDI, PRIO-GRID static, and GAUL admin boundaries, all with schema validation, drift detection, and JSONL provenance ledgers
- PRIO-GRID spatial backbone — pure-numpy generation of the standard 259,200-cell global grid at 0.5° resolution, validated cell-by-cell against the official PRIO reference shapefile
- Temporal backbone with VIEWS month_id adapters (epoch: January 1980)
- Vintage-aware consolidation — lossless, append-only, bitemporal event store from all three UCDP sources with content-digest deduplication
- Configurable viewpoints — survivorship strategies (annual_wins, dot9_wins), temporal distribution (even_split, ceil_split), production filtering via named profiles
- Deterministic compilation of viewpoint output onto the spatiotemporal grid as npy arrays
- Consumer adapters — grid-to-DataFrame and grid-to-FeatureFrame conversions for downstream consumers
- End-to-end provenance — every value in a compiled grid is traceable back to the specific source records and compilation config that produced it
views-datafactory is the upstream data provider for the VIEWS forecasting pipeline. It produces the compiled grid arrays that models consume for training and inference.
- views-metric-lab: First consumer — uses compiled grids for evaluation metric research and model stress-testing
- views-hydranet: Will consume compiled grids for production neural forecasting (HydraNet v2)
- views-pipeline-core: Central pipeline infrastructure — may depend on views-datafactory for standardised data access
- views-evaluation: May consume compiled grids for benchmarking
No layer imports from the layer above or below. The filesystem is the API between layers. Each dashed line below represents zero code coupling — only a file path convention. This is enforced by an import-enforcement test that AST-scans every package.
UCDP/GED API ──→ datafactory_harvester ──→ data/raw/ucdp_*/*.parquet
ACLED API ──────→ datafactory_harvester ──→ data/raw/acled/*.parquet
(event sources)
─ ─ ─ ─ ─ ─ ─ ─ filesystem boundary
│
datafactory_consolidation ──→ data/consolidated/*.parquet
│
─ ─ ─ ─ ─ ─ ─ ─ filesystem boundary
│
datafactory_viewpoint ──→ data/viewpoint/*.parquet
│
─ ─ ─ ─ ─ ─ ─ ─ filesystem boundary
│
PRIO-GRID ──────→ datafactory_priogrid ─┐ │
PRIO-GRID API ──→ datafactory_harvester ──→ data/raw/priogrid_static/
FAO GAUL API ───→ datafactory_harvester ──→ data/raw/gaul_admin/
JRC GHSL ───────→ datafactory_harvester ──→ data/raw/ghspop/ + ghsbuilts/
V-Dem ──────────→ datafactory_harvester ──→ data/raw/vdem/
GDL SHDI ───────→ datafactory_harvester ──→ data/raw/shdi/
(GHS-POP, GHS-BUILT-S, V-Dem, SHDI skip consolidation — single releases)
│ │
datafactory_compilation ──→ data/compiled/grid.npy [T,H,W,C]
│
─ ─ ─ ─ ─ ─ ─ ─ filesystem boundary
│
assemble_grid.py ──→ data/assembled/grid.npy [T,H,W,C]
(compiled UCDP + ACLED + GHS-POP + GHS-BUILT-S + V-Dem + SHDI + static + admin)
│
─ ─ ─ ─ ─ ─ ─ ─ filesystem boundary
│
datafactory_query ──→ FeatureFrame / DataFrame
(load_dataset: region, time, feature subsetting)
│
generate_consumer_data.py ──→ {run_type}_viewser_df.parquet
(factory → VIEWSER vocabulary)
│
┌───────────┴───────────┐
│ Consumers │
│ (metric lab, │
│ hydranet) │
└───────────────────────┘
Not all paths traverse all layers. GHS-POP, GHS-BUILT-S, V-Dem, and SHDI skip consolidation (single releases; ADR-029, ADR-034, ADR-035, ADR-036).
The architecture is a graph, not a pipeline. Each node is a self-contained package with explicit dependencies:
Layer 0 — Foundation (no internal imports):
datafactory_provenance Content digests, JSONL ledger operations
datafactory_http HTTP request utilities: retry with exponential backoff
Layer 1 — Source nodes (import provenance + http):
datafactory_priogrid PRIO-GRID spatial + temporal backbone
datafactory_harvester Data ingestion with pluggable sources
Layer 2 — Consolidation (imports provenance, reads harvester files):
datafactory_consolidation Lossless event store from raw snapshots
Layer 3 — Viewpoint (imports provenance, reads consolidation files):
datafactory_viewpoint Opinionated materialized views
Layer 4 — Compilation (imports provenance + priogrid, reads viewpoint files):
datafactory_compilation Viewpoint output → populated grid arrays
Consumer-facing (no datafactory_* imports):
datafactory_adapters Grid → DataFrame, Grid → FeatureFrame
Assembly (script, not package — combines compiled sources):
scripts/assemble_grid.py Compiled UCDP + ACLED + GHS-POP + GHS-BUILT-S + V-Dem + SHDI + static + admin → [T,H,W,C]
Consumer entry point (imports priogrid + adapters, reads assembled files):
datafactory_query load_dataset() — region/time/feature subsetting, npy or zarr
Dependency rules (ADR-012):
datafactory_provenanceimports nothing internaldatafactory_priogridanddatafactory_harvesterimport only fromdatafactory_provenance(anddatafactory_httpfor network operations)datafactory_consolidationimports only fromdatafactory_provenance; reads harvester output as filesdatafactory_viewpointimports only fromdatafactory_provenance; reads consolidation output as filesdatafactory_compilationimportsdatafactory_provenanceanddatafactory_priogrid; reads viewpoint output as filesdatafactory_adaptersimports nothing fromdatafactory_*; sits alongside the graph, not inside itdatafactory_queryimportsdatafactory_priogridanddatafactory_adapters; reads assembled output as files (npy or zarr)- Independence is enforced by an import-enforcement test that AST-scans every package
| Principle | Description |
|---|---|
| Graph, not pipeline | Sources don't know about consumers. Compilation edges are consumer-driven. |
| Provenance all the way through | Every node that produces data writes a JSONL ledger entry. Mission-critical. |
| Config-driven, fail-loud | Frozen dataclasses with __post_init__ validation. No silent defaults. |
| npy now, zarr-ready | Dimension order: [T, H, W, C] (time, height, width, channels). Coordinate arrays shipped alongside data. |
| Screaming architecture | Package and file names communicate domain, not programming taxonomy. |
Layers are decoupled by the filesystem, not by imports. Each layer computes SHA-256 of its input file (source_digest) and output file (output_digest), recording both in its JSONL ledger. This creates a cryptographic provenance chain from raw API response to consumer parquet — any tampering or re-run is detectable by comparing digests across ledgers.
| Boundary | Writer | Reader | File Path | Format | Key Columns / Shape |
|---|---|---|---|---|---|
| Harvest → Consolidation | harvest_ucdp.py |
consolidate_ucdp.py |
data/raw/ucdp_*/ |
Parquet | id, date_start, date_end, best, low, high, type_of_violence |
| Consolidation → Viewpoint | consolidate_ucdp.py |
build_viewpoint.py |
data/consolidated/ |
Parquet | All harvest columns + _source_type, _source_version, _ingested_at |
| Viewpoint → Compilation | build_viewpoint.py |
compile_grid.py |
data/viewpoint/*.parquet |
Parquet | latitude, longitude, date_month, best, type_of_violence |
| Compilation → Assembly | compile_grid.py |
assemble_grid.py |
data/compiled/ |
npy [T,H,W,C] |
grid.npy + pgids.npy + time_steps.npy + feature_names.json |
| Assembly → Query/Consumer | assemble_grid.py |
load_dataset() |
data/assembled/ or .zarr |
npy or zarr | grid.npy [T,H,W,C] with same sidecars |
The generate_consumer_data.py script produces parquet files for VIEWS training scripts (purple_alien, white_ranger, etc.). Column names are the factory's source names — consumer-side renaming (if any) is the model's responsibility via config_queryset.py.
| Column | Meaning |
|---|---|
ged_sb_best |
State-based conflict fatalities (best estimate) |
ged_ns_best |
Non-state conflict fatalities |
ged_os_best |
One-sided violence fatalities |
gaul0_code |
Country identifier (GAUL Level 0) |
row, col |
Grid position (derived from priogrid_gid) |
Training partitions use VIEWS month_id encoding (epoch: January 1980):
- Calibration: train 121–444, test 445–492
- Validation: train 121–492, test 493–540
- Forecasting: train 121–current, test current+1 to current+36
Output: {calibration,validation,forecasting}_viewser_df.parquet with MultiIndex (month_id, priogrid_gid).
| Package | Layer | Purpose | Status |
|---|---|---|---|
datafactory_provenance |
0 | Content digests, JSONL ledgers, file locking, rotation | Done |
datafactory_http |
0 | HTTP request utilities: retry with exponential backoff | Done |
datafactory_priogrid |
1 | PRIO-GRID spatial + temporal backbone (259,200 cells, monthly) | Done |
datafactory_harvester |
1 | Data harvesting: UCDP (annual/candidate/.9), ACLED, GHS-POP, GHS-BUILT-S, V-Dem, SHDI, PRIO-GRID static, GAUL admin | Done |
datafactory_consolidation |
2 | Lossless, vintage-aware consolidation of UCDP and ACLED sources | Done |
datafactory_viewpoint |
3 | Opinionated views: survivorship, temporal distribution, profiles | Done |
datafactory_compilation |
4 | Viewpoint output → grid npy with coordinate sidecars | Done |
datafactory_adapters |
— | Consumer-facing: grid → DataFrame, grid → FeatureFrame | Done |
datafactory_query |
— | Consumer entry point: load_dataset(), region/time subsetting, dual npy/zarr backend |
Done |
views-datafactory/
├── pyproject.toml # hatchling + uv
├── CLAUDE.md # AI assistant context
├── .github/workflows/ci.yml # CI: ruff + mypy + pytest
├── src/
│ ├── datafactory_provenance/ # Layer 0 — digests + ledgers
│ │ ├── digests_and_ledgers.py
│ │ ├── constants.py VIEWS_EPOCH_YEAR (shared constant)
│ │ └── source_registry.py SourceEntry, PIPELINE_SOURCES
│ ├── datafactory_http/ # Layer 0 — HTTP retry utilities
│ │ └── retry.py request_with_retry (shared)
│ ├── datafactory_priogrid/ # Layer 1 — PRIO-GRID backbone
│ │ ├── grid_config.py GridConfig (spatial params)
│ │ ├── temporal_config.py TemporalConfig (year/month range)
│ │ ├── cell_generator.py generate_grid, pgid_to_latlon
│ │ ├── temporal_generator.py generate_time_steps, VIEWS month_id
│ │ ├── spatiotemporal.py SpatioTemporalGrid (composition)
│ │ ├── parity_validation.py validate_parity (vs. reference shapefile)
│ │ ├── shapefile_reader.py PyShpReader (pluggable Protocol)
│ │ └── shapefile_harvester.py fetch_shapefile (reference data)
│ ├── datafactory_harvester/ # Layer 1 — data ingestion
│ │ ├── event_validation.py ValidationResult, compare_snapshots
│ │ ├── snapshot_storage.py save_event_snapshot, archive_snapshot
│ │ └── sources/ Source plugin registry
│ │ ├── ucdp_annual.py UCDP/GED Annual source
│ │ ├── ucdp_candidate.py UCDP/GED Candidate Monthly source
│ │ ├── ucdp_dot9.py UCDP/GED .9 Consolidated Monthly
│ │ ├── acled.py ACLED event data (year-by-year)
│ │ ├── ghspop.py GHS-POP R2023A (JRC GHSL raster)
│ │ ├── ghsbuilts.py GHS-BUILT-S R2023A (JRC GHSL raster)
│ │ ├── vdem.py V-Dem v16 democracy indicators
│ │ ├── shdi.py SHDI subnational human development
│ │ ├── priogrid_static.py PRIO-GRID static covariates
│ │ └── gaul_admin.py GAUL 2024 admin boundaries
│ ├── datafactory_consolidation/ # Layer 2 — lossless event stores
│ │ ├── event_store.py Append-only Parquet store
│ │ ├── tagging.py tag_table (shared metadata tagging)
│ │ └── consolidators/
│ │ ├── ucdp.py Three-source UCDP consolidation
│ │ └── acled.py ACLED event consolidation
│ ├── datafactory_viewpoint/ # Layer 3 — opinionated views
│ │ ├── survivorship.py Strategy registry (annual_wins, dot9_wins)
│ │ ├── temporal_distribution.py Strategy registry (even_split, ceil_split)
│ │ ├── profiles.py Named presets (production_parity)
│ │ ├── raster_io.py read_geotiff (shared GeoTIFF reader)
│ │ ├── temporal.py interpolate_temporal (shared interpolation)
│ │ └── builders/
│ │ ├── ucdp_v1.py UCDP viewpoint builder
│ │ ├── ghspop_v1.py GHS-POP viewpoint builder
│ │ ├── ghsbuilts_v1.py GHS-BUILT-S viewpoint builder
│ │ ├── vdem_v1.py V-Dem viewpoint builder
│ │ └── shdi_v1.py SHDI viewpoint builder
│ ├── datafactory_compilation/ # Layer 4 — grid compilation
│ │ ├── compilation_config.py CompilationConfig
│ │ ├── grid_compilation.py compile_grid (event sources)
│ │ ├── pregridded_compilation.py compile_pregridded (raster sources)
│ │ ├── output.py write_compilation_output (shared writer)
│ │ └── aggregation.py count, sum_field, max_field strategies
│ ├── datafactory_adapters/ # Consumer-facing conversions
│ │ ├── feature_frame.py FeatureFrame dataclass
│ │ ├── grid_to_dataframe.py Grid → pandas DataFrame
│ │ └── grid_from_feature_frame.py FeatureFrame → Grid (inverse)
│ └── datafactory_query/ # Consumer entry point
│ ├── dataset.py load_dataset (npy + zarr backends)
│ ├── regions.py Region name → PRIO-GRID cell set
│ └── temporal.py Flexible time range parsing
├── tests/ # ~2000 tests
├── scripts/ # Operational scripts
│ ├── harvest_ucdp.py UCDP harvest (annual + candidate + .9)
│ ├── harvest_acled.py ACLED event harvest (year-by-year)
│ ├── harvest_ghspop.py GHS-POP raster harvest (12 epochs)
│ ├── harvest_ghsbuilts.py GHS-BUILT-S raster harvest (12 epochs)
│ ├── harvest_priogrid.py PRIO-GRID static covariates
│ ├── harvest_shapefile.py PRIO-GRID shapefile download
│ ├── harvest_gaul.py GAUL admin boundary harvest + spatial join
│ ├── consolidate_ucdp.py Three-source UCDP consolidation
│ ├── build_viewpoint.py Viewpoint from profile
│ ├── compile_grid.py Event source grid compilation
│ ├── compile_acled.py ACLED grid compilation
│ ├── run_acled_pipeline.py ACLED end-to-end pipeline
│ ├── run_ghspop_pipeline.py GHS-POP end-to-end pipeline
│ ├── run_ghsbuilts_pipeline.py GHS-BUILT-S end-to-end pipeline
│ ├── harvest_vdem.py V-Dem CSV harvest
│ ├── compile_vdem.py V-Dem grid compilation
│ ├── run_vdem_pipeline.py V-Dem end-to-end pipeline
│ ├── harvest_shdi.py SHDI subnational HDI harvest
│ ├── run_shdi_pipeline.py SHDI end-to-end pipeline
│ ├── assemble_grid.py UCDP + ACLED + GHS-POP + GHS-BUILT-S + V-Dem + SHDI + static + admin
│ ├── generate_consumer_data.py Factory → VIEWSER parquet for training
│ ├── export_dataframe.py Grid → DataFrame export
│ ├── export_zarr.py Grid → zarr store (HTTP-servable)
│ ├── preflight.py Pre-pipeline disk + env checks
│ ├── check_health.py System health check
│ ├── refresh_pipeline.sh Full pipeline orchestrator (14 steps)
│ ├── visualize_audit.py Data audit visualization
│ └── ... verify_parity, verify_remote, etc.
├── docs/ # ADRs, CICs, protocols, standards
│ ├── ADRs/ 10 constitutional + 39 project-specific
│ ├── CICs/ 34 active class intent contracts
│ ├── contributor_protocols/ carbon, silicon, hardened
│ └── standards/ logging & observability
├── reports/ # Strategic documents + audit outputs
│ ├── rd_roadmap11.md R&D roadmap (v11)
│ ├── product_development_plan11.md Product development plan (v11)
│ ├── technical_risk_register.md Technical risk register (ADR-020)
│ └── dot9_investigation/ .9 data stream research findings
├── provenance/ # JSONL ledgers (gitignored)
└── data/ # Raw + compiled data (gitignored)
As a consumer (you want load_dataset() and the adapters):
pip install views-datafactoryAs a developer (you want the tests, scripts, and docs):
# Clone the repository
git clone https://github.com/views-platform/views-datafactory.git
cd views-datafactory
# Install with uv
uv sync
# Verify
uv run pytest -vNote: This project uses uv for dependency management and hatchling as the build backend.
| You want to... | Credentials needed |
|---|---|
| Develop / run the test suite | None. A bare checkout runs the full suite. |
| Read served data (remote zarr) | An entry in ~/.netrc for the VIEWS data server — ask the data factory administrator for the password. |
| Run the harvest pipeline yourself | UCDP_API_TOKEN, ACLED_USERNAME/ACLED_PASSWORD, GDL_API_TOKEN (free registrations with the respective providers). GHS-POP, GHS-BUILT-S, V-Dem, PRIO-GRID static, and GAUL need no credentials. |
Setup instructions for all of these: docs/guides/credential_setup.md. New to the whole platform and just want to run a model? Start at docs/guides/model_consumer_quickstart.md.
# Full pipeline (automated): runs all 14 steps end-to-end
bash scripts/refresh_pipeline.sh
# Or run individual steps:
uv run python scripts/harvest_ucdp.py # fetch all UCDP sources
uv run python scripts/harvest_acled.py # fetch ACLED events
uv run python scripts/harvest_ghspop.py # fetch GHS-POP rasters
uv run python scripts/harvest_ghsbuilts.py # fetch GHS-BUILT-S rasters
uv run python scripts/consolidate_ucdp.py # three-source UCDP consolidation
uv run python scripts/build_viewpoint.py # apply production_parity profile
uv run python scripts/compile_grid.py # UCDP → grid.npy
uv run python scripts/compile_acled.py # ACLED → grid.npy
uv run python scripts/run_ghspop_pipeline.py # GHS-POP viewpoint + compile
uv run python scripts/run_ghsbuilts_pipeline.py # GHS-BUILT-S viewpoint + compile
uv run python scripts/run_vdem_pipeline.py # V-Dem viewpoint + compile
uv run python scripts/run_shdi_pipeline.py # SHDI viewpoint + compile
uv run python scripts/assemble_grid.py \ # assemble all sources into one grid
--acled-grid data/compiled/acled \ # (each --*-grid flag is optional;
--ghspop-grid data/compiled/ghspop \ # omitted sources are silently skipped,
--ghsbuilts-grid data/compiled/ghsbuilts \ # producing a partial grid)
--vdem-grid data/compiled/vdem \
--shdi-grid data/compiled/shdi
uv run python scripts/generate_consumer_data.py # factory → VIEWSER training parquets
# Visualize the assembled grid
uv run python scripts/visualize_audit.py # 15 audit plots → reports/audit/
# Export for downstream consumers
uv run python scripts/export_zarr.py # grid → zarr store (HTTP-servable)
uv run python scripts/export_dataframe.py # grid → DataFrame CSV# Programmatic usage — load_dataset() is the primary consumer entry point
from datafactory_query import load_dataset
# Load Ethiopia 2020-2024 as FeatureFrame (temporal + geographic subsetting)
ff = load_dataset(region="Ethiopia", start="2020", end="2024")
# Load global land cells as DataFrame
df = load_dataset(region="land", output_format="dataframe")
# Load from the remote zarr store (requires ~/.netrc — see docs/guides/credential_setup.md)
from datafactory_query.defaults import DEFAULT_REMOTE
ff = load_dataset(
region="africa_me",
start="2020",
data_dir=DEFAULT_REMOTE.zarr_url,
)# Low-level access — direct npy loading
from datafactory_adapters import FeatureFrame, grid_to_dataframe
import json, numpy as np
data = np.load("data/assembled/grid.npy", mmap_mode="r")
pgids = np.load("data/assembled/pgids.npy")
time_steps = np.load("data/assembled/time_steps.npy")
feature_names = json.loads(open("data/assembled/feature_names.json").read())
df = grid_to_dataframe(data, pgids, time_steps, feature_names, month_id_epoch=1980)
ff = FeatureFrame.from_grid(data, pgids, time_steps, feature_names)The reports/ directory contains living documents that define the project's direction:
- R&D Roadmap — Research questions, hypotheses, data agenda, milestones. Focuses on what must be discovered.
- Product Development Plan — Users, requirements, architecture, release plan. Focuses on what must be built.
- Technical Risk Register — 298 concerns tracked, 258 resolved, 37 open with trigger conditions (ADR-020).
- .9 Investigation — Empirical findings on UCDP .9 data stream characteristics.
This system separates information by ownership and change frequency (ADR-003). Different questions have different authoritative sources:
| If you need to know… | Look in… | Why there |
|---|---|---|
| How to set up credentials, get data, run a model, or operate the server | docs/guides/ — how-to guides (see its README for an index) |
Task-oriented walkthroughs |
| What data sources we use, who publishes them, what license/format/coverage | docs/sources/ — catalog cards |
Upstream facts, rarely change |
| Why we chose a specific source, what alternatives we considered | docs/ADRs/ — source selection ADRs (028, 029) |
One-time decisions |
| How a derived feature is constructed (aggregation, survivorship, distribution) | Viewpoint CICs (docs/CICs/Viewpoint*.md) and profile configs (src/datafactory_viewpoint/profiles.py) |
Volatile — changes as research evolves |
| What SLO, env vars, or feature list a source has | src/datafactory_provenance/source_registry.py — SourceEntry |
Operational config, changes on redeploy |
| Current harvest status, data versions, content digests | provenance/*.jsonl ledgers, or run uv run python scripts/check_health.py |
Volatile state, changes every pipeline run |
| What a config class guarantees | docs/CICs/ — Class Intent Contracts |
Code contracts |
| Why the system is designed this way | docs/ADRs/ — Architecture Decision Records |
Immutable decisions |
| Known technical risks and their status | reports/technical_risk_register.md |
Living risk tracking |
Rule of thumb: if the information changes when the upstream provider changes something → catalog card. When we make a decision → ADR. When we redeploy → code config. When the pipeline runs → provenance ledger. When in doubt, the provenance ledger is the most authoritative source for anything operational.
Contributions are welcome. Please consult the VIEWS Platform contributing guidelines before submitting pull requests.
When adding new packages, follow the existing datafactory_* naming convention and ensure the new package:
- Has an
__init__.pywith__all__defined - Is listed in
pyproject.tomlunder[tool.hatch.build.targets.wheel] packages - Imports only from
datafactory_provenance(or provenance + priogrid for compilation nodes) - Has corresponding tests in
tests/ - Has an
ARCHITECTURE.mddescribing purpose, boundaries, and invariants
This project is licensed under the MIT License — see LICENSE. The license covers the software in this repository only; upstream data sources (UCDP, ACLED, GHS-POP, GHS-BUILT-S, V-Dem, SHDI, PRIO-GRID, GAUL, WDI) are subject to their respective providers' terms of use.
Built as part of the VIEWS (Violence & Impacts Early-Warning System) project, providing early warning for conflict to support humanitarian response and prevention.
Data sources:
