┌──────────────────────────────────────────────┐
│ Frontend (Streamlit) │
│ Upload → Extract → Validate → Convert │
├──────────────────────────────────────────────┤
│ Backend │
│ ┌─────────┐ ┌──────────┐ ┌──────────────┐ │
│ │Extractor│ │Converter │ │ Compressor/ │ │
│ │ (.txt → │ │(.asc → │ │ Encryptor/ │ │
│ │ JSON) │ │ CSV) │ │ Headers │ │
│ └─────────┘ └──────────┘ └──────────────┘ │
├──────────────────────────────────────────────┤
│ Data layer │
│ 01-input/ + 00-metadata/ → 02-output/ │
│ (.asc) (.json, .xlsx) (.csv, .parquet) │
└──────────────────────────────────────────────┘
Dependencies: polars, pandas, streamlit, pyjanitor, rich, fastexcel
| Aspect | Local | In Cloud |
|---|---|---|
| Data storage | Local filesystem | MinIO (S3) or SURF Object Store |
| PII handling | SHA256 hashing in code | + platform-level encryption at rest? |
| User access | Single user, localhost | Admin and end-user |
| Pipeline trigger | Manual (Streamlit click) | Orchestrated (Dagster?) |
| Monitoring | Console output, JSON logs | Grafana + Loki |
Notes:
No Fairness Without Awareness — R package for analyzing bias in student data.
Developed by LTA lectorate (The Hague University). Trains ML models to predict student retention, then evaluates fairness across sensitive variables (gender, prior education, socioeconomic status).
Pipeline:
Student data → Transform & enrich → Train models → Fairness analysis → PDF report
(1CijferHO) (join, impute, (logistic reg (DALEX explainer (Quarto +
add SES/APCG) + random forest) + fairmodels) LaTeX)
Models trained:
Tuning:
Data split: 60/20/20 (stratified)
Fairness evaluation:
Config-driven:
variabelen.xlsx — which vars to include, which are sensitiveconfig.yml — program name, year, formlevels.xlsx — factor level ordering┌──────────────────────────────────────────────────┐
│ Orchestration (main.R → run_nfwa) │
├──────────────────────────────────────────────────┤
│ Transform layer │
│ transform_ev_data → transform_vakhavw │
│ → transform_1cho_data → add_apcg → add_ses │
├──────────────────────────────────────────────────┤
│ ML layer │
│ ┌───────────────┐ ┌────────────────────┐ │
│ │ run_models() │ │ Fairness pipeline │ │
│ │ logistic + RF │→ │ DALEX → fairmodels │ │
│ │ tuning + CV │ │ per sensitive var │ │
│ └───────────────┘ └────────────────────┘ │
├──────────────────────────────────────────────────┤
│ Output: PDF report (Quarto + LaTeX) │
│ Fairness plots, density plots, conclusions │
└──────────────────────────────────────────────────┘
Dependencies: tidymodels, ranger, glmnet, DALEX, fairmodels, ggplot2, quarto
| Aspect | Local | On SDP |
|---|---|---|
| Data input | Local parquet/CSV | MinIO or PostgreSQL |
| Model artifacts | Local files | Model registry (MLflow?) |
| Experiment tracking | None | MLflow Tracking Server? |
| Fairness reports | Local PDF | Stored + versioned |
| Compute | Single machine | K8s job with resource limits |
| Reproducibility | renv lockfile | + container image (Harbor?) |
Test questions:
1cijferho NFWA
┌──────────────┐ ┌──────────────┐
│ Streamlit │ │ PDF report │
│ ┌────────┐ │ │ ┌────────┐ │
│ │Pipeline│ │ │ │run_nfwa│ │
│ │┌──────┐│ │ │ │┌──────┐│ │
│ ││ func ││ │ │ ││ func ││ │
│ │└──────┘│ │ │ │└──────┘│ │
│ └────────┘ │ │ └────────┘ │
└──────────────┘ └──────────────┘
Functions work standalone → composed in pipeline → wrapped in UI → deployed on platform
Config switch determines where each level connects:
Operational load
Tooling decisions
These repos surface 4 of the 14 gaps identified in the SDP gap analysis — practical test cases instead of abstract requirements.
Caspar and Tomer have hands-on experience with both repositories.
| Role | CEDA | Team SDA |
|---|---|---|
| Domain knowledge | Educational analytics, data pipelines | Platform engineering, deployment |
| Code ownership | R/Python packages, business logic | Infrastructure, CI/CD, monitoring |
| Experiment focus | "Does it produce correct results?" | "Does it run reliably at scale?" |
Joint experiments (Sep–Dec 2026):
| Period | Activity | With |
|---|---|---|
| May–Aug 2026 | Prepare containerized versions of both repos | CEDA |
| Document current config patterns and data flows | CEDA | |
| Align on SDP namespace and access setup | CEDA + Team SDA | |
| Sep–Dec 2026 | Deploy 1cijferho on SDP (Streamlit + pipeline) | CEDA + Team SDA |
| Explore PII handling requirements with SDP team | CEDA + Team SDA | |
| Deploy NFWA as K8s batch job | CEDA + Team SDA | |
| Experiment with MLflow, OpenLineage on real workloads | CEDA + Team SDA | |
| Evaluate audit logging options for ML decisions | CEDA + Team SDA | |
| Document operational load and findings | CEDA + Team SDA |
Findings feed into tooling choices and gap resolution for the broader Educational Data Space platform.