omicau

July 12, 2026 · View on GitHub

A reproducible, leakage-safe, platform-agnostic command-line tool that audits multi-omic datasets: it ingests non-standardized matrices, aligns them, locks their provenance with a value-level cryptographic hash, tests for missingness bias and batch effects, benchmarks classical and neural data-fusion models under group-aware cross-validation, attributes predictive signal to individual features, and compiles an interactive dual clinical/research dashboard alongside machine-readable JSON/CSV assets.

The core runs fully offline with no LLM connection and no orchestration framework. Optional tiers (an LLM interpretation plugin, remote data hubs) degrade gracefully when their dependencies or network are absent.

📖 Documentation: https://tunabirgun.github.io/omicau/

Contents: What it answers · Installation · Quickstart · Step-by-step guide · Use your own data · Workflow · Methodology · Plain-language verdict · Reporting · Validation · Data hubs · CLI & config reference · Software versions · Reproducibility · HPC · Troubleshooting

New here? Clinicians / PIs: skim What it answers, then Reporting. Bioinformaticians: QuickstartUse your own dataMethodology. Power users: the CLI & configuration reference.


What it answers

For a set of omic layers and a clinical endpoint, omicau answers three questions:

  1. Does combining the layers actually help, or does one layer already carry the signal? (marginal gain from adding each modality)
  2. Is the data trustworthy, or is the apparent signal an artifact of batch effects, target-linked missingness, or information leakage? (adversarial diagnostics + control baselines)
  3. What drives the prediction, biologically? (leakage-safe feature attribution)

The four design pillars are universality (runs out of the box on laptops, Apple Silicon, Intel, and headless HPC nodes; MPS/CUDA/CPU auto-selected), reproducibility (seeded loops, pinned versions, an immutable data hash), usability (a dual clinical/research dashboard), and scientific validity (leakage prohibition, masked missing-value handling, control baselines).


Installation

omicau is on PyPI — one line installs it. Python ≥ 3.10; every option below gives the same omicau command.

pipx install omicau            # simplest: isolated CLI install (recommended for a tool)
pip install omicau             # or into your current environment
pip install "omicau[all]"      # + every cross-platform runtime feature

pipx keeps omicau in its own environment (nothing to conflict with your other packages); plain pip works too. The base install is flawless on Windows, macOS, and Linux — only the optional extras touch the network or a platform-specific dep.

Debian / Ubuntu (24.04+): the system Python is "externally managed" (PEP 668), so pip install omicau into it errors with externally-managed-environment. Use pipx (recommended) or a virtual environment — not system pip:

sudo apt install -y pipx && pipx ensurepath
pipx install omicau            # then restart the shell; `omicau` is on your PATH

Or a venv: python3 -m venv ~/.venvs/omicau && ~/.venvs/omicau/bin/pip install omicau. (apt install python3-omicau does not exist — omicau ships via PyPI/conda, not apt.)

Linux torch size: on Linux, pip/pipx pull PyTorch's default CUDA build (~2.5 GB of NVIDIA libraries), which is slow and unnecessary on a CPU-only box (e.g. WSL). For a small, fast CPU-only install, install torch from the CPU index first, then omicau reuses it:

python3 -m venv ~/.venvs/omicau && source ~/.venvs/omicau/bin/activate
pip install torch --index-url https://download.pytorch.org/whl/cpu   # ~200 MB, not ~2.5 GB
pip install omicau

conda users get the CPU build automatically via environment.yml. Keep the default CUDA wheel only if you have an NVIDIA GPU and want it.

With conda / mamba (installs CPU PyTorch as a prebuilt binary — no compiler step):

git clone https://github.com/tunabirgun/omicau.git && cd omicau
conda env create -f environment.yml     # or: mamba env create -f environment.yml
conda activate omicau

(Once the conda-forge feedstock lands, conda install -c conda-forge omicau will work directly.)

From source (development / latest main): git clone https://github.com/tunabirgun/omicau.git && cd omicau && pip install ".[all,dev]".

Optional extras — combine freely: omicau[x] from PyPI, .[x] from a checkout, or — if you installed with pipxpipx inject omicau <deps> (e.g. pipx inject omicau fastapi uvicorn python-multipart for [ui]; omicau's error messages print the exact command for your install):

ExtraAdds
[llm]AI plain-language verdict (Claude / ChatGPT / Gemini / local)
[data]remote data hubs (TCGA, CCLE, Xena, OpenPBTA, Metabolomics, Expression Atlas)
[ui]local no-code web UI (FastAPI + uvicorn)
[cptac]CPTAC hub only (needs a build toolchain; no Windows / py3.12 wheel)
[all]every cross-platform runtime feature (everything except cptac + dev)
[dev]pytest + coverage

Core dependencies: numpy, pandas, scipy, scikit-learn, torch, plotly, click, jinja2, tqdm. Releases are automatic: bumping version in pyproject.toml and merging to main publishes to PyPI via OIDC (publish-pypi.yml, no stored token), and conda-forge auto-follows PyPI. See RELEASING.md.

Quickstart

omicau check-env                                   # CPU/GPU + dependency status
omicau bootstrap --dataset mock --out-dir demo     # write a synthetic dataset
omicau run --config demo/config.json --cores 8     # full audit -> demo/run/report.html
omicau verify --config demo/config.json            # recompute the provenance hash

Open demo/run/report.html for the dashboard; demo/run/audit.json and the *.csv files are the machine-readable assets.

Optional no-code web UI

omicau is a CLI tool first. For non-coders there is an opt-in local UI:

pip install ".[ui]"
omicau ui                     # opens a browser wizard on 127.0.0.1

omicau ui starts a local, single-user server (auto-selected free port, bound to 127.0.0.1 behind a one-time token), opens a browser, and walks you through a wizard: drop your files (or load a public cohort from a data hub — TCGA, Xena, CCLE, CPTAC, OpenPBTA, Expression Atlas, Metabolomics Workbench, or an offline synthetic demo, via the same connectors the CLI bootstrap uses), confirm the auto-guessed omic role and orientation of each, map the clinical columns (target / patient-group / batch — each shows its live consequence, e.g. class balance or the leakage implication of the grouping), check cross-file alignment, optionally add a plain-language AI verdict (Claude, ChatGPT, Gemini, or a local model — the API key stays in memory for that one run and is never stored), then run the identical run_audit pipeline and read the report in-page. All data stays on the machine — nothing is uploaded. The UI only builds the config and shows the result; it never re-implements the science, so the CLI and UI always give the same answer.

For a no-install experience, a desktop app (double-click → opens the UI) can be packaged from packaging/: a PyInstaller onedir bundle with a private CPU-only PyTorch runtime, wrapped as a signed Windows installer (Inno Setup), a Linux AppImage, and a notarized macOS (arm64) .dmg, built per-OS in CI (.github/workflows/release.yml). See packaging/README.md for sizes and signing. HPC/headless clusters use the pip package + CLI, not the GUI.

Verifying a run's provenance hash

Every run prints and stores a SHA-256 hash of the aligned data. It is deterministic, so anyone can recompute it from the same inputs and confirm no drift:

omicau verify --config demo/config.json --expected <hash>   # exit 1 on mismatch
omicau verify --config demo/config.json --audit demo/run/audit.json  # compare stored vs recomputed

Step-by-step guide

The whole path, from a clean machine to a finished audit, with a pointer into the detailed section for each step.

  1. Install — see Installation (conda, from source, or pip/pipx once published).
  2. Sanity-check the machineomicau check-env prints CPU/GPU, folder access, and which optional tiers are present.
  3. Get data, one of three ways:
  4. Run the auditomicau run --config <config.json> --cores 8.
  5. Read the report — open <out>/run/report.html (executive tab first); audit.json + *.csv are the machine-readable exports.
  6. (optional) Add a written AI verdict — see Plain-language verdict.
  7. Verify provenanceomicau verify --config <config.json> --audit <out>/run/audit.json.

Finding a public dataset

The [data] extra lets omicau bootstrap assemble a ready-to-run dataset from a public hub. Browse the hub, copy an identifier, pass it as the argument.

HubOrganism(s)Where to find an identifierExample
Expression Atlasany organismebi.ac.uk/gxa/experiments — filter by species; the E-… accession is on the experiment page, target = one of its listed experimental factors--dataset expression_atlas --study E-GEOD-66290 --target infect
TCGA / cBioPortalhumanstudy ids at cbioportal.org/datasets (e.g. brca_tcga_pub); target = a clinical column--dataset tcga --study brca_tcga_pub --target PAM50_SUBTYPE
UCSC Xenahumanpreset cohorts (e.g. brca); target = a phenotype column--dataset xena --preset brca --target PAM50Call_RNAseq
CCLE / DepMaphuman cell linesany gene symbol (CRISPR dependency, regression)--dataset ccle --target SOX10
Metabolomics Workbenchvariousstudy ids ST0000xx at metabolomicsworkbench.org--dataset metabolomics --study ST000009 --target gender

Pick a target with ≥2 well-populated groups (or a continuous column for regression) that is plausibly encoded in the molecular data. Omit --target for a sensible default. Each bootstrap writes a config.json next to the matrices, so omicau run --config <out-dir>/config.json works immediately.

Build omicau-ready data from wet-lab results

Turn instrument/pipeline outputs into the layout below (details and a per-omics value-type table are in Use your own data):

  1. One matrix per modality — a sample × feature CSV/TSV: sample ids in the header row and first column, numbers in the body, blanks/NA for missing. Export the normalized matrix (log2 TPM/CPM for RNA-seq; log2 abundance for proteomics; beta/M for methylation). Do not impute or scale — omicau masks missing values and standardizes inside each CV fold. Genes-as-rows is fine (orientation is auto-detected).
  2. One clinical table — one row per sample: a sample-id column, the outcome, and optional group (patient/animal/plot id) and batch (run/plate) columns.
  3. A config.json naming them (shown below), then omicau run.

Use your own data

The remote data hubs are optional shortcuts; the primary workflow is to point omicau at your own matrices. You provide one file per omic modality, one clinical table with the outcome, and a small config that names them.

Ingestion is deliberately forgiving, so you rarely have to reshape anything:

  • Delimiter is auto-detected (comma, tab, semicolon, or whitespace).
  • Orientation is auto-detected by sample-id overlap — a samples × features matrix and a features × samples matrix (e.g. genes-as-rows expression) both work.
  • Sample names are fuzzily matched across files (whitespace, case, and common batch/aliquot suffixes are normalized).
  • Dirty headers, mixed whitespace, common NA tokens (NA, null, ., ?, …), European decimals, and ±inf are healed automatically.
  • Missing values stay masked — never imputed at ingest.

Minimal layout (any of CSV/TSV; matrices can be either orientation):

mystudy/
  rna.csv          # samples × genes  (or genes × samples — auto-detected)
  protein.csv      # samples × proteins
  clinical.csv     # one row per sample: sample_id, outcome[, patient_id, batch]
  config.json

config.json (JSON shown; .toml is also accepted, and .yaml with the [yaml] extra):

{
  "run_name": "my_study",
  "output_dir": "run",
  "modalities": [
    {"name": "rna",     "path": "rna.csv",     "description": "RNA-seq log-TPM"},
    {"name": "protein", "path": "protein.csv", "description": "Proteomics"}
  ],
  "clinical": {
    "path": "clinical.csv",
    "target": "outcome",
    "sample_id": "sample_id",
    "group": "patient_id",
    "batch": "batch",
    "task": "auto"
  },
  "cv": {"n_splits": 5, "seed": 42},
  "neural": {"enabled": true, "epochs": 60},
  "llm": {"enabled": false}
}
omicau run --config mystudy/config.json --cores 8

Field notes:

  • sample_id — the clinical column holding sample identifiers (omit to use the table's first column / index). The same ids must appear (after normalization) in each modality matrix; samples are intersected across all files.
  • target — the clinical column to predict. task: "auto" infers classification vs regression; set "classification", "regression", or "survival" to force it.
  • time / event (survival only) — for task: "survival", the numeric time-to-event column and the event indicator (1 = event, 0 = right-censored); target is ignored. Scored by Harrell's concordance index.
  • group (optional, recommended) — a column such as patient id, so multiple samples from one patient never split across train and test (leakage-safe CV). Accepts a list of columns (e.g. ["animal", "run"]) that combine into the coarsest independent unit for nested / repeated-measure designs.
  • batch (optional) — a column such as sequencing batch or site; drives the batch-effect diagnostics (and the opt-in cv.batch_blocked cross-site stress test and cv.batch_adjust_sensitivity in-fold sensitivity probe).
  • Only rows with a non-missing target are kept. Paths are resolved relative to the config file, so keep the config next to your matrices.

Start from a working template by generating the mock dataset and editing its config.json + CSV headers to match your files:

omicau bootstrap --dataset mock --out-dir template   # writes a runnable config.json + example CSVs

Add as many modalities as you like (methylation, metabolomics, CNV, …). A single modality is allowed — the audit simply skips the cross-modality comparisons.

Input format for each omics layer

Every modality is a single sample × feature numeric matrix: identifiers live in the header row and the index column, and the body is numbers only (blanks or NA/null/. for missing). omicau does not care which normalization you used — it standardizes inside each cross-validation fold — but the values must be numeric and comparable within a column. The table below is guidance, not a hard schema; the ingester auto-detects delimiter and orientation, so any of these read as-is.

Omics layerMatrix body (value type)Feature id (columns)Notes for omicau
RNA-seq / expressionlog-normalized expression (log2 TPM/FPKM/CPM, or RSEM/log2(x+1))gene symbol / Ensembl / EntrezPrefer a log scale; raw counts work but log-transform heavy-tailed counts first. One value per gene per sample.
Proteomicsnormalized abundance / intensity (usually log2)UniProt id or gene symbolMissing values are common and informative (MNAR) — leave them blank/NA, do not impute. Collapse isoforms to one column per protein.
DNA methylationbeta values in [0,1] or M-valuesIllumina probe id (cg…) or geneBeta and M are both fine as numeric features; don't mix the two in one matrix.
Metabolomicspeak area / concentration (often log-transformed)metabolite name / RefMet / HMDB idKeep below-detection as blank/NA. One column per metabolite.
Copy number (CNV)log2 copy-ratio (continuous) or GISTIC discrete -2…2gene symbolEither continuous or integer levels; keep one convention per matrix.
miRNAlog-normalized expressionmiRBase id (hsa-miR-…)Same shape as RNA-seq.
Somatic mutationbinary 0/1 (mutated) or a mutation countgene symbolPresented as numeric features; a near-constant column (almost all 0) is dropped automatically.

Universal rules across layers:

  • One file per modality, samples aligned by id across all files and the clinical table. Feature names must be unique within a modality; the same name in two modalities is fine (omicau namespaces them as modality::feature).
  • Do not pre-impute or pre-scale. Missing entries stay masked (the neural fuser ignores them; classical models median-impute inside the training fold only), and standardization happens inside each fold — pre-scaling across all samples leaks.
  • Log-transform skewed counts (RNA-seq/miRNA/metabolomics) before ingest; omicau standardizes but does not log for you.
  • Orientation and delimiter are auto-detected, so genes-as-rows or samples-as-rows, CSV or TSV, all load without a flag.

Workflow

flowchart TB
    S0["Multi-modal ingestion — auto-delimiter, orientation, fuzzy sample-name match"]
    S1["Alignment & masking — sample intersection, drop missing endpoints, NaN masks"]
    S2["Provenance SHA-256 — value-level hash of matrices + sample index + target"]
    S3["Cost / runtime estimate — N×P_m, K folds, E epochs, device, cores"]
    S4["Nested group-aware CV — impute + scale + select fitted inside train folds only"]
    S5["Fusion benchmarks — classical concat + masked global-pooling neural network"]
    S6["Leakage-safe XAI — permutation importance on held-out folds"]
    S7["Utility & redundancy audit — marginal gain, CKA, batch/missingness/control gates"]
    S8["Dual reporting — clinical + research dashboard, JSON/CSV assets"]
    S0 --> S1 --> S2 --> S3 --> S4 --> S5 --> S6 --> S7 --> S8

Methodology

1. Flexible ingestion and alignment

Matrices lack a standard layout, so ingestion adapts per file:

  • Delimiter is inferred with csv.Sniffer, falling back to header-frequency counting across the candidate set {tab, comma, semicolon, pipe} and to whitespace splitting.
  • Orientation is resolved by overlap scoring. For a matrix with row labels R and column labels C, and reference sample ids S, the fraction of each axis that matches the reference is overlap(L, S) = |normalize(L) ∩ S| / |L|. If overlap(C, S) > overlap(R, S) the matrix is transposed so samples are rows. This automatically corrects genes-as-rows expression matrices.
  • Fuzzy sample-name normalization strips whitespace, upper-cases, and removes configurable prefix/suffix regexes. The default suffix rule collapses a TCGA aliquot barcode to its patient stem.
  • Numeric self-repair coerces text columns, masks common NA tokens (NA, null, ., ?, …), heals European decimals by comparing the parse yield of the raw text vs a 1.234,5 → 1234.5 repair and keeping the interpretation that recovers more finite values, and maps ±∞ to NaN.
  • Samples are intersected across all modalities and the clinical table; records with a missing endpoint are dropped; duplicate features are namespaced per modality (modality::feature) to block cross-modality collisions.

Missing values are kept as NaN (true masks), never imputed at ingest.

2. Cryptographic provenance

Immediately after alignment, an immutable signature is computed as the SHA-256 of a canonical JSON manifest of the sorted sample index, each modality's sorted feature list and shape, a content digest of each aligned matrix (the numeric values in canonical row/column order), a digest of the encoded target, the target name, and the task. It is tamper-evident at the value level: changing any single measurement — not only which samples or features enter the study — changes the hash, locking asset provenance across the study lifetime.

3. Missingness-bias diagnostics

Tests whether missingness itself carries information (missing-not-at-random):

  • Kruskal-Wallis of each sample's per-modality missing fraction across outcome classes (or Spearman correlation for a continuous target).
  • Chi-squared of a binary "any-missing" indicator against the outcome.
  • Kruskal-Wallis of the missing fraction across batches.

p-values are corrected across all tests with the Benjamini-Hochberg FDR: sort p_(1) ≤ … ≤ p_(n), set p̃_(i) = min_{k≥i} min(1, p_(k)·n/k).

4. Batch-effect and confounding diagnostics

Each modality is projected onto principal components (standardized; mean-imputed for the projection only — never for modeling). Batch structure is quantified by (i) the silhouette of batch labels in PC space, (ii) one-way ANOVA and Kruskal-Wallis of PC1 across batches, and (iii) the fraction of PC1 variance explained by batch, η² = SS_between / SS_total. Batch/outcome confounding — the regime in which a batch effect masquerades as signal — is tested with a chi-squared test + Cramér's V for a categorical target and a one-way ANOVA with η² for a continuous (regression) target. This confounding result, not the PC-variance structure alone, gates the batch-confounded verdict.

5. Leakage-safe preprocessing (nested)

All preprocessing is an sklearn Pipeline fitted inside each training fold only and applied to the held-out fold:

  1. median imputation (train-fold medians),
  2. zero-variance filtering,
  3. standardization z = (x − μ_train) / σ_train,
  4. optional univariate selection (SelectKBest, ANOVA F-test for classification, F-regression for regression).

No validation-fold statistic ever influences training.

6. Group-aware cross-validation

When a group column is present (e.g. patient id with multiple samples), splits keep all of a group's samples on one side, prohibiting identity leakage: StratifiedGroupKFold (classification) or GroupKFold (regression); otherwise StratifiedKFold / KFold. The fold count is clamped so every fold is populated and, for classification, contains both classes. Metrics are computed on pooled out-of-fold predictions.

7. Fusion benchmarks across the integration axis

omicau benchmarks three integration regimes so you can see which one a dataset actually wants: early (feature concatenation, this section), intermediate (the masked global-pooling neural network, §8), and late (stacking — a meta-learner cross-validated over the single-modality out-of-fold predictions, reported as stacking::FUSION; leakage-safe because the base predictions are out-of-fold and the meta-learner is itself cross-validated).

Early fusion concatenates the (namespaced) modality matrices. For each estimator — L2-regularized logistic regression / ridge, random forest, or histogram gradient boosting — omicau cross-validates:

  • each modality alone,
  • the full fusion,
  • each leave-one-modality-out subset.

The marginal gain of modality m is Δ_m = score(FUSION) − score(FUSION∖m). Because k-fold training sets overlap, its significance uses the Nadeau–Bengio corrected resampled t-test (inflating the fold-difference variance by 1/k + 1/(k−1)) rather than a naive paired t-test, whose Type-I error is badly inflated; the same corrected test is applied to the neural fuser's gain.

8. Masked Global Pooling Fusion network (PyTorch)

The custom neural fuser is agnostic to feature counts and to which features are missing per sample. Each modality m has a learned per-feature embedding table E_m ∈ ℝ^{P_m × d}. For a sample with standardized values x and observed mask o ∈ {0,1}^{P_m}:

  • token per feature: t_j = E_m[j] · x_j,
  • masked mean pooling over observed features only: pooled_m = (Σ_j o_j · t_j) / max(1, Σ_j o_j) (a max variant is available),
  • followed by LayerNorm.

Missing features (o_j = 0) contribute nothing to the pooled embedding — no artificial variance is injected. Per-modality embeddings are concatenated and passed to an MLP head (Linear → ReLU → Dropout → Linear). Loss is cross-entropy (classification) or MSE (regression). Standardization statistics are computed from the training fold only (masked over observed entries), keeping the neural path leakage-safe. Early stopping uses an internal split of the training fold. Training is wrapped in an out-of-memory self-repair loop that halves the batch size, clears the device cache, and retries, then falls back to CPU.

9. Leakage-safe feature attribution (XAI)

Primary attribution is permutation importance computed on each held-out validation fold with the model trained only on that fold's training data. The across-fold mean and ±SD are reported, so an unstable (high-SD) attribution is visible rather than hidden. This is unconditional permutation importance, which over-credits correlated predictors and can split or inflate importance among collinear features (Hooker et al. 2021; Strobl et al. 2008) — omic layers are highly collinear, so the report flags this and the ranking should be read as indicative, not causal. The neural fuser additionally exposes a native score from its per-feature embedding norms weighted by observed feature variance.

10. Modality-utility ledger and redundancy

Representational redundancy between modalities X and Y is the linear centered kernel alignment CKA(X, Y) = ‖YᵀX‖_F² / (‖XᵀX‖_F · ‖YᵀY‖_F) ∈ [0, 1] on column-standardized matrices; high CKA with a stronger-alone modality marks a layer as redundant. Each modality receives a verdict — predictive (adds marginal signal), not additive (predictive alone, but its leave-one-out gain is not distinguishable from zero, or it is subsumed by a stronger layer), batch-confounded, or control-like. The predictive verdict is gated on significance, not just a positive point estimate: a modality is only badged as adding signal when its leave-one-out gain clears both a minimum effect size and the Nadeau–Bengio paired test (a positive but non-significant gain reads "not additive", so a noisy gain is never rendered as an established contribution). A layer is called batch-confounded only when its variance is batch-structured and batch is genuinely confounded with the outcome; a batch effect orthogonal to the outcome is harmless and is not flagged (Nygaard et al. 2016).

11. Control baselines

The identical pipeline is run on three corrupted inputs: shuffled target, column-shuffled features, and random Gaussian noise. A well-behaved harness scores at chance on all three. The leakage alarm gates the whole ledger and is CI-aware: it fires when a control's bootstrap 95% CI lower bound clears chance (i.e. the control is significantly above chance), falling back to a task-aware margin (chance + 0.12) when no CI is available. These controls detect target/pipeline leakage (a preprocessing step or the target sneaking into training); they cannot by themselves detect group leakage from a mis-set or missing group column, so the report and the dashboard word the pass as "no target/pipeline leakage" and surface the grouping status (whether samples were keyed to independent subjects) as its own trust-checklist row. AUPRC is reported against its prevalence baseline so improvement over chance is visible.

12. Metrics

Classification: AUROC, AUPRC (average precision), accuracy, balanced accuracy, F1, Matthews correlation. Regression: R², RMSE, MAE, Spearman ρ, Pearson r. Survival: Harrell's concordance index (a ridge-penalised Cox model with in-fold PCA reduction; dependency-light — no scikit-survival/lifelines, so it ships in the base install). All guard against degenerate folds and return NaN rather than raising. The primary metric is AUROC (classification), R² (regression), or the C-index (survival).

12b. Task types and optional capabilities

  • Survival / time-to-event (task: "survival") — Cox + C-index under the same leakage-safe group-aware CV, controls, and bootstrap CIs. Set clinical.time and clinical.event (not target); the event indicator is parsed tolerantly (numeric 1/0, word labels DECEASED/LIVING or Dead/Alive, cBioPortal 1:DECEASED, booleans, yes/no), and a column that parses to zero events fails loudly rather than silently treating every subject as censored.
  • Single-modality runs are a first-class honesty check: with one layer there is no fusion to benchmark, so the report drops the inert fusion/redundancy panels and verifies the single layer's signal is real and leakage-free.
  • Composite grouping (group: [ ... ]) for nested / repeated-measure designs.
  • Cross-organism data and the Expression Atlas hub, with an opt-in --normalization tmm|median_of_ratios (whole-matrix, caveated; log2cpm default).
  • Batch-adjustment sensitivity probe (cv.batch_adjust_sensitivity) — an opt-in, in-fold "does the signal survive batch removal?" check, hard-gated off when batch is confounded with the outcome. It never corrects your data ("diagnose, don't correct").

All of the above are exposed in both the CLI/config and the no-code web UI.

Every model's primary metric carries a group-level bootstrap 95% confidence interval (resampling whole patient groups, consistent with the group-aware CV); the per-fold spread is reported only as "fold dispersion," since correlated folds bias it low. For binary classification the report adds calibration — a reliability curve, Brier score, and expected calibration error — flagging that class_weight="balanced" miscalibrates the predicted probabilities by construction.

13. Fairness, generalization, and governance

  • Per-site / subgroup performance — the best model's pooled out-of-fold predictions are re-scored within each level of the batch/site column, since a global metric can hide disparities across sites or subgroups (Obermeyer et al., Science 2019). Pure re-aggregation, no retraining. Each per-stratum metric and the max-minus-min gap carry a bootstrap 95% CI, and strata below n=20 are shown but not scored — so the "investigate site/subgroup bias" call fires only when the gap interval excludes a negligible difference, not on a max-minus-min point estimate that small-sample noise inflates.
  • Cross-site stress test (opt-in cv.batch_blocked) — the reference fusion is additionally cross-validated with folds blocked on batch (leave-one-batch-out), giving an honest new-batch generalization estimate and an optimism gap against the standard CV.
  • Governance — a DOME methods block (Data / Optimization / Model / Evaluation, with explicit limitations; Walsh et al., Nat Methods 2021) is written into audit.json, and a MODEL_CARD.md (intended use = RUO audit, data, metrics with CIs, caveats; Mitchell et al., FAccT 2019) is emitted alongside the report.

14. Pre-flight cost estimation

Before heavy loops, wall-time is estimated from N × P_m per modality, the fold count K, neural epochs E, the device (MPS/CUDA/CPU), and the core count C. A single live RandomForest fit calibrates the per-fit cost on the actual machine; it is scaled by the model/fold counts, and the neural cost is modelled from the epoch budget and feature footprint. The estimate is deliberately conservative so HPC allocations are safe.


Plain-language verdict (optional AI endpoint)

omicau always writes a deterministic, rule-based verdict. On top of that you can point it at an LLM to turn the numbers into a short plain-English verdict. This is optional and additive — skip it and nothing changes. Only aggregate, de-identified diagnostics are sent (scores, CIs, flags, ledger — never raw matrices, sample ids, or clinical values); the key is used for one call and never stored, logged, or written to any output or the provenance hash.

Install the SDK tier once (pip install ".[llm]" — Anthropic + OpenAI SDKs; the OpenAI SDK also drives ChatGPT, Gemini, and local servers), then add an llm block to the config. The key is read from a named environment variable, so the config stores only the variable name:

providermodel exampleEndpoint / key
anthropic (Claude)claude-sonnet-5export ANTHROPIC_API_KEY=…
openai (ChatGPT)gpt-4oexport OPENAI_API_KEY=…, api_key_env: "OPENAI_API_KEY"
gemini (Google)gemini-2.0-flashexport GEMINI_API_KEY=…, api_key_env: "GEMINI_API_KEY" (base URL set for you)
local (Ollama / LM Studio / vLLM)llama3.1no key; base_url: "http://localhost:11434/v1"
"llm": {
  "enabled": true, "provider": "anthropic", "model": "claude-sonnet-5",
  "api_key_env": "ANTHROPIC_API_KEY", "base_url": null
}
export ANTHROPIC_API_KEY=sk-...              # cloud providers read the named env var
omicau run --config my_study/config.json --llm     # --llm / --no-llm overrides the config

Fully offline example — a local model via Ollama (ollama pull llama3.1), no key, nothing leaves the machine:

"llm": { "enabled": true, "provider": "local", "model": "llama3.1",
         "base_url": "http://localhost:11434/v1" }

In the web UI, pick the provider and paste the key into the verdict step (held in memory for that one run). If the SDK, key, or network is missing, omicau silently falls back to the rule-based verdict — the audit never fails because of it.

Reporting

  • Dashboard (report.html): a single, fully offline self-contained file (Plotly bundled inline; fonts embedded as data URIs — zero network calls). Humanist sans typography (IBM Plex Sans + IBM Plex Mono), a color-blind-safe Okabe-Ito palette (cobalt #0072B2 = standard, vermillion #D55E00 = warning), an executive tab for PIs/clinicians and a research tab for computational biologists, five interactive figures, and tables that are sortable, text-filterable, and CSV/TSV-exportable via dependency-free vanilla JavaScript.
  • Machine-readable: audit.json (full state, including a DOME methods block), model_metrics.csv (per model: primary metric + AUPRC + balanced accuracy for classification, with the 95% CI), modality_ledger.csv, missingness_tests.csv, and MODEL_CARD.md (a research-use-only model card).

Validation

omicau was re-run end-to-end on eight public datasets spanning model and non-model organisms and every task type, and each result was checked against the published finding. Numbers below are the fresh runs (group-aware CV, feature attribution off for speed); the shuffled-target control and the CI-aware leakage gate are shown so the honesty behaviour is visible, not just the headline score.

OrganismUse caseTarget · layersnMetricScore95% CIShuffled controlLeakage gate
Human (BRCA)clinicalPAM50 subtype · RNA+CNV493AUROC0.9530.94–0.960.49clean
Human (BRCA)clinicalPAM50 subtype · RNA956AUROC0.9880.98–0.990.50clean
Human (BRCA)clinicalOverall survival · RNA+CNV499C-index0.6560.55–0.760.54clean
Cell linesdrug / biotechSOX10 dependency · expression11030.7230.66–0.77−0.07clean
MousebiologyDlx3-KO genotype · RNA12AUROC0.7780.46–1.000.53clean (wide CI)
ZebrafishbiologyIL-KO genotype · RNA20AUROC1.0000.69flagged (n≪p)
YeastbiotechNaCl osmostress · RNA15AUROC1.0000.42clean
Arabidopsisnon-model plantBotrytis infection · RNA12AUROC1.0000.89flagged (n≪p)

Each score matches the literature: PAM50 subtypes are transcriptionally defined (Parker et al. 2009; TCGA 2012), expression predicts SOX10 dependency in melanoma lineage (DepMap; Broad self-expression ρ ≈ −0.82), the Dlx3 knockout is a moderate osteogenic signature (Duverger et al. 2014, so AUROC 0.778 with a wide CI at n=12 is the honest read), the yeast HOG/ESR osmostress program is a large clean response (Babazadeh et al. 2017), and Botrytis infection reprograms roughly a third of the Arabidopsis genome (Liu et al. 2015). Where n ≪ features (zebrafish, Arabidopsis), a perfect AUROC is a small-sample artifact, and the control baselines catch it — the shuffled-target control also beats chance (0.69, 0.89), so omicau flags leakage and refuses to certify the score rather than reporting a misleading 1.000. Where the cohort is adequate (human, yeast) or the effect is moderate (mouse), the controls sit at chance and the confidence interval tells the honest story.

Earlier releases were additionally validated on non-hub multi-omics benchmarks: ROSMAP and BRCA (MOGONET; three omics — fusion AUROC ≈ 0.80 and ≈ 0.95, with methylation and miRNA correctly flagged redundant with RNA), a GEO lung tumour-vs-normal microarray (AUROC ≈ 0.99), and a whole-blood methylation epigenetic clock (regression, R² ≈ 0.88, MAE ≈ 3.6 years). Across all of these the shuffled controls scored at chance and no leakage was flagged.


Data hubs

All clients run structural gates (numeric-only, non-finite healing, constant-feature dropping, sample-extension matching) and use retry-with-jitter. Network access is optional and isolated; the core is unaffected if it is absent.

Only hubs whose live connection was verified are shipped. Each client was probed directly against its real endpoint (the connection column reflects that check).

ClientSourceModalities → targetConnection
tcgacBioPortal public REST APImRNA + copy-number + merged sample/patient clinicalverified (laml_tcga: sex from RNA+CNV, AUROC ≈ 0.90)
ccleDepMap 24Q4 (figshare)RNA-seq → CRISPR gene-effect dependencyverified (1103 lines × 19k genes; SOX10 R² ≈ 0.74)
xenaUCSC Xena hubs (no auth)RNA-seq / methylation / CNV / protein + phenotypeverified (TCGA-BRCA PAM50, 1247 samples)
openpbtaPublic AWS S3 (anonymous)putative-fusion matrix + histologiesverified (open-targets/v15 S3 listing + TSV headers)
metabolomics_workbenchMetabolomics Workbench RESTmetabolomics + study factorsverified (ST000009: gender AUROC ≈ 0.88)
cptaccptac PyPI packagematched proteomics + transcriptomicsneeds the cptac package (a build toolchain; no Windows/py3.12 wheel) — verified by docs only
allofusAll of Us Researcher WorkbenchWGS / proteomics / RNA-seqWorkbench-only by design; off-platform it raises a clear error (cannot be externally connected)

gdsc and linkedomics were evaluated and dropped: their public download endpoints do not respond to a scripted client (HTTP 410 and 403 respectively).

The All of Us client runs only inside the secure Researcher Workbench and reads the managed WORKSPACE_CDR / WORKSPACE_BUCKET / GOOGLE_PROJECT variables; data cannot be exported and off-platform sessions raise a clear error. No participant-level data is ever transmitted off-platform.

Reproduce any row of the Validation table in two commands, e.g. omicau bootstrap --dataset expression_atlas --study E-GEOD-66290 --target infect --out-dir d then omicau run --config d/config.json.


Command-line and configuration reference

Commands

CommandKey optionsWhat it does
omicau check-envCPU/GPU, folder access, and which optional tiers ([data], [ui], [llm]) are installed.
omicau bootstrap--dataset*, --out-dir*, --study / --target / --preset / --cancer, --task, --seed, --normalizationAssemble a runnable dataset (matrices + clinical.csv + config.json) from the mock generator or a public hub.
omicau run--config*, --cores/--threads, --device {auto,cpu,cuda,mps}, --llm/--no-llm, --deterministic, --out-dirFull audit → report.html + audit.json + CSVs + MODEL_CARD.md.
omicau verify--config, --audit, --expected <hash>Recompute the provenance SHA-256 and compare (exit 1 on drift).
omicau ui--port, --host, --no-browserLocal no-code wizard (needs [ui]).

* required. --cores/--threads honor cgroup limits; --deterministic turns on strict PyTorch determinism (torch.use_deterministic_algorithms + CUBLAS_WORKSPACE_CONFIG).

Configuration

A config is JSON, TOML, or YAML ([yaml] extra). Unknown keys warn and are ignored, every omitted key takes the default below, and the fully-resolved config is written back into audit.json. Top-level: run_name, output_dir, organism, seed (the master seed — every component derives from it), modalities, plus these blocks:

KeyDefaultPurpose
clinical.target"target"Outcome column (classification / regression).
clinical.sample_idindexSample-id column (omit → table index / first column).
clinical.groupnullIndependent unit for leakage-safe CV; a list combines columns to the coarsest unit.
clinical.batchnullBatch/site column → hygiene diagnostics + opt-in stress tests.
clinical.task"auto"auto | classification | regression | survival.
clinical.time / clinical.eventnullSurvival only: numeric time + event indicator (tolerant parsing).
cv.n_splits5CV folds (auto-clamped so every fold is populated).
cv.n_bootstrap1000Group-level bootstrap resamples for the 95% CI.
cv.batch_blockedfalseOpt-in leave-one-batch-out cross-site stress test.
cv.batch_adjust_sensitivityfalseOpt-in in-fold batch-centering probe (auto-off if batch is confounded).
classical.models["linear","random_forest"]Estimators; also gradient_boosting.
classical.max_features2000In-fold SelectKBest cap (null = keep all).
neural.enabled / neural.epochs / neural.poolingtrue / 60 / "mean"Fusion network; pooling mean | max.
xai.enabledtruePermutation-importance attribution — set false to skip on high-dimensional data (much faster).
controls.enabledtrueShuffled-target / shuffled-feature / random-noise baselines.
compute.cores / compute.device / compute.deterministicnull / "auto" / falseWorkers, backend, strict determinism.
normalization.preset"none"none (ids verbatim) | tcga (uppercase + collapse aliquot barcode).
llm.enabled / llm.provider / llm.modelfalse / "anthropic" / "claude-sonnet-5"Optional plain-language verdict (key is never part of the config).

Output files

FileContents
report.htmlSelf-contained interactive dashboard (executive + research tabs; opens with no network).
audit.jsonFull run state — dataset, models, utility ledger, diagnostics, calibration, subgroups, DOME block, resolved config, environment, provenance hash.
model_metrics.csvPer model: primary metric (+ AUPRC + balanced accuracy for classification), 95% CI, features, folds.
modality_ledger.csvPer layer: standalone score, marginal gain + p-value, redundancy CKA, verdict.
missingness_tests.csvMissingness-bias tests with BH-adjusted p-values.
MODEL_CARD.mdResearch-use-only model card (intended use, data, metrics, caveats).
runtime_log.txtWall-time per step with a hostname/device/cores tag.

Python API

The CLI is a thin wrapper — the identical audit runs in-process, returning the full audit dict with asset paths under _assets:

from omicau.config import OmicauConfig
from omicau.cli import run_audit

cfg = OmicauConfig.from_file("mystudy/config.json")
audit = run_audit(cfg, cores=8, device="auto", llm=False)
print(audit["_assets"]["html"], audit["_assets"]["json"])   # report.html + audit.json
print(audit["utility"]["best_model"])                       # results as plain dicts

On a cluster (SLURM)

The core is headless and honors cgroup CPU limits; no GUI, network, or interactive prompt is required.

#SBATCH -c 16
#SBATCH --mem=32G
omicau run --config study/config.json --cores $SLURM_CPUS_PER_TASK --device cpu --deterministic

Install torch from the CPU wheel first (see Installation) to avoid the ~2.5 GB CUDA download on CPU-only nodes, and set xai.enabled=false in the config for very high-dimensional matrices where feature attribution dominates the runtime.


Software versions

Frozen Python stack (development + test reference)

PackageVersionPackageVersion
Python3.12.10scikit-learn1.9.0
numpy2.5.0torch2.12.1 (CPU)
pandas3.0.3plotly6.8.0
scipy1.18.0click8.4.1
jinja23.1.6tqdm4.68.3
requests2.34.2pytest9.1.1

Optional tiers (pinned floors in pyproject.toml): anthropic ≥ 0.39, openai ≥ 1.0, cptac ≥ 1.5 (tested against 1.5.14), google-cloud-storage ≥ 2.10, pyyaml ≥ 6.0. The optional plain-language verdict is provider-agnostic: Claude (Anthropic Messages API, default claude-sonnet-5), ChatGPT (OpenAI), Gemini (Google's OpenAI-compatible endpoint), and local models (Ollama / LM Studio / vLLM) — the openai SDK drives the last three. Only the audit's numeric diagnostics are sent (never raw data); the key is used for one call and never stored.

Upstream database / atlas releases (pinned)

ResourceRelease pinnedRoute
cBioPortalREST API v3 (public); example study laml_tcgahttps://www.cbioportal.org/api
DepMap / CCLEDepMap Public 24Q4 (figshare article 27993248)figshare ndownloader
CPTACvia cptac 1.5.14 (Zenodo-hosted, on-demand)cptac cancer classes
OpenPedCanrelease v15 (open-targets/v15)public S3 d3b-openaccess-us-east-1-prd-pbta
OpenPBTArelease-v23-20230115public S3 (same bucket)
GDSC (optional target)release 8.4 (24Jul22)
PRISM (optional target)Repurposing 19Q4figshare article 9393293
All of UsCDR v7/v8 Workbench variable conventionsin-Workbench BigQuery + GCS

Open-access release identifiers move over time; each client docstrings its verified-as-of date, and the endpoints are best-effort — re-verify before production.


Reproducibility log

  • Determinism: every stochastic step is seeded (numpy, random, torch.manual_seed; per-fold seeds derive from the master seed). Cross-validation splits depend only on the target, groups, and seed, so leave-one-out and single-vs-fusion comparisons are exactly paired. Strict bit-level determinism (torch.use_deterministic_algorithms + CUBLAS_WORKSPACE_CONFIG) is opt-in via --deterministic / compute.deterministic; it is off by default because some ops lack deterministic kernels (enabled with warn_only so it never hard-fails).
  • Provenance: the value-level SHA-256 of the aligned matrices, sample index, features, and target is written to audit.json and re-checkable with omicau verify (exit 1 on any drift).
  • Environment capture: audit.json → environment records the Python and platform plus the versions of the packages that actually drive the numbers — numpy, scikit-learn, scipy, pandas, plotly, jinja2, and torch — so a run's results can be tied to the exact stack that produced them; runtime_log.txt records the wall-time of every step with a device tag (hostname/device/cores).
  • Built and tested on: Python 3.12.10, Windows 11 (10.0.26200), x86-64 (AMD64), CPU-only torch. The package is cross-platform (pathlib throughout, defensive newline="" I/O, no shell invocation) and headless-HPC ready (--cores/--threads honor cgroup limits; no interactive prompts).

Decoupling and HPC

The omicau core is 100% functional with no internet access, no LLM connection, and no multi-agent framework. The LLM interpretation layer (interpretation/llm_summary.py) is an optional plugin; when absent it falls back to a deterministic rule-based summary filling the identical JSON schema, so the report never breaks. Remote data hubs are optional and isolated behind lazy imports. Thread/worker counts are set explicitly via --cores / --threads; the PyTorch device is chosen with --device (MPS/CUDA/CPU, auto by default).


Troubleshooting

SymptomCause / fix
error: externally-managed-environment on pip installDebian/Ubuntu 24.04+ system Python (PEP 668). Use pipx install omicau or a virtual environment — not system pip.
omicau: command not found after pip installThe script dir isn't on PATH. Use pipx (runs pipx ensurepath), or call python -m omicau.
A feature errors asking for an extra (e.g. [ui], [data], [llm])Install it: pip install "omicau[ui]", or on a pipx install pipx inject omicau fastapi uvicorn python-multipart. The error prints the exact command for your environment.
The run seems to "hang" after "Checking for batch effects"On high-dimensional data, permutation-importance attribution re-scores per feature and is the bottleneck. Set xai.enabled=false in the config — the audit finishes in seconds with full metrics/CIs/controls/calibration, just no feature ranking.
Linux install pulls ~2.5 GB of NVIDIA librariesThat's torch's default CUDA wheel. On a CPU-only box install torch from the CPU index first (see Installation), then pip install omicau.
All modalities were dropped during alignmentSample ids don't overlap across your files. Check that the same ids appear in every matrix and the clinical table; for raw TCGA barcodes set normalization.preset: "tcga".
Survival run: Target column 'target' not foundA survival config uses clinical.time + clinical.event, not target. Set task: "survival" with both columns.
Survival run reports no events / a degenerate C-indexThe event column parses to all-censored. omicau now errors loudly on this; check the encoding (numeric 1/0, DECEASED/LIVING, or cBioPortal 1:DECEASED are all handled).
AUROC 1.000 but the report flags leakage / a very wide CINot a bug — this is the honesty gate. With few samples relative to features (n ≪ p) a perfect score is a small-sample artifact; the shuffled control also beats chance, so omicau refuses to certify it. Collect more independent samples.
UnicodeEncodeError in the terminal (Windows)The CLI reconfigures stdout to UTF-8 at startup; if a wrapper suppresses that, set PYTHONUTF8=1.

License

MIT.