Jev training stack (two repos)

September 19, 2026 · View on GitHub

Use both repos in order when building a dataset for a locally hosted model.

                    ┌─────────────────────────────────────┐
                    │  Raw corpus (JSONL, transcripts, …) │
                    └──────────────────┬──────────────────┘

                    ┌──────────────────▼──────────────────┐
                    │  jev-curate  —  keep / drop          │
                    │  Output: curated.jsonl               │
                    └──────────────────┬──────────────────┘

                    ┌──────────────────▼──────────────────┐
                    │  jev-triage  —  accept / teacher / human │
                    │  Output: queues + soft_labels.jsonl  │
                    └──────────────────┬──────────────────┘

              ┌────────────────────────┼────────────────────────┐
              ▼                        ▼                        ▼
        accepted (cheap)         teacher (LLM/VLM)         human (annotators)
              │                        │                        │
              └────────────────────────┼────────────────────────┘

                    ┌─────────────────────────────────────┐
                    │  Train — targets = REAL outcomes       │
                    │  (not “Jev was confident”)             │
                    └─────────────────────────────────────┘
RepoQuestion it answersSuccess in one line
jev-curate“Should this row exist in our dataset at all?”Curated split audited; bad rows in rejected.jsonl
jev-triage“How much labeling money do we spend on this row?”Queues processed; model eval on human/outcome labels

Shared rule: filter and route with Jev; learn from real outcomes.