hush

September 20, 2026 · View on GitHub

Issue and pull request triage that stays quiet when it isn't sure.

Marketplace npm Tests License: MIT

A GitHub Action that reads every new issue and labels it — but only when it can say how sure it is. When it can't, it does nothing and tells you why.

hush answering four questions about an issue and acting on the two it was sure of

■ "Crash on save when filename is very long"
  bug            100% ≥ 80%, confidence 100%     → applied `bug`
  duplicate       97% ≥ 85%                      → applied `possible-duplicate`
  spam             3% < 90%                      → stayed quiet
  needs info      14% < 85%                      → stayed quiet

■ "🔥 BUY CHEAP FOLLOWERS 🔥"
  spam            99% ≥ 90%                      → applied `spam`
  label          best fit was "none" (100%)      → stayed quiet

Real output, 202–530 ms per issue.

And the part that matters, from its own first issue — a report that could reasonably be a bug or a documentation problem:

label | stayed quiet | bug 72% / confidence 63% — below 80% / 60%

It had an answer. It wasn't sure enough. So it said nothing.

Why another triage bot

Because the others talk when they shouldn't.

LLM triage bots return prose with no calibrated notion of certainty, so a 55% hunch and a 99% read look identical coming out. You get confident-sounding labels on issues the model never understood, and after the third wrong one you turn it off.

hush runs on Jev, a decision model that returns a probability instead of a sentence. Every verdict comes with the number that produced it, and every threshold is yours:

label-threshold: 0.80        # how likely the label must be
label-confidence-threshold: 0.60  # and how much the model must trust its own read
spam-threshold: 0.90
needs-info-threshold: 0.85
duplicate-threshold: 0.85

Below the line, it abstains. Silence is the default behaviour, not the failure mode.

It is also cheap enough to leave on: one request per issue, around $0.00002. Ten thousand issues cost about twenty cents.

Try it on a backlog you already have

Before adding anything to a repository, point it at one:

export TYPESAFE_API_KEY=...          # console.typesafe.ai/keys
npx hush-triage owner/repo
remotion-dev/remotion · 6 open issues · nothing will be written

#11474 Codemods: It always creates a full copy of the enti…  needs-more-info, bug
      · spam       spam 4% < 90%
      · duplicate  duplicate 10% < 85%
      ✓ needs_info needs info 90% ≥ 85%
      ✓ label      bug 92% ≥ 80%, confidence 89% ≥ 60%
#11469 Video matting: Allow setting default…                 feature
      · spam       spam 3% < 90%
      · duplicate  duplicate 22% < 85%
      · needs_info needs info 79% < 85%
      ✓ label      feature 100% ≥ 80%, confidence 100% ≥ 60%

6 of 6 would get a label · 0 left alone
237 ms average · \$0.00033 total

Those two rows are the whole design: 90% cleared the needs-info threshold, 79% did not, and the second issue was left alone on that question.

It reads issues and writes nothing — there is no --apply. Use --limit, --labels with your own taxonomy, and --json to pipe it somewhere.

Install

# .github/workflows/triage.yml
name: triage
on:
  issues:
    types: [opened, edited, reopened]

permissions:
  issues: write

jobs:
  hush:
    runs-on: ubuntu-latest
    steps:
      - uses: emreozyoruk/hush@v1
        with:
          typesafe-api-key: ${{ secrets.TYPESAFE_API_KEY }}
          apply: false   # watch it first

Get a key at console.typesafe.ai/keys and add it as a repository secret.

Start with apply: false. hush writes a table into the job summary of every run showing what it would have done. Read a week of those, move your thresholds, then set apply: true.

Is 0.80 a good threshold?

Measured, not asserted. 800 closed issues from fourteen repositories across four ecosystems, each carrying exactly one label its own maintainers applied — then hush is asked the same question and the answers are compared.

thresholdacts onprecisionstays quiet
0.70701/80076%12%
0.80 (default)649/80078%19%
0.90573/80080%28%
0.95521/80083%35%

Precision is not evenly spread, and the shape of that is the useful part:

bugfeaturedocsquestion
97%86%82%40%

Almost every disagreement is one class. Repositories don't use question as a category — they use it as a workflow state, meaning "this is support, not our bug tracker". Several issues their maintainers labelled question are textbook defect reports with version numbers and reproduction steps.

Excluding question: 88% across 508 decisions. The label that gets applied most, bug, is the one it is best at — 97%.

And the default travels. Excluding question, by ecosystem:

JavaScriptRustGoPython
90%90%89%85%

Four ecosystems, five points apart. Whatever the default is doing, it is not fitted to one community's way of writing issues.

Two things that matter for configuring it. The curve is shallow — moving from 0.80 to 0.95 buys five points of precision and costs sixteen points of coverage. And thresholds are the blunt instrument while the label descriptions are the sharp one: sharpening the four defaults to name what each label is not moved the benchmark three points and bug to 97%.

Method, dataset and raw results: bench/.

Your labels, your words

The default set is bug, feature, docs, question. Replace it with your own — the description is what the model judges against, so write it the way you would explain the label to a new maintainer:

- uses: emreozyoruk/hush@v1
  with:
    typesafe-api-key: ${{ secrets.TYPESAFE_API_KEY }}
    apply: true
    comment: true
    labels: |
      {
        "bug": "A defect: the library does something other than what the docs say.",
        "performance": "It works, but it is too slow or uses too much memory.",
        "platform/windows": "Specific to Windows; does not reproduce on Linux or macOS.",
        "good first issue": "Small, well-understood, and does not need project context."
      }

A label is only ever applied if it already exists in the repository. hush never creates labels, never removes one, and never touches a label a human added.

What it decides

questiontypewhat it does
labelone of yours, or noneApplies the category when both the probability and the model's confidence clear your thresholds.
spamyes/noApplies spam.
needs more infoyes/noApplies needs-more-info when a maintainer would have to ask before acting.
duplicateyes/noCompares against the 40 most recent open issues and applies possible-duplicate.

All four travel in one request, which is why triage costs a fraction of a cent and finishes before the page reloads.

Pull requests

Point it at pull_request and it asks a different set of questions — the ones a reviewer has before opening the diff:

on:
  pull_request_target:
    types: [opened, edited, reopened, ready_for_review]

permissions:
  pull-requests: write
questionappliesasks
kindfix feature docs refactor chore testWhat sort of change is this?
riskyneeds-careful-reviewDoes it touch migrations, auth, permissions, payments, deletion paths, public API or deploy config?
untestedneeds-testsDid behaviour change with no test in the diff?
undescribedneeds-descriptionWould a reviewer have to read the diff to learn what it does?

It sees the title, the description, the file list with per-file line counts, the commit count and whether this is the author's first contribution — not the diff body. On three real pull requests from supabase, prisma and vite it called the kind correctly every time and raised none of the three flags, which is the right answer for well-tested changes from mature repositories.

Thresholds: risky-threshold (0.80), untested-threshold (0.85), undescribed-threshold (0.85). The kind choice reuses label-threshold and label-confidence-threshold.

Inputs

inputdefault
typesafe-api-keyRequired.
github-token${{ github.token }}Needs issues: write to apply anything.
applyfalseAct, rather than only report.
commentfalseExplain the labels in a comment.
labelsthe four aboveJSON object of name → meaning.
check-duplicatestrueCompare against other open issues.
label-threshold0.80
label-confidence-threshold0.60
spam-threshold0.90
needs-info-threshold0.85
duplicate-threshold0.85

What it will not do

  • Close, lock, delete or edit anything. It labels, and optionally comments.
  • Apply a label that does not already exist in your repository.
  • Remove or overwrite a label a person set.
  • Act at all while apply is false.

Where it is unsure, it leaves the issue exactly as it found it. That is the entire idea.

Development

npm test     # the decision layer is pure; the thresholds are tested offline

No build step and no dependencies — the action runs the files in src/ directly on Node 20.

Licence

MIT