Jev Voice

September 20, 2026 · View on GitHub

A macOS menu-bar app that turns spoken commands into actions using TypeSafe AI's Jev — a System One model that returns typed decisions with probabilities instead of generated text.

Press ⌥Space, say something, and Jev decides what to do:

  • "open chrome and go to google.com"
  • "quit spotify"
  • "type hello world"
  • "set volume to 30"
  • "search for swift concurrency"
  • "maximize chrome"
  • "full screen safari"
  • "restore chrome"

App names support common aliases such as “chrome”, “code”, “cmux”, and “settings”. The app can speak short execution replies using the selected macOS voice, and Settings can enable Ask before running commands for a voice-confirmation step.

Mixed commands combine deterministic local actions with Jev-guided UI work: “open Devin and start a new session” = local open + Jev-guided click. The app keeps those typed decisions in order, carries the opened app forward as the UI task target, and asks for confirmation before actions that sound hard to undo.

Speech is transcribed on-device (SFSpeechRecognizer). Only the transcript is sent to the Jev API.

How it works

How it decides

Jev first interprets deterministic local commands, then picks the next action from the live accessibility elements on screen. DeepSeek Flash is an optional fallback for open-ended tasks or screens without accessible controls.

Computer use settings let you choose the Jev step planner or DeepSeek Flash, and control DeepSeek thinking (off, low, or high). In Jev mode, the DeepSeek key is optional and is only used for fallback.

Learned shortcuts

After a successful UI task, Jev remembers the app control that worked for that goal. Similar requests can use those learned labels as additional context without bypassing the live-screen decision.

Chrome DevTools fallback

When a Chromium or Electron window has a thin Accessibility tree, Jev can read visible web controls through Chrome DevTools Protocol and use those controls for clicks and text entry. Enable remote debugging with open -a "Google Chrome" --args --remote-debugging-port=9222, then enable the Chrome DevTools option in Computer use settings.

Written replies

Requests such as “write a note apologising for the delay” are composed by a generative text model rather than typing the spoken words verbatim. DeepSeek is used by default, or a local oMLX OpenAI-compatible endpoint can be selected in Computer use settings. Jev classifies the request; it does not generate the prose. Generated text is previewed before typing by default.

microphone ──> SFSpeechRecognizer (on-device) ──> transcript

                                          ClauseSplitter (multi-verb clauses)

                              POST /v1/systemone  ──> Jev (typed answers)
                                            per clause: action?   target_app?
                                            system_action? mentions_url?
                                            refers_to_frontmost? destructive?

                                 Decision + SlotExtractor (url/query/text/%)

                              Executor (NSWorkspace / CGEvent / osascript / pmset)

Jev never produces free text — each clause is one systemOne call with a state payload (clause, full_transcript, frontmost_app, installed_apps) and six questions:

questiontypeshape
actionchoiceone of openApp, closeApp, openURL, webSearch, dictate, uiTask, system, none
target_appchoiceinstalled app names (≤254) + none
system_actionchoicevolumeSet, mute, lockScreen, screenshot, …
mentions_urlnoulprobability the clause names a website
refers_to_frontmostnoul"quit it" → the frontmost app
composesnoulprobability the clause asks Jev to write the wording
destructivenoulprobability the clause is hard to undo

Arguments (URLs, queries, dictation text, percentages) can't come from Jev — it's decision-only — so they're extracted deterministically by SlotExtractor heuristics.

Most commands execute locally without a network call. Commands Jev handles remain typed decisions; low-confidence or ambiguous commands ask for confirmation before anything executes.

Install with Homebrew

This repository doubles as a Homebrew tap (see Casks/jev-voice.rb):

brew tap chris-wozniczek/jev-voice https://github.com/chris-wozniczek/jev-voice-control
brew install --cask jev-voice

Homebrew 7+ refuses untrusted third-party taps; if prompted, run brew trust chris-wozniczek/jev-voice first.

Releases are produced by .github/workflows/release.yml: bump CFBundleShortVersionString in Info.plist, push a matching vX.Y.Z tag, and the workflow builds Jev-Voice-X.Y.Z.zip, publishes a GitHub release, and commits the new version/sha256 into the cask.

Build from source

make app    # builds build/Jev Voice.app (ad-hoc signed)
make run    # launches it
make test   # unit tests (pure logic: clause splitting, slot extraction, codecs)
make dist   # zips the app into build/Jev-Voice-<version>.zip and prints its sha256

Then set your API key in the popover's Settings (gear icon), or:

defaults write com.chriswozniczek.jevvoice typesafeAPIKey <key>
# or export TYPESAFE_API_KEY=... before launching

Permissions

  • Microphone — hear commands
  • Speech Recognition — transcribe (on-device when supported)
  • Accessibility — required for the Cmd+V paste used by dictation
  • Automation — required for osascript volume/brightness actions

Teaching Jev app shortcuts

Optional app shortcuts provide a fast path for common actions such as opening a new session or tab. They are stored in:

~/Library/Application Support/Jev Voice/app-actions.json

The file contains an actions array. Each entry has an app name (or *), an optional bundleId, a display name, matching phrases, and key steps:

{
  "actions": [
    {
      "app": "Devin",
      "name": "new session",
      "phrases": ["new session"],
      "steps": [
        {"kind": "key", "key": "n", "modifiers": ["command"]}
      ]
    }
  ]
}

Disable or edit these shortcuts in Settings. Everything still works through the generic observe → Jev → act → verify loop without this file.

Native Accessibility can read the target window's tree in-process for lower latency and uses Cua when the tree is too thin. Disable it in Computer use settings to use the Cua observer exclusively.

When an app exposes almost no accessible controls, the optional OCR fallback uses Apple Vision locally to read visible labels and click their screen coordinates. It requires Screen Recording permission and can be disabled in Settings › Computer use.

Debugging

Inspect recent Jev Voice logs with:

log show --last 5m --predicate 'subsystem == "com.chriswozniczek.jevvoice"' --info

The step list includes a Clear button for removing completed local steps.

Limitations

  • Jev is text-only and never generates strings, so URLs/queries/dictation come from heuristics over the transcript, not from the model.
  • System actions are AppleScript/pmset based; brightness key codes may vary on some hardware.
  • App launching scans standard macOS application directories and running regular applications, with aliases and user-defined aliases from Settings.

Privacy

  • Speech recognition runs on-device (requiresOnDeviceRecognition when supported).
  • Only the transcript, frontmost app name, and installed app list are sent to api.typesafe.ai. Your API key stays in local UserDefaults.