I built this for my “run your entire computer with your voice” video. You talk to your Mac and it acts: it opens apps, plays music on Spotify, runs a terminal command, reads what’s on screen out loud, controls Premiere or starts an OBS recording.
It’s meant to be taken apart. The brain is OpenAI’s gpt-realtime-2, the hands are agent-desktop, and this repo is the glue: a handful of small voice tools. Clone it, hand it to your coding agent, and say “add my app.”
What it does
- Opens apps by voice.
- Plays music on Spotify.
- Opens Terminal and your command-line coding agent.
- Reads what’s on your screen back to you.
- Starts an OBS recording, and plays, pauses or nudges the Premiere playhead by a frame.
- Lets you add any app as a new tool, each one a short Python function of about 15 lines.
How it works
- You speak, using push-to-talk, a hold-to-talk hotkey or a wake word.
- OpenAI’s gpt-realtime-2 hears the request and picks which tool fits.
- The tool runs from actions.py, clicking and reading apps through agent-desktop or using AppleScript for apps that support it.
- The app does the thing and the model speaks a confirmation back to you.
What you need
- A Mac
- Node, to install agent-desktop
- Python 3.10 or newer
- An OpenAI API key with Realtime access
- Accessibility permission for agent-desktop
- Input Monitoring permission, only for the hold-to-talk hotkey
Get the code
Download it, or clone it and follow the steps below. You can also hand the repo to Claude Code or Codex and ask it to set it up for you.
git clone https://github.com/per-simmons/voice-os.git
cd voice-os
npm install -g agent-desktop
cp .env.example .env
./run.sh- After installing agent-desktop, grant it Accessibility in System Settings, Privacy & Security, Accessibility.
- Paste your OPENAI_API_KEY into the .env file.
- ./run.sh creates a Python environment, installs what it needs and launches. Press Enter, talk, and it acts.
- For a hotkey instead, run ./ptt.sh and hold Right Control anywhere to talk. For the local wake word “hey chat,” run ./run.sh --local.
The full setup guide is in the README on GitHub.
Try it
- Say “open Spotify.”
- Say “play some Tchaikovsky.”
- Say “what’s on my screen?”
Security
- With the local wake word, your mic is processed on your Mac and thrown away. Nothing leaves the machine until you say the wake word.
- Push-to-talk sends audio only while you hold the key.
- The cloud wake word mode streams your mic to OpenAI the whole time it runs.
Good to know
- It costs money per use: roughly a few cents per command on gpt-realtime-2. Push-to-talk and the local wake word cost nothing while idle, but the cloud wake word streams your mic constantly at about $1 an hour.
- The local “hey chat” wake word works in a quiet room and gets misheard in noise.
- The built-in tools are examples. Anything beyond them you add yourself (or have your coding agent write).
- The optional waveform overlay is experimental and can block clicks near the top of the screen, so only run it while demoing.