Talking to Yumii

The conversation model: continuous listening, barge-in, sessions, and the window itself.

The conversation model

Yumii listens continuously while the orb is active — no wake word, no button. The loop:

  1. You speak. Local voice detection (Silero, on your CPU) notices when speech starts and ends.
  2. She transcribes. Local Whisper by default; Groq's cloud Whisper if configured (faster, more accurate).
  3. She thinks. Your words + her personality + her memory of you go to the LLM. Tools may trigger the permission gate.
  4. She speaks. Synthesized locally, sentence-by-sentence — first words in about a second.

Interrupting her (please do)

Start talking while she's speaking and she stops. Mid-sentence, mid-word, doesn't matter — her unspoken words are dropped and she listens. Conversations flow much better once you stop treating her turns as sacred.

The window

  • Drag the orb itself to move her anywhere on screen. The window is frameless, always-on-top, and skips the taskbar — a presence, not an app.
  • The mic button beside the orb mutes her ears; the ⚙ gear opens the mode menu and dashboard.
  • Ctrl+Shift+Space toggles her from anywhere.
  • The tray icon has Show/Hide, Dashboard, and Quit — quitting also shuts the backend down cleanly.

Sessions: conversations that persist

Every conversation is a session, stored locally and titled from your first words:

  • The current session survives reconnects and restarts — no fragmented history.
  • Dashboard → Chats lists them all: read transcripts, resume any thread, rename, delete.
  • Long-term memory is shared across all sessions — a new chat doesn't mean she forgot you. And she can search every past session by content: just ask "what did we say about the trip?" See Memory.

Languages

She speaks the language you speak to her, defaulting to English. Her local voice (Kokoro) is English-first; other languages sound best through a cloud voice — see Providers → TTS.