Umm: Voice Typing

Personal Project — Sep 2026

Umm is an open-source/source-available Android voice keyboard and accessibility floating dictation assistant that transcribes speech and types a cleaned-up version directly into any app. When dictating, spoken filler words (“um”, “uh”, “like”), false starts, and mid-sentence self-corrections are automatically removed, while punctuation, paragraph breaks, and formatting (such as bullet lists in emails) are inferred from context. Mixed speech such as Hinglish is preserved as spoken without unwanted translation.

Unlike subscription-based dictation services or device-locked alternatives, Umm operates entirely on a pay-per-use model powered by the user's own OpenRouter account. Dictations are recorded on-device in mono AAC with local silence and speech detection. Audio is transcribed via OpenRouter's audio transcription endpoint (defaulting to openai/gpt-4o-mini-transcribe with fallbacks), and transcripts are structured and refined through a lightweight LLM completion stage (defaulting to google/gemini-3.5-flash-lite). A typical short dictation costs about $0.0004 and resolves in 1.2 to 1.8 seconds.

In addition to the standard input method editor (IME) keyboard, Umm includes an optional floating button built on Android's Accessibility API, enabling hold-to-talk, tap-to-stop, or continuous dictation over any editable text field without switching away from the user's primary keyboard. Privacy is foundational: Umm has no centralized backend, accounts, advertising, or telemetry; all requests communicate directly from the phone to OpenRouter with credentials encrypted in Android Keystore, and optional zero-data-retention routing ensures upstream providers never store speech payloads.

Highlights

  • Real-time voice dictation with automated cleanup of filler words, false starts, and speech disfluencies
  • Four configurable cleanup levels (Raw, Light, Formatted, Polished) mapped per application category
  • Optional screen-docked floating button using Android Accessibility API for single-tap or hold dictation across apps
  • Two-stage pipeline: on-device silence detection, OpenRouter audio transcription, and LLM text formatting
  • Pay-per-use economics with personal OpenRouter keys (~$0.0004/dictation), monthly spending caps, and in-app usage tracking
  • Privacy-first architecture: no intermediate servers, encrypted local key storage, and zero-data-retention provider enforcement
  • Multilingual and code-mixed speech support (such as Hinglish) preserved accurately without translation

Download the latest release on GitHub (opens in a new tab).

Technologies Applied

KotlinAndroidJetpack ComposeOpenRouterSpeech-to-TextLLMMaterial 3Accessibility API