Find a file
Sourav Anand 11ca1ea1b3
Replace demo video links with new URLs
Updated video links in the demo section to point to new URLs.
2026-07-04 04:58:16 +05:30
.github Replaced YouTube iframes with thumbnails in README and updated CI workflow triggers 2026-07-04 04:38:32 +05:30
app Added camera permission handling, fixed image conversion row strides, and added demo videos to README. 2026-07-04 04:27:35 +05:30
docs moved Edge TTS key to settings page. 2026-06-30 21:14:18 +05:30
gradle added API key import through QR feature 2026-06-30 22:18:35 +05:30
scratch adding more options on setting screen with verification and media support 2026-06-05 00:58:38 +05:30
screenshots Replaced YouTube demo links with local videos and updated README table layout 2026-07-04 04:50:06 +05:30
.gitattributes Initial commit 2025-09-23 01:01:18 +05:30
.gitignore code refactor started with step 1 2026-05-30 13:57:06 +05:30
build.gradle.kts Up to 101 2026-05-02 09:24:20 +05:30
CONTRIBUTING.md Added contribution guidelines 2025-12-04 14:14:26 +05:30
gradle.properties Refactor step 4 migrated to Room DB 2026-05-30 21:56:13 +05:30
gradlew Initial commit 2025-09-23 01:01:18 +05:30
gradlew.bat Initial commit 2025-09-23 01:01:18 +05:30
LICENSE Initial commit 2025-09-23 01:01:18 +05:30
README.md Replace demo video links with new URLs 2026-07-04 04:58:16 +05:30
settings.gradle.kts Up to 101 2026-05-02 09:24:20 +05:30

AI Assistant for Android

License Kotlin Compose MediaPipe TensorFlow Lite Sherpa ONNX

An offline-first Android assistant that runs locally on your device. It supports screen vision, live camera sharing with real-time voice chat, encrypted data storage, and lets you connect your own LLM API keys. You can set it as your default phone assistant.

App Overview

Conversational Chat UI Hands Free Floating Orb System Settings

Demo


Table of Contents


Overview

This app is a private voice assistant for Android. It runs offline as much as possible, encrypts your conversations, and reads your screen when asked. You can connect it to any OpenAI-compatible LLM endpoint.

Unlike standard assistants, it does not send your voice data to external servers. It replaces Android's default assistant, so you can launch it with a long press on the home button.


Features

AI & LLM

  • Custom LLM Providers: Works with Groq, OpenAI, Anthropic, Gemini, Ollama, LM Studio, etc.
  • Vision Auto-Detection: Automatically tests if your model supports images and enables screen vision and live camera features if it does.
  • Fast Speech Output: Starts speaking sentences before the model finishes generating the full response.

Voice Pipeline

  • Voice Activation (VAD): Detects when you start and stop talking locally on your device.
  • Interruption Support: You can speak over the assistant to interrupt it.
  • Hands-Free Mode: Automatically listens for your next response after speaking.
  • Audio Routing: Supports Bluetooth headsets and speakerphone.

Screen Vision

  • Screen Capture: Captures screen frames locally in the background.
  • Smart Frame Syncing: Matches the captured frame to the exact moment you started speaking.
  • Camera Sharing: Share your live camera feed to discuss your real-world surroundings with the assistant.
  • Privacy Controls: Easily pause or stop screen capture from a system notification.

Privacy & Encryption

  • Secure Storage: Encrypts conversations and attachments using AES-256-GCM.
  • Safe Deletion: Securely deletes files from the device when you delete a message or chat.

Offline Tools

  • Offline Actions: Processes commands (like calling or navigation) offline using an on-device model.
  • Offline Speech & Voice: Handles speech-to-text and text-to-speech without internet.
  • Offline Translation: Translates input to English locally using Google ML Kit.

Android Integration

  • Default Assistant: Launch by long-pressing the home button or swiping.
  • Lock Screen Support: Answer voice queries without unlocking your phone.
  • System Actions: Set alarms, dial contacts (with fuzzy name matching), search YouTube music, and check the weather.

Architecture Overview

The app is built around offline first, modular components. Each layer can be swapped independently.

graph TD
    A[User — Voice / Text / File] --> B[MainViewModel]

    subgraph Input Processing
        B --> C[TranslatorManager\nML Kit Offline]
        C --> D[TextClassifierHelper\nMediaPipe TFLite]
    end

    subgraph Screen Context
        SC[ScreenCaptureService\n~6 FPS] --> VB[VisualBufferManager\nCircular Frame Buffer]
        VB --> B
    end

    subgraph Intent Routing
        D -->|CALL / SONGS / ALARM\nNAVIGATION / WEATHER / REMINDER| E[Native Action Executor]
        D -->|OTHER / SETTINGS| F[FlexibleLlmAdapter\nStreaming REST]
    end

    subgraph Lock State Machine
        E -->|Missing detail| LS[LockState\nMulti-Turn Collector]
        LS --> B
    end

    subgraph TTS Layer
        F -->|Streaming tokens| SB[Sentence Boundary Detector]
        E --> SB
        SB --> TTS[TtsEngineSelector]
        TTS --> T1[EdgeTTS — Free\nNeural WebSocket]
        TTS --> T2[Google Cloud TTS]
        TTS --> T3[Offline VITS\nSherpa-ONNX]
        TTS --> T4[System Native TTS]
    end

    subgraph Storage
        B --> DB[(Room DB\nAES-256-GCM\nEncrypted)]
    end

    subgraph Voice Pipeline
        MIC[Microphone] --> AH[AudioHygieneProcessor\nAEC · NS · AGC]
        AH --> VAD[VadIntelligenceProcessor\nSilero VAD]
        VAD --> VSM[VoiceStateMachine\nIDLE · LISTENING\nPROCESSING · BOT_SPEAKING]
        VSM --> B
    end

Design details:

  • Flexible LLM Adapter: Uses templates to support any OpenAI-compatible API endpoint.
  • Intent Verification: Checks user queries against keyword lists to prevent incorrect system actions.
  • Local Encryption: All conversation history and files are encrypted before saving to disk.

Request Processing Flow

Voice / Text Input

When you speak or type, the app processes the audio locally:

  1. It records and cleans up the audio.
  2. It detects when you start and stop speaking.
  3. If you speak while the assistant is answering, it stops talking and listens to you.
Audio Processing & STT

The app converts your speech to text using either Android's built-in system, an offline voice model, or a cloud API.

Translation & Classification
  1. If you speak a non-English language, the app translates it to English offline using Google ML Kit.
  2. It classifies the text to determine your intent (e.g., call, alarm, navigation, or general question).
  3. It double-checks keywords to make sure it doesn't run a system command by mistake.
Native Actions & Lock States

When the app identifies a command, it runs the matching phone action:

  • Calls: Searches contacts and dials the number.
  • Music: Searches and plays songs on YouTube.
  • Alarms/Reminders: Sets system alarms using natural time phrasing.
  • Navigation: Launches Google Maps navigation.
  • Weather: Gets your location and displays the forecast.

If a command is missing information (like a name for a call), the app will ask you for it directly on the next turn.

LLM & Vision

General questions go to your selected LLM. If you use screen vision or live camera sharing, the app attaches the latest captured frame from your screen or camera when you start speaking, provided your selected LLM supports images.

Streaming TTS

As the LLM generates a response, the app parses it sentence-by-sentence. Once a sentence is ready, the voice engine starts speaking it immediately so you don't have to wait for the full response to load.

Conversation Storage

All messages are encrypted and saved locally. Attached files are also encrypted before saving. Deleting a message or chat deletes all associated files permanently.


Technologies Used

Category Technology
Language Kotlin
UI Jetpack Compose, Material You
Architecture MVVM, ViewModel, Coroutines, Flow
On-Device ML MediaPipe Text Classifier, TensorFlow Lite, Sherpa-ONNX (Silero VAD, Parakeet STT, VITS TTS)
Translation Google ML Kit On-Device Translation
Database Room (encrypted via Android KeyStore + AES-256-GCM)
Networking OkHttp, Retrofit
TTS Edge TTS (WebSocket), Google Cloud TTS, Sherpa-ONNX VITS, Android TextToSpeech
APIs Groq (default LLM), YouTube Data API v3, Open-Meteo (weather)
Location GMS Fused Location Provider

Project Structure

app/src/main/java/com/app/assistant/
│
├── AssistActivity.kt            # Entry point for system assistant integration
├── MainActivity.kt              # Main UI activity and lifecycle management
│
├── api/                         # HTTP clients and API services
│
├── camera/
│   ├── ScreenCaptureService.kt  # Captures screenshots periodically
│   └── VisualBufferManager.kt   # Stores captured frames with timestamps
│
├── classifier/
│   └── TextClassifierHelper.kt  # Local TFLite intent classifier
│
├── db/
│   ├── AppDatabase.kt           # Room database setup
│   ├── EncryptionUtil.kt        # Encryption helpers (AES-GCM + KeyStore)
│   ├── DynamicConversationRepo  # Handles encrypted database operations
│   └── *Entity.kt               # Database schemas
│
├── llm/
│   ├── FlexibleLlmAdapter.kt    # API adapter for various LLMs
│   ├── ModelCapabilityProber.kt # Detects model vision capability
│   └── LlmMessage.kt            # LLM data schemas
│
├── repository/
│   ├── ContactsRepository.kt    # Matches contact names offline
│   ├── SettingsRepository.kt    # App settings manager
│   └── WeatherRepository.kt     # Location and weather logic
│
├── speech/
│   ├── AudioHygieneProcessor.kt    # Audio recording and processing
│   ├── VadIntelligenceProcessor.kt # On-device speech detection (VAD)
│   ├── VoiceStateMachine.kt        # Manages voice states (listening, speaking)
│   └── SpeechRecognizerManager.kt  # Manages speech-to-text engines
│
├── translation/
│   └── TranslatorManager.kt     # Manages offline translation
│
├── tts/
│   ├── TtsEngineSelector.kt     # Switches between TTS engines
│   ├── EdgeTtsApiManager.kt     # Microsoft Edge TTS integration
│   ├── GoogleTtsApiManager.kt   # Google Cloud TTS integration
│   ├── OfflineTtsManager.kt     # On-device TTS integration
│   └── NativeTtsManager.kt      # Default system TTS fallback
│
└── ui/
    ├── screen/
    │   ├── ChatScreen.kt         # Main chat screen logic
    │   ├── ChatLayout.kt         # Layout definitions for chat
    │   ├── ConversationItem.kt   # Chat item UI elements
    │   ├── HandsFreeBar.kt       # Voice UI overlay
    │   ├── SettingsScreen.kt     # Settings layout and configuration
    │   └── UserInputField.kt     # Text input and file picker UI
    └── theme/                    # Theme, styling, and typography

Setup

Prerequisites

  • Android Studio
  • Device or emulator (Android 8.0+ / API 26+)
  • Groq API key (or another OpenAI-compatible endpoint key)
  • YouTube Data API v3 key

Steps

  1. Clone this repository:
    git clone https://github.com/your-username/ai-assistant-android.git
    cd ai-assistant-android
    
  2. Open the project in Android Studio and let files sync.
  3. Configure your API keys (see details in the API Configuration section).
  4. Run the app on your device.
  5. (Optional) Set the app as your default assistant under Settings → Apps → Default Apps → Digital assistant app.

API Configuration

Choose one of these methods to configure your API keys:

Option 1: Using local.properties Add the keys to your local.properties file:

YOUTUBE_API_KEY=YOUR_YOUTUBE_API_KEY
GROQ_API_KEY=YOUR_GROQ_API_KEY
EDGE_TTS_SUBSCRIPTION_KEY=YOUR_EDGE_TTS_SUBSCRIPTION_KEY

Gradle will load these during the build. (Make sure this file is not committed to git).

Option 2: In-App Settings Open the app, go to Settings, and paste your keys directly into the input fields. This does not require rebuilding the app.

Customizing the LLM Endpoint In the settings page, you can change the default URL to any OpenAI-compatible server (like Ollama, LM Studio, Anthropic, or Gemini). The app automatically checks if the custom endpoint supports images.


Permissions

Permission Required For
RECORD_AUDIO Listening to voice input
CAMERA Accessing the camera for live camera sharing
READ_CONTACTS Searching contact names to make calls
CALL_PHONE Making phone calls
ACCESS_FINE_LOCATION Checking weather and generating navigation routes
MEDIA_PROJECTION Recording the screen for screen vision
BLUETOOTH / BLUETOOTH_CONNECT Supporting Bluetooth headsets
POST_NOTIFICATIONS Showing the screen capture status bar notification

Supported Features Table

Feature Offline Cloud Native Android AI / LLM
Voice activity detection (VAD) ✅
Speech-to-text (Parakeet) ✅
Speech-to-text (Cloud) ✅
Intent classification ✅
Language translation ✅
Offline TTS (VITS) ✅
Neural TTS (Edge) ✅
Encrypted storage ✅
Phone calls ✅
Alarms & reminders ✅
Google Maps navigation ✅
YouTube music playback ✅
Weather forecasts ✅
Contact matching ✅ ✅
Screen vision ✅ ✅
Live camera sharing ✅ ✅
General chat ✅
Custom LLM endpoints ✅ ✅
Image & file attachments ✅

Privacy & Security

Your data stays on your device unless it is sent to your selected LLM provider.

Database Encryption The app generates an AES-256 encryption key inside Android's secure hardware Keystore. All messages and chat titles are encrypted before saving. Each record uses a unique initialization vector.

File Encryption All saved images, documents, and screenshots are encrypted before they are written to disk.

Secure Deletion When you delete a message or chat, the app deletes the database records and removes the associated files from disk.

Voice Processing Your raw voice audio is never stored or sent anywhere. Voice detection is handled entirely in memory on the device, and only the transcribed text is processed.


Known Issues

  • Barge-In Interruption: Sometimes the barge-in/interruption feature does not trigger correctly when speaking over the assistant. Work is underway to refine this.
  • Translation Disabled: The translation feature is currently disabled due to conflicts with streaming text output. A new approach is being developed.
  • Settings Classification: Settings-related commands (unlike Call, Play Song, etc., which work) do not have a separate action yet. Instead, they are classified under OTHER and sent directly to the LLM. A new approach is being worked on to support settings actions natively.

Roadmap

  • Android XR support for smart glasses
  • Wear OS support for wrist control
  • Category for sending text messages through actions
  • Automatic summary of older conversations
  • Screen capture using Accessibility Services to avoid permission prompts

Contributing

Contributions are welcome. Please open an issue to discuss new features before making changes, or open a pull request directly for bug fixes.

  1. Fork the repository and create a branch.
  2. Follow the project's coding style (standard Kotlin, MVVM architecture).
  3. Submit a pull request explaining your changes.

Make sure your changes do not store unencrypted user data or bypass security features.


License

This project is licensed under the MIT License.

Copyright © 2026 Sourav Anand.