ScreenMind: Vision, Audio & Chat with Local AI | Open Source Friday
Ayush Shekhar introduces ScreenMind, an open source project that combines vision, audio transcription, and Q&A/reasoning using a single local model.
Overview
ScreenMind is presented as a local (on-device) AI setup that aims to:
- Understand what is happening on a user’s screen (vision)
- Transcribe audio such as meetings (speech-to-text)
- Answer questions about what the system has seen (chat/Q&A with reasoning)
A key point highlighted in the video description is that ScreenMind uses a single Gemma 4 model and is intended to run locally on a GPU, with a stated target of as little as 4GB of VRAM.
Project link
- GitHub repository: https://github.com/ayushh0110/ScreenMind