Provides local vision and audio perception for MCP-compatible agents, enabling them to read images, transcribe text fro…
Provides local vision and audio perception for MCP-compatible agents, enabling them to read images, transcribe text from visual media, analyze videos, and convert speech to text entirely on-device. It is privacy-focused with no cloud upload or API keys by default, using Ollama for inference.
Scheme0
mcp
free