
omniparser-autogui-mcpAutomatic operation of on-screen GUI.
Overview
Key features
- Screen analysis
- GUI automation
- Configurable target window
- Server address specification
- SSE communication
Pricing
- Model
- Free
- Category
- MCP Servers
- Rating
- No reviews yet
Use cases
Automated browser search
Search for text in an on-screen browser
GUI automation
Automatically operate the GUI based on screen analysis
Pros & Cons
Pros
- Automatic GUI operation
- Screen analysis with OmniParser
- Configurable via environment variables
Cons
- Limited to Windows
- Requires configuration of claude_desktop_config.json file
- Dependent on OmniParser models with different licenses
Reviews
Sign in to leave a review.
No reviews yet. Be the first!
Q&A
Can I customize the target window?
Yes, you can specify the target window via the TARGET_WINDOW_NAME environment variable.
Asked by Gunnar Eriksson · Oct 9, 2025
What configuration is required?
The tool requires installation via git clone and configuration of the claude_desktop_config.json file.
Asked by Ximenez Alvarado · Aug 28, 2025
How is it licensed?
It is licensed under MIT, excluding submodules and subpackages.
Asked by Esther Adeyemi · Aug 21, 2025
What platforms does it support?
It is confirmed to work on Windows.
Asked by Ola Eriksen · Jul 2, 2025
Ask a question
MCP Servers alternatives

Open-source MCP server that lets LLMs drive real browsers via Playwright and accessibility snapshots.

Python agent framework from the Pydantic team for building type-safe GenAI apps.

Adaptive memory layer that helps AI agents learn from context over time.

AI email assistant that organizes, drafts replies, and helps you reach inbox zero faster.

Open-source 24/7 local screen and audio recording for building context-aware AI apps

TypeScript library for building and orchestrating AI agents with tools, memory, and multi-agent workflows.

Bringing the bankless onchain API to MCP

Python tool for converting files and office documents to Markdown.
Trending now

Document intelligence API that parses, splits, OCRs, and extracts structured data from complex PDFs, slides, and spreadsheets.

Sponsored answers, paid per click.

Accurate Homework Help with Full Explanations

Open multimodal 12B model handling interleaved images and text with a 128K context window.
