What is Appium MCP? Appium MCP is an open source Model Context Protocol server from the Appium project that exposes Appium's mobile automation and driver capabilities as tools, so coding agents such as Claude Code and Cursor can start device sessions, find elements, tap, type, and capture screens on Android and iOS.
Mobile engineers search for appium mcp for a simple reason. Their coding agent can now write a new checkout screen in minutes, and nobody wants to wait a day for someone to tap through it on a phone. If the agent can drive a device directly, the loop gets shorter.
This guide walks through that loop honestly. You will set up the official Appium MCP server in Claude Code and Cursor, see what its tools actually cover, and learn where those tools stop. Then we look at peer servers and at where a verification layer like QApilot MCP fits on top when the question changes from "can my agent tap this button?" to "did this change break the release?"
Why coding agents need a mobile MCP server at all
The Model Context Protocol is an open standard that lets an AI client call external tools through a common interface. On the web, agents already have browser tools. On mobile, the agent has nothing to look at unless something connects it to an emulator, a simulator, or a physical device.
That gap shows up in three places:
- Writing UI code blind. The agent edits a Jetpack Compose, SwiftUI, or Flutter screen without seeing how it renders on a device.
- Debugging without state. A crash report says "button not found," but the agent cannot inspect the live view hierarchy.
- Checking work before the PR. The agent claims a flow works. Someone still has to prove it on a real build.
An appium mcp server solves the first two directly. It gives the agent eyes and hands on the device through Appium's mature UiAutomator2 and XCUITest drivers. The third problem is where it gets more interesting, and we will come back to it.
What you need before you install
Appium MCP's README lists a short set of prerequisites. Check these first, because most setup failures trace back to a missing SDK path rather than the MCP server itself.
| Requirement | Why it matters |
|---|---|
| Node.js 22 or higher | The server runs through npx |
| Java JDK 8 or higher | Needed by the Android toolchain |
Android SDK with adb on PATH |
Android emulators and USB devices |
ANDROID_HOME set |
The server reads it to find the SDK |
| Xcode and simulators (macOS only) | iOS simulator and real device sessions |
One useful detail: for local Android and iOS sessions, Appium MCP bundles the UiAutomator2 and XCUITest drivers. You do not need a separately installed global Appium server for that mode. If you already run Appium on a grid or device farm, you can point the MCP server at it with a remote server URL instead.
Confirm your device is visible with adb devices (Android) or by booting a simulator in Xcode before you connect the agent.
How to set up Appium MCP in Claude Code
How do you set up Appium MCP with Claude Code? Register the official server with the claude mcp add command, pass your Android SDK path as an environment variable, then restart the session and confirm the Appium tools appear.
The command from the project README is:
claude mcp add appium-mcp -- npx -y appium-mcp@latest
To set the SDK path at the same time, add an environment flag:
claude mcp add appium-mcp -e ANDROID_HOME=/path/to/android/sdk -- npx -y appium-mcp@latest
The Claude Code MCP documentation explains scope options, so you can register the server for one project or for your whole user profile. After adding it, start a new session and ask the agent to list available devices. If it calls select_device and returns your emulator, the connection works.
How to set up Appium MCP in Cursor
Cursor supports MCP servers through its settings or a JSON config file, as described in the Cursor MCP docs. The Appium README offers a one-click install button, or you can add it manually under Cursor Settings, then MCP, then Add new MCP Server.
A project-level config looks like this:
{
"mcpServers": {
"appium-mcp": {
"type": "stdio",
"command": "npx",
"args": ["appium-mcp@latest"],
"timeout": 100,
"env": {
"ANDROID_HOME": "/Users/you/Library/Android/sdk",
"CAPABILITIES_CONFIG": "/path/to/capabilities.json"
}
}
}
}
CAPABILITIES_CONFIG is optional. It points at a JSON file with per-platform capability presets such as app path, device name, platform version, and UDID. Teams that share a repo usually commit a capabilities file so every engineer's agent starts sessions the same way.
What the Appium MCP tools cover
At the time of writing, the Appium MCP README documents 30 tools. Three of them are opt-in: appium_ai (vision-based element finding, which needs your own vision model API key) and two documentation tools. Many tools bundle several actions behind one name, so the real surface is wider than the count suggests. The list changes as the project ships, so check the repository before you rely on a specific tool.
Here is how the tools group by job:
| Category | Example tools | What the agent can do |
|---|---|---|
| Platform and device setup | select_device, prepare_ios_simulator, appium_prepare_ios_real_device |
Pick a device, boot a simulator, sign WebDriverAgent for a real iPhone |
| Sessions | appium_session_management, appium_driver_settings |
Create, attach, list, and delete sessions, local or remote |
| Contexts | appium_context |
Switch between native and webview contexts |
| Elements and gestures | appium_find_element, appium_gesture, appium_set_value, appium_alert |
Find by accessibility id or other locators, tap, swipe, scroll to an element, type, handle alerts |
| Screen and device state | appium_screenshot, appium_get_page_source, appium_screen_recording, appium_geolocation |
Capture screens and XML, record video, set GPS and orientation |
| App management | appium_app_lifecycle, appium_mobile_permissions |
Install, launch, terminate, clear data, open deep links, change permissions |
| Test generation | generate_locators, appium_generate_tests |
Suggest locators for the current screen and generate test code from a scenario |
That is a strong set of automation primitives. The README is also opinionated in a good way. It tells agents to prefer accessibility ids and native locators before XPath, and it keeps vision-based finding switched off by default so the model does not reach for a slow, paid call when a stable locator would work.
For a developer who wants the agent to open the app, reproduce a bug, grab the page source, and propose a fix, this is often exactly enough.
Where Appium MCP stops
What can Appium MCP's tools not do alone? They give an agent device control, not a release check. They do not decide what to test after a change, build lasting context about your app's journeys, or return a verdict on whether the build is safe to ship.
None of this is a flaw in Appium MCP. It is a scope boundary, and it matters once you try to use the agent as your first line of verification.
The agent has to plan every step. Appium MCP executes what it is told. If you ask "check that checkout still works," your coding agent must work out the screens, the order of taps, the test data, and what counts as success. That planning happens inside the agent's context window, which competes with the code it is writing.
There is no memory of the app. Each session starts from the current screen. The tools do not keep a model of how screens connect, which journeys matter most, or what changed since the last build.
Locators are still locators. Generated tests and element finding rely on accessibility ids, resource ids, predicates, or XPath. When a redesign moves or renames elements, the same maintenance problem that slows scripted Appium suites comes back. Vision finding helps for single elements, but it is opt-in and depends on an external model.
Results are raw. You get screenshots, XML, recordings, and tool responses. Turning those into "pass, fail, and why" is still the agent's job, or yours.
If your goal is interactive control, these limits rarely matter. If your goal is "the agent wrote this UI, who checked the binary?", they matter a lot.
How Appium MCP compares to other mobile MCP servers
Appium MCP is not the only option. Two peers come up often in the same searches, and each makes a different trade.
Mobile Next mobile-mcp is a platform-agnostic server that lets agents interact with iOS and Android apps through accessibility snapshots or coordinate taps from screenshots. It installs the same way (claude mcp add mobile-mcp -- npx -y @mobilenext/mobile-mcp@latest) and offers an optional hosted device cloud. It suits developers who want quick exploration without Appium-specific knowledge.
Maestro MCP connects coding agents to the Maestro CLI so they can write, run, and debug Maestro YAML flows. It suits teams already invested in Maestro's flow format.
| Server | Best at | What you still own |
|---|---|---|
| Appium MCP | Deep device and driver control on the Appium stack you may already run | Test planning, app context, interpreting results |
| Mobile Next mobile-mcp | Fast, platform-agnostic exploration | Test planning, app context, interpreting results |
| Maestro MCP | Writing and running Maestro YAML flows | Defining flows, maintaining them |
| QApilot MCP | Intent-level verification with an agent-readable report | Deciding what must hold for the change |
All three peers are good explorers and flow runners. The open question is what sits between "the agent can drive the device" and "the team trusts the change."
Where QApilot MCP fits on top
How does QApilot MCP relate to Appium MCP? Appium MCP gives agents automation primitives. QApilot MCP is a local-first verification loop for coding agents: you state intent, it runs the check on your device, and it returns a markdown report your agent can read in the same session. Teams can use both.
The idea is captured in one line: Your coding agent writes mobile code faster than anyone can check it. QApilot MCP checks it.
In practice the loop looks like this:
- Intent in. You or your agent say what must hold, for example "verify checkout still works after this change."
- Run on your device. QApilot plans the steps on its own infrastructure, so planning does not eat your agent's context, and runs them live on your emulator or physical device.
- Report out. A markdown report lands in the session. Your agent can query it, fix what failed, and run again before the PR.
- Keep it. Tests are saved as YAML or Appium code in your repo, so you own the suite.
QApilot's planning draws on the same app understanding behind its autonomous testing platform, which builds a Knowledge Graph of screens and journeys rather than starting from a blank screen each time. That is the piece raw automation tools do not try to provide.
It is also local-first. Your app stays on your machine by default, and the current QApilot MCP CLI guide runs Android checks through a local Appium server. If you already have Appium MCP working, most of your toolchain is in place.
A few honest boundaries:
- QApilot MCP is in early access. You can follow the guide or request a walkthrough.
- Local MCP ships first. Cloud MCP is planned for later and is not generally available.
- It is not another Appium IDE or a wrapper that replaces Appium MCP. Appium MCP stays useful for hands-on device control. QApilot answers a different question.
When the job grows from one change on one device to certifying a release across a device matrix, the same engine powers the wider QApilot release readiness suite.
Choosing the right setup for your team
Use the job you need done to pick the layer, not the tool with the most stars.
Pick Appium MCP alone if your agent mainly needs to reproduce bugs, inspect the view hierarchy, and drive a device you already configure for Appium. It is free, open source, and close to the Appium docs your team knows.
Pick mobile-mcp or Maestro MCP if you want lighter exploration without Appium setup, or your team already writes Maestro flows. See how QApilot differs in the QApilot vs Maestro comparison.
Add QApilot MCP when the bottleneck is verification: agents ship UI changes faster than QA can check them, and you want a pass or fail verdict with reasons before merge. The QApilot vs Appium comparison covers the broader platform differences.
A simple rule works for most teams. Use Appium MCP when the agent needs hands. Use QApilot MCP when the team needs an answer.






