Native iOS and Android Recipes
Native iOS and Android development with a coding agent works when each platform has one script that builds the app, runs its tests on a pinned simulator or emulator, and saves the results as evidence. Claude Code, Codex, and Cursor iterate against that script, and an MCP server such as MobileBuildMCP or mobile-mcp lets the agent see the running app.
You ask the agent for a new SwiftUI settings screen. It writes clean Swift, announces “done”, and the project does not build, because the new file never reached the Xcode target. On the Android side, the agent runs ./gradlew clean build on every change, waits minutes for a result, and pastes thousands of lines of Gradle output into its own context. Neither agent ever looked at the screen it built. This page fixes those three gaps: build feedback, speed, and proof.
What you get from an agent-ready native project
Section titled “What you get from an agent-ready native project”- Two copy-paste scripts,
scripts/ios-check.shandscripts/android-check.sh. Each one boots a pinned device, runs the tests, writes a short summary for the agent, and saves the full evidence underbuild/evidence/. - An instructions block for
AGENTS.md,CLAUDE.md, or a Cursor rule that makes “the script exits 0” the definition of done. - Verified install commands for the two mobile MCP servers, plus the Codex plugins for iOS and Android, per tool.
- Build-time rules that keep the agent’s inner loop short: narrow test selection, warm devices, and caches.
- Recipes for a SwiftUI screen, a Jetpack Compose screen, and a crash reproduction, each ending in UI-test evidence rather than a diff you have to read.
Why native apps need a different agent loop
Section titled “Why native apps need a different agent loop”Web agents get cheap feedback: a dev server reloads in a second and a browser MCP shows the page. Native projects take longer to give that feedback, for three reasons:
- The project file is part of the build. Xcode decides what compiles from the target membership in
project.pbxproj, and Gradle from the module layout. An agent that writes a correct file in the wrong place gets a build error or, worse, a file that silently never compiles. - Builds are slow and loud. A clean
xcodebuildor Gradle build prints thousands of lines. Streamed into the agent’s context, that output costs tokens and pushes the task instructions out of view. - The UI lives in a separate process. The simulator and the emulator are separate programs. Without a test runner or an MCP server, the agent has no way to know whether the screen renders, and neither do you.
So the recipes on this page do not start with a feature prompt. They start by giving the agent a single, quiet, reliable command per platform, and they judge the agent by what that command proves.
Set up the build-and-verify loop
Section titled “Set up the build-and-verify loop”-
Pin one simulator and one emulator. List what exists with
xcrun simctl list devices availableandemulator -list-avds(theemulatorbinary lives in$ANDROID_HOME/emulator). Pick one of each and write the names into the scripts below. A pinned device means a test failure is about your code, not about which phone the agent happened to boot. -
Add
scripts/ios-check.sh. It keeps DerivedData inside the repository, writes the full log to a file, and prints only the lines the agent needs:#!/usr/bin/env bash# scripts/ios-check.sh — build, test and collect evidence on one pinned simulator.# Usage: scripts/ios-check.sh [TestTarget/TestClass ...] (no argument = all tests)set -euo pipefailSCHEME="${SCHEME:-App}"SIM="${SIM:-iPhone 17}"OUT="build/evidence/ios"rm -rf "$OUT" && mkdir -p "$OUT"ONLY=()for t in "$@"; do ONLY+=("-only-testing:$t"); donexcrun simctl bootstatus "$SIM" -b >/dev/null # boots the simulator if needed, waits until readyset +excodebuild test \-scheme "$SCHEME" \-destination "platform=iOS Simulator,name=$SIM" \-derivedDataPath build/DerivedData \-resultBundlePath "$OUT/tests.xcresult" \${ONLY[@]+"${ONLY[@]}"} > "$OUT/xcodebuild.log" 2>&1STATUS=$?set -excrun xcresulttool get test-results summary --path "$OUT/tests.xcresult" > "$OUT/summary.json" 2>/dev/null || truegrep -E "error:|\*\* (BUILD|TEST) (SUCCEEDED|FAILED)|failed" "$OUT/xcodebuild.log" | tail -40echo "exit=$STATUS evidence=$OUT"exit $STATUSThe
${ONLY[@]+...}form matters: macOS still ships Bash 3.2, where an empty array underset -uaborts the script.xcresulttool get test-resultsneeds Xcode 16 or later. -
Add
scripts/android-check.sh. It reuses a running emulator, clears logcat so the log covers only this run, and copies the Gradle test reports into the evidence folder:#!/usr/bin/env bash# scripts/android-check.sh — unit + instrumented tests on one pinned emulator, evidence collected.# Usage: scripts/android-check.sh [com.example.app.SomeUiTest] (no argument = all tests)set -euo pipefailAVD="${AVD:-Pixel_8_API_35}"MODULE="${MODULE:-app}"OUT="build/evidence/android"rm -rf "$OUT" && mkdir -p "$OUT"if ! adb get-state >/dev/null 2>&1; thenemulator -avd "$AVD" -no-window -no-audio -no-boot-anim > "$OUT/emulator.log" 2>&1 &adb wait-for-deviceuntil [ "$(adb shell getprop sys.boot_completed | tr -d '\r')" = "1" ]; do sleep 2; donefiadb logcat -cFILTER=()if [ $# -gt 0 ]; then FILTER=("-Pandroid.testInstrumentationRunnerArguments.class=$1"); fiset +e./gradlew ":$MODULE:testDebugUnitTest" ":$MODULE:connectedDebugAndroidTest" \${FILTER[@]+"${FILTER[@]}"} > "$OUT/gradle.log" 2>&1STATUS=$?set -eadb logcat -d > "$OUT/logcat.txt"cp -R "$MODULE/build/reports" "$OUT/reports" 2>/dev/null || truecp -R "$MODULE/build/outputs/androidTest-results" "$OUT/androidTest-results" 2>/dev/null || truegrep -E "FAILED|BUILD (SUCCESSFUL|FAILED)|e: " "$OUT/gradle.log" | tail -40echo "exit=$STATUS evidence=$OUT"exit $STATUSThe optional argument narrows only the instrumented run to one class; unit tests always run in full, because they are cheap.
-
Tell the agent the rules. Paste this block into
AGENTS.mdfor Codex,CLAUDE.mdfor Claude Code, or a project rule for Cursor. Claude Code readsAGENTS.mdwhen the repository has noCLAUDE.md(v2.1.277 or later, on thelatestrelease channel as of 2026-09-26;stableusers still need aCLAUDE.mdthat points toAGENTS.md), so one file can serve both CLIs. If you have not written an instructions file before, start with AGENTS.md and CLAUDE.md:## Native build and test loop- iOS: run `scripts/ios-check.sh`, optionally with a test class such as `AppUITests/SettingsUITests`.Never call `xcodebuild` without `-derivedDataPath build/DerivedData`.- Android: run `scripts/android-check.sh`, optionally with a fully qualified test class.Never run `./gradlew clean` unless the user asks.- A task is done only when the relevant script prints `exit=0`. Quote its summary lines in your final message.- Every control a test touches gets a stable ID: `.accessibilityIdentifier(...)` in SwiftUI,`Modifier.testTag(...)` in Compose. Never locate elements by visible text alone.- Never edit snapshot references (`__Snapshots__/`, `src/test/snapshots/`) or weaken an existingassertion to make a test pass. Stop and report the failure instead.- New Swift files go inside an existing synchronized folder; do not edit `project.pbxproj` by hand. -
Commit the scripts and run them once yourself. A script that fails on a clean checkout teaches the agent to ignore it. Run both, open
build/evidence/, and addbuild/to.gitignoreif it is not there yet.
Connect the agent to the simulator and emulator
Section titled “Connect the agent to the simulator and emulator”Tests prove the behaviour you wrote down. For exploratory work (does the screen look right, what happens when I tap here), the agent needs to see and drive the device. Two MCP servers cover this. If you have never added an MCP server, read the introduction to MCP first:
- MobileBuildMCP (formerly XcodeBuildMCP, maintained by Sentry) builds, runs, and debugs iOS and macOS projects on simulators and devices. It needs macOS 14.5+ and Xcode 16+. Its most-used tool is
build_run_sim, which builds the scheme, installs the app on a booted simulator, launches it, and captures logs. - mobile-mcp (Mobile Next) drives iOS simulators, Android emulators, and real devices through one tool set:
mobile_list_elements_on_screenfor an accessibility snapshot,mobile_take_screenshot,mobile_click_on_screen_at_coordinates,mobile_get_device_logs,mobile_get_crash, and more. Version 1.0.5 advertises 32 tools.
Popularity: on 2026-09-26 MobileBuildMCP had 6.4k GitHub stars and npm mobilebuildmcp 2.7.1 (published 2026-09-23); on 2026-09-28 mobile-mcp had 7,958 GitHub stars, with npm @mobilenext/mobile-mcp 1.0.5 (published 2026-09-23). Sources: the GitHub repositories getsentry/MobileBuildMCP and mobile-next/mobile-mcp, and the npm registry.
Add the servers from the terminal. The default local scope keeps them out of teammates’ sessions until you decide to share them with -s project:
claude mcp add mobilebuild -- npx -y mobilebuildmcp@latest mcpclaude mcp add mobile-mcp -- npx -y @mobilenext/mobile-mcp@latestThe Claude Code desktop app on macOS also has an iOS Simulator pane (beta since July 2026), so you can watch the simulator next to the session. To keep the agent away from snapshot references, add deny rules to .claude/settings.json:
{ "permissions": { "deny": [ "Edit(**/__Snapshots__/**)", "Edit(**/src/test/snapshots/**)" ] }}Add the servers with codex mcp add (the -- separates the server command):
codex mcp add mobilebuild -- npx -y mobilebuildmcp@latest mcpcodex mcp add mobile-mcp -- npx -y @mobilenext/mobile-mcp@latestSimulator builds can outlast the default MCP tool timeout. MobileBuildMCP’s client docs tell Codex users to raise it per server in ~/.codex/config.toml, because Codex ignores a top-level tool_timeout_sec for MCP calls:
[mcp_servers.mobilebuild]command = "npx"args = ["-y", "mobilebuildmcp@2.7.1", "mcp"] # pinned, like the committed configstool_timeout_sec = 600Codex’s official plugin marketplace also lists build-ios-apps (skills such as ios-debugger-agent, swiftui-ui-patterns, and swiftui-performance-audit, with an XcodeBuildMCP configuration) and test-android-apps (skills android-emulator-qa and android-performance); checked in openai/plugins on 2026-09-26. Install them from the TUI: run /plugins and search by name.
Two details from the plugin files (version 0.1.2 of each): build-ios-apps starts the server under its old package name, xcodebuildmcp@latest, with four workflows switched on (simulator,ui-automation,debugging,logging), and it can mirror the simulator in the Codex in-app browser. If you also added mobilebuild by hand, disable one of the two, or the agent sees every build tool twice.
Add a project-scoped .cursor/mcp.json so every teammate who opens the repository gets the same servers:
{ "mcpServers": { "mobilebuild": { "command": "npx", "args": ["-y", "mobilebuildmcp@2.7.1", "mcp"] }, "mobile-mcp": { "command": "npx", "args": ["-y", "@mobilenext/mobile-mcp@1.0.5"] } }}The versions are pinned because this file is committed: with @latest, every teammate’s next session silently runs whatever was published overnight. Bump the pins in a pull request, like any other dependency.
Cursor’s agent runs scripts/ios-check.sh and scripts/android-check.sh in its terminal like any other command. For Cursor-specific mobile tips, see Cursor mobile development.
When to use which. Use MobileBuildMCP when the problem is the Xcode build itself (schemes, signing, a Swift macro, a crash on launch). Use mobile-mcp when the problem is behaviour on screen, on either platform. Enable only the server a session needs: every advertised tool definition takes context on every turn. mobile-mcp 1.0.5 advertises all 32 tools at once. MobileBuildMCP exposes only its simulator workflow by default; widen it with the MOBILEBUILDMCP_ENABLED_WORKFLOWS environment variable (for example simulator,ui-automation,debugging,logging) only when the task needs UI automation or the debugger. In Claude Code that is claude mcp add mobilebuild -e MOBILEBUILDMCP_ENABLED_WORKFLOWS=simulator,ui-automation -- npx -y mobilebuildmcp@latest mcp. Codex’s build-ios-apps plugin sets XCODEBUILDMCP_ENABLED_WORKFLOWS instead only because it runs the old xcodebuildmcp@latest package; mobilebuildmcp 2.7.1 ignores that name.
Working inside Xcode instead of a terminal. MobileBuildMCP’s client documentation (checked 2026-09-26) describes Xcode 26.3 or later hosting the Codex and Claude Code agents under Xcode Settings > Intelligence. Those in-Xcode agents start with a minimal PATH, so a server launched with npx often fails to start; the documentation wraps the command in /bin/zsh -lc with Homebrew and nvm on the PATH. The check scripts on this page work the same from either place.
How to keep native build times inside the agent’s loop
Section titled “How to keep native build times inside the agent’s loop”A slow loop changes agent behaviour: the agent batches many edits between builds, and a failure then has many possible causes. Keep three loops, each with its own cost:
| Loop | What runs | Command | When the agent runs it |
|---|---|---|---|
| Inner | Unit tests of the changed module (Swift Testing or XCTest; JUnit) | xcodebuild test ... -only-testing:AppTests/CartTests · ./gradlew :feature:cart:testDebugUnitTest --tests '*CartViewModelTest' | After every edit |
| Middle | One UI test class on the pinned device | scripts/ios-check.sh AppUITests/CartUITests · scripts/android-check.sh com.example.cart.CartScreenTest | Before it claims “done” |
| Outer | Every suite, snapshot verification, both platforms | The same two scripts without arguments, in CI | On every pull request |
Rules that keep the inner and middle loops fast:
- Never clean by reflex. Tell the agent in its instructions file that
cleanis forbidden unless you ask. Most “fix it with a clean build” instincts come from stale DerivedData or Gradle caches, and the failure list below has the targeted fix for them. - Keep devices warm. Both scripts boot the device only when it is not running. Leave the simulator and the emulator up for the whole session.
- Turn on Gradle’s caches. Add
org.gradle.caching=trueandorg.gradle.configuration-cache=truetogradle.properties. The configuration cache skips re-running build scripts when nothing in them changed. - Split build from test on iOS when you rerun.
xcodebuild build-for-testingfollowed byxcodebuild test-without-buildingreruns a flaky test without recompiling. Use it for triage only: after a code change the agent must rebuild. - Measure before tuning.
xcodebuild -showBuildTimingSummaryprints where the time went, and./gradlew --profilewrites a local HTML report underbuild/reports/profile/. Ask the agent to read the report and propose one change, then measure again.
Recipe: a SwiftUI screen with UI-test evidence
Section titled “Recipe: a SwiftUI screen with UI-test evidence”The prompt below asks for the behaviour first, as tests, and makes the script’s exit code the finish line.
The screenshots land inside build/evidence/ios/tests.xcresult, next to the test that took them. The attachment code the agent should produce looks like this:
let app = XCUIApplication()
func snap(_ name: String) { let shot = XCTAttachment(screenshot: app.screenshot()) shot.name = name shot.lifetime = .keepAlways // keep it even when the test passes add(shot)}
func testEnablingReminderEnablesTimePicker() { app.launchArguments = ["-uiTesting"] app.launch() app.buttons["settings.open"].tap() snap("1-settings-open") app.switches["settings.reminderToggle"].tap() snap("2-reminder-on") XCTAssertTrue(app.datePickers["settings.reminderTime"].isEnabled) snap("3-time-picker-enabled")}Check one thing the agent often gets wrong: the -uiTesting launch argument has to change app behaviour (an in-memory UserDefaults suite, no onboarding). If nothing in the app reads it, the UI test depends on whatever state the last run left behind and will flake.
Recipe: a Jetpack Compose screen with UI-test evidence
Section titled “Recipe: a Jetpack Compose screen with UI-test evidence”A correct Compose test finds nodes by tag, not by text, so a copy change does not break it:
@get:Rule val composeRule = createAndroidComposeRule<MainActivity>()
@Testfun incrementUpdatesTotal() { composeRule.onNodeWithTag("cart.line.0.increment").performClick() composeRule.onNodeWithTag("cart.total").assertTextEquals("Total: 24.00")}For visual evidence on Android without an emulator, add a screenshot-test library such as Paparazzi: ./gradlew :app:recordPaparazziDebug records reference images, and ./gradlew :app:verifyPaparazziDebug fails when the rendering changes. On iOS, swift-snapshot-testing plays the same role with assertSnapshot(of: view, as: .image). Both record references on the first run, so review and commit those references yourself.
Recipe: reproduce a crash the agent cannot see
Section titled “Recipe: reproduce a crash the agent cannot see”A crash report with a stack trace is a spec for a failing test. Hand it over with the device server connected:
The order matters. A test that failed before the fix and passes after it is evidence. A fix with no failing test first is a guess that happened to compile.
How to verify the agent’s native work without reading every line
Section titled “How to verify the agent’s native work without reading every line”The scripts turn every change into the same bundle of artifacts, and review starts from that bundle, not from the diff:
| Evidence | Where it lives | What a reviewer checks |
|---|---|---|
| Test verdict | Last lines of the script output, summary.json, JUnit XML under androidTest-results/ | exit=0 on both platforms; the count of tests went up, never down |
| UI screenshots | tests.xcresult attachments (open in Xcode), Paparazzi or snapshot-testing images | The screens match the spec; snapshot diffs are intentional |
| Runtime logs | xcodebuild.log, logcat.txt | No new crashes, ANRs, or error spam during the UI tests |
| Test changes | The test files in the diff | New tests assert the acceptance criteria; no assertion was weakened, no snapshot re-recorded without a reason |
Three rules make this trustworthy:
-
CI runs the same scripts. A green result in the agent’s session is a claim; the same script in CI on a clean macOS runner is the proof. Keep the scripts in the repository so the two cannot drift.
-
The agent cannot edit the oracle. Snapshot references and the tests that encode the acceptance criteria are protected, as described in protecting the oracle. Claude Code gets per-path deny rules (the Claude Code tab above). For Codex and Cursor this page relies on a tool-neutral guard instead, and it backs up Claude Code too: CODEOWNERS on the reference folders, plus a CI step that fails when they change without a label.
# .github/CODEOWNERS**/__Snapshots__/ @your-org/mobile-leads**/src/test/snapshots/ @your-org/mobile-leads# CI job step (check out with fetch-depth: 0 so origin/main is available)- name: Block unlabelled snapshot changesif: ${{ !contains(github.event.pull_request.labels.*.name, 'snapshot-update') }}run: |if git diff --name-only origin/main... | grep -E '__Snapshots__/|src/test/snapshots/'; thenecho "Snapshot references changed without the snapshot-update label"; exit 1fi -
A named person signs off on the evidence. The developer who assigned the task reviews the evidence bundle and the test diff. The code diff gets a full human read only where risk requires it: signing, entitlements, payments, and anything touching
Info.plistpermissions orAndroidManifest.xml.
The pull-request contract for this bundle is described in the evidence bundle, and the reasoning for reviewing evidence instead of diffs in reading evidence instead of code.
When native agent loops break
Section titled “When native agent loops break”Model and effort for native work
Section titled “Model and effort for native work”Start on your tool’s default model and raise effort before you switch model; current defaults and prices are on the models hub. Native tasks are more often limited by build feedback than by the model, so fix the loop first.