Skip to content

Native iOS and Android Recipes

Native iOS and Android development with a coding agent works when each platform has one script that builds the app, runs its tests on a pinned simulator or emulator, and saves the results as evidence. Claude Code, Codex, and Cursor iterate against that script, and an MCP server such as MobileBuildMCP or mobile-mcp lets the agent see the running app.

You ask the agent for a new SwiftUI settings screen. It writes clean Swift, announces “done”, and the project does not build, because the new file never reached the Xcode target. On the Android side, the agent runs ./gradlew clean build on every change, waits minutes for a result, and pastes thousands of lines of Gradle output into its own context. Neither agent ever looked at the screen it built. This page fixes those three gaps: build feedback, speed, and proof.

What you get from an agent-ready native project

Section titled “What you get from an agent-ready native project”
  • Two copy-paste scripts, scripts/ios-check.sh and scripts/android-check.sh. Each one boots a pinned device, runs the tests, writes a short summary for the agent, and saves the full evidence under build/evidence/.
  • An instructions block for AGENTS.md, CLAUDE.md, or a Cursor rule that makes “the script exits 0” the definition of done.
  • Verified install commands for the two mobile MCP servers, plus the Codex plugins for iOS and Android, per tool.
  • Build-time rules that keep the agent’s inner loop short: narrow test selection, warm devices, and caches.
  • Recipes for a SwiftUI screen, a Jetpack Compose screen, and a crash reproduction, each ending in UI-test evidence rather than a diff you have to read.

Why native apps need a different agent loop

Section titled “Why native apps need a different agent loop”

Web agents get cheap feedback: a dev server reloads in a second and a browser MCP shows the page. Native projects take longer to give that feedback, for three reasons:

  1. The project file is part of the build. Xcode decides what compiles from the target membership in project.pbxproj, and Gradle from the module layout. An agent that writes a correct file in the wrong place gets a build error or, worse, a file that silently never compiles.
  2. Builds are slow and loud. A clean xcodebuild or Gradle build prints thousands of lines. Streamed into the agent’s context, that output costs tokens and pushes the task instructions out of view.
  3. The UI lives in a separate process. The simulator and the emulator are separate programs. Without a test runner or an MCP server, the agent has no way to know whether the screen renders, and neither do you.

So the recipes on this page do not start with a feature prompt. They start by giving the agent a single, quiet, reliable command per platform, and they judge the agent by what that command proves.

  1. Pin one simulator and one emulator. List what exists with xcrun simctl list devices available and emulator -list-avds (the emulator binary lives in $ANDROID_HOME/emulator). Pick one of each and write the names into the scripts below. A pinned device means a test failure is about your code, not about which phone the agent happened to boot.

  2. Add scripts/ios-check.sh. It keeps DerivedData inside the repository, writes the full log to a file, and prints only the lines the agent needs:

    #!/usr/bin/env bash
    # scripts/ios-check.sh — build, test and collect evidence on one pinned simulator.
    # Usage: scripts/ios-check.sh [TestTarget/TestClass ...] (no argument = all tests)
    set -euo pipefail
    SCHEME="${SCHEME:-App}"
    SIM="${SIM:-iPhone 17}"
    OUT="build/evidence/ios"
    rm -rf "$OUT" && mkdir -p "$OUT"
    ONLY=()
    for t in "$@"; do ONLY+=("-only-testing:$t"); done
    xcrun simctl bootstatus "$SIM" -b >/dev/null # boots the simulator if needed, waits until ready
    set +e
    xcodebuild test \
    -scheme "$SCHEME" \
    -destination "platform=iOS Simulator,name=$SIM" \
    -derivedDataPath build/DerivedData \
    -resultBundlePath "$OUT/tests.xcresult" \
    ${ONLY[@]+"${ONLY[@]}"} > "$OUT/xcodebuild.log" 2>&1
    STATUS=$?
    set -e
    xcrun xcresulttool get test-results summary --path "$OUT/tests.xcresult" > "$OUT/summary.json" 2>/dev/null || true
    grep -E "error:|\*\* (BUILD|TEST) (SUCCEEDED|FAILED)|failed" "$OUT/xcodebuild.log" | tail -40
    echo "exit=$STATUS evidence=$OUT"
    exit $STATUS

    The ${ONLY[@]+...} form matters: macOS still ships Bash 3.2, where an empty array under set -u aborts the script. xcresulttool get test-results needs Xcode 16 or later.

  3. Add scripts/android-check.sh. It reuses a running emulator, clears logcat so the log covers only this run, and copies the Gradle test reports into the evidence folder:

    #!/usr/bin/env bash
    # scripts/android-check.sh — unit + instrumented tests on one pinned emulator, evidence collected.
    # Usage: scripts/android-check.sh [com.example.app.SomeUiTest] (no argument = all tests)
    set -euo pipefail
    AVD="${AVD:-Pixel_8_API_35}"
    MODULE="${MODULE:-app}"
    OUT="build/evidence/android"
    rm -rf "$OUT" && mkdir -p "$OUT"
    if ! adb get-state >/dev/null 2>&1; then
    emulator -avd "$AVD" -no-window -no-audio -no-boot-anim > "$OUT/emulator.log" 2>&1 &
    adb wait-for-device
    until [ "$(adb shell getprop sys.boot_completed | tr -d '\r')" = "1" ]; do sleep 2; done
    fi
    adb logcat -c
    FILTER=()
    if [ $# -gt 0 ]; then FILTER=("-Pandroid.testInstrumentationRunnerArguments.class=$1"); fi
    set +e
    ./gradlew ":$MODULE:testDebugUnitTest" ":$MODULE:connectedDebugAndroidTest" \
    ${FILTER[@]+"${FILTER[@]}"} > "$OUT/gradle.log" 2>&1
    STATUS=$?
    set -e
    adb logcat -d > "$OUT/logcat.txt"
    cp -R "$MODULE/build/reports" "$OUT/reports" 2>/dev/null || true
    cp -R "$MODULE/build/outputs/androidTest-results" "$OUT/androidTest-results" 2>/dev/null || true
    grep -E "FAILED|BUILD (SUCCESSFUL|FAILED)|e: " "$OUT/gradle.log" | tail -40
    echo "exit=$STATUS evidence=$OUT"
    exit $STATUS

    The optional argument narrows only the instrumented run to one class; unit tests always run in full, because they are cheap.

  4. Tell the agent the rules. Paste this block into AGENTS.md for Codex, CLAUDE.md for Claude Code, or a project rule for Cursor. Claude Code reads AGENTS.md when the repository has no CLAUDE.md (v2.1.277 or later, on the latest release channel as of 2026-09-26; stable users still need a CLAUDE.md that points to AGENTS.md), so one file can serve both CLIs. If you have not written an instructions file before, start with AGENTS.md and CLAUDE.md:

    ## Native build and test loop
    - iOS: run `scripts/ios-check.sh`, optionally with a test class such as `AppUITests/SettingsUITests`.
    Never call `xcodebuild` without `-derivedDataPath build/DerivedData`.
    - Android: run `scripts/android-check.sh`, optionally with a fully qualified test class.
    Never run `./gradlew clean` unless the user asks.
    - A task is done only when the relevant script prints `exit=0`. Quote its summary lines in your final message.
    - Every control a test touches gets a stable ID: `.accessibilityIdentifier(...)` in SwiftUI,
    `Modifier.testTag(...)` in Compose. Never locate elements by visible text alone.
    - Never edit snapshot references (`__Snapshots__/`, `src/test/snapshots/`) or weaken an existing
    assertion to make a test pass. Stop and report the failure instead.
    - New Swift files go inside an existing synchronized folder; do not edit `project.pbxproj` by hand.
  5. Commit the scripts and run them once yourself. A script that fails on a clean checkout teaches the agent to ignore it. Run both, open build/evidence/, and add build/ to .gitignore if it is not there yet.

Connect the agent to the simulator and emulator

Section titled “Connect the agent to the simulator and emulator”

Tests prove the behaviour you wrote down. For exploratory work (does the screen look right, what happens when I tap here), the agent needs to see and drive the device. Two MCP servers cover this. If you have never added an MCP server, read the introduction to MCP first:

  • MobileBuildMCP (formerly XcodeBuildMCP, maintained by Sentry) builds, runs, and debugs iOS and macOS projects on simulators and devices. It needs macOS 14.5+ and Xcode 16+. Its most-used tool is build_run_sim, which builds the scheme, installs the app on a booted simulator, launches it, and captures logs.
  • mobile-mcp (Mobile Next) drives iOS simulators, Android emulators, and real devices through one tool set: mobile_list_elements_on_screen for an accessibility snapshot, mobile_take_screenshot, mobile_click_on_screen_at_coordinates, mobile_get_device_logs, mobile_get_crash, and more. Version 1.0.5 advertises 32 tools.

Popularity: on 2026-09-26 MobileBuildMCP had 6.4k GitHub stars and npm mobilebuildmcp 2.7.1 (published 2026-09-23); on 2026-09-28 mobile-mcp had 7,958 GitHub stars, with npm @mobilenext/mobile-mcp 1.0.5 (published 2026-09-23). Sources: the GitHub repositories getsentry/MobileBuildMCP and mobile-next/mobile-mcp, and the npm registry.

Add the servers from the terminal. The default local scope keeps them out of teammates’ sessions until you decide to share them with -s project:

Terminal window
claude mcp add mobilebuild -- npx -y mobilebuildmcp@latest mcp
claude mcp add mobile-mcp -- npx -y @mobilenext/mobile-mcp@latest

The Claude Code desktop app on macOS also has an iOS Simulator pane (beta since July 2026), so you can watch the simulator next to the session. To keep the agent away from snapshot references, add deny rules to .claude/settings.json:

{
"permissions": {
"deny": [
"Edit(**/__Snapshots__/**)",
"Edit(**/src/test/snapshots/**)"
]
}
}

When to use which. Use MobileBuildMCP when the problem is the Xcode build itself (schemes, signing, a Swift macro, a crash on launch). Use mobile-mcp when the problem is behaviour on screen, on either platform. Enable only the server a session needs: every advertised tool definition takes context on every turn. mobile-mcp 1.0.5 advertises all 32 tools at once. MobileBuildMCP exposes only its simulator workflow by default; widen it with the MOBILEBUILDMCP_ENABLED_WORKFLOWS environment variable (for example simulator,ui-automation,debugging,logging) only when the task needs UI automation or the debugger. In Claude Code that is claude mcp add mobilebuild -e MOBILEBUILDMCP_ENABLED_WORKFLOWS=simulator,ui-automation -- npx -y mobilebuildmcp@latest mcp. Codex’s build-ios-apps plugin sets XCODEBUILDMCP_ENABLED_WORKFLOWS instead only because it runs the old xcodebuildmcp@latest package; mobilebuildmcp 2.7.1 ignores that name.

Working inside Xcode instead of a terminal. MobileBuildMCP’s client documentation (checked 2026-09-26) describes Xcode 26.3 or later hosting the Codex and Claude Code agents under Xcode Settings > Intelligence. Those in-Xcode agents start with a minimal PATH, so a server launched with npx often fails to start; the documentation wraps the command in /bin/zsh -lc with Homebrew and nvm on the PATH. The check scripts on this page work the same from either place.

How to keep native build times inside the agent’s loop

Section titled “How to keep native build times inside the agent’s loop”

A slow loop changes agent behaviour: the agent batches many edits between builds, and a failure then has many possible causes. Keep three loops, each with its own cost:

LoopWhat runsCommandWhen the agent runs it
InnerUnit tests of the changed module (Swift Testing or XCTest; JUnit)xcodebuild test ... -only-testing:AppTests/CartTests · ./gradlew :feature:cart:testDebugUnitTest --tests '*CartViewModelTest'After every edit
MiddleOne UI test class on the pinned devicescripts/ios-check.sh AppUITests/CartUITests · scripts/android-check.sh com.example.cart.CartScreenTestBefore it claims “done”
OuterEvery suite, snapshot verification, both platformsThe same two scripts without arguments, in CIOn every pull request

Rules that keep the inner and middle loops fast:

  • Never clean by reflex. Tell the agent in its instructions file that clean is forbidden unless you ask. Most “fix it with a clean build” instincts come from stale DerivedData or Gradle caches, and the failure list below has the targeted fix for them.
  • Keep devices warm. Both scripts boot the device only when it is not running. Leave the simulator and the emulator up for the whole session.
  • Turn on Gradle’s caches. Add org.gradle.caching=true and org.gradle.configuration-cache=true to gradle.properties. The configuration cache skips re-running build scripts when nothing in them changed.
  • Split build from test on iOS when you rerun. xcodebuild build-for-testing followed by xcodebuild test-without-building reruns a flaky test without recompiling. Use it for triage only: after a code change the agent must rebuild.
  • Measure before tuning. xcodebuild -showBuildTimingSummary prints where the time went, and ./gradlew --profile writes a local HTML report under build/reports/profile/. Ask the agent to read the report and propose one change, then measure again.

Recipe: a SwiftUI screen with UI-test evidence

Section titled “Recipe: a SwiftUI screen with UI-test evidence”

The prompt below asks for the behaviour first, as tests, and makes the script’s exit code the finish line.

The screenshots land inside build/evidence/ios/tests.xcresult, next to the test that took them. The attachment code the agent should produce looks like this:

let app = XCUIApplication()
func snap(_ name: String) {
let shot = XCTAttachment(screenshot: app.screenshot())
shot.name = name
shot.lifetime = .keepAlways // keep it even when the test passes
add(shot)
}
func testEnablingReminderEnablesTimePicker() {
app.launchArguments = ["-uiTesting"]
app.launch()
app.buttons["settings.open"].tap()
snap("1-settings-open")
app.switches["settings.reminderToggle"].tap()
snap("2-reminder-on")
XCTAssertTrue(app.datePickers["settings.reminderTime"].isEnabled)
snap("3-time-picker-enabled")
}

Check one thing the agent often gets wrong: the -uiTesting launch argument has to change app behaviour (an in-memory UserDefaults suite, no onboarding). If nothing in the app reads it, the UI test depends on whatever state the last run left behind and will flake.

Recipe: a Jetpack Compose screen with UI-test evidence

Section titled “Recipe: a Jetpack Compose screen with UI-test evidence”

A correct Compose test finds nodes by tag, not by text, so a copy change does not break it:

@get:Rule val composeRule = createAndroidComposeRule<MainActivity>()
@Test
fun incrementUpdatesTotal() {
composeRule.onNodeWithTag("cart.line.0.increment").performClick()
composeRule.onNodeWithTag("cart.total").assertTextEquals("Total: 24.00")
}

For visual evidence on Android without an emulator, add a screenshot-test library such as Paparazzi: ./gradlew :app:recordPaparazziDebug records reference images, and ./gradlew :app:verifyPaparazziDebug fails when the rendering changes. On iOS, swift-snapshot-testing plays the same role with assertSnapshot(of: view, as: .image). Both record references on the first run, so review and commit those references yourself.

Recipe: reproduce a crash the agent cannot see

Section titled “Recipe: reproduce a crash the agent cannot see”

A crash report with a stack trace is a spec for a failing test. Hand it over with the device server connected:

The order matters. A test that failed before the fix and passes after it is evidence. A fix with no failing test first is a guess that happened to compile.

How to verify the agent’s native work without reading every line

Section titled “How to verify the agent’s native work without reading every line”

The scripts turn every change into the same bundle of artifacts, and review starts from that bundle, not from the diff:

EvidenceWhere it livesWhat a reviewer checks
Test verdictLast lines of the script output, summary.json, JUnit XML under androidTest-results/exit=0 on both platforms; the count of tests went up, never down
UI screenshotstests.xcresult attachments (open in Xcode), Paparazzi or snapshot-testing imagesThe screens match the spec; snapshot diffs are intentional
Runtime logsxcodebuild.log, logcat.txtNo new crashes, ANRs, or error spam during the UI tests
Test changesThe test files in the diffNew tests assert the acceptance criteria; no assertion was weakened, no snapshot re-recorded without a reason

Three rules make this trustworthy:

  1. CI runs the same scripts. A green result in the agent’s session is a claim; the same script in CI on a clean macOS runner is the proof. Keep the scripts in the repository so the two cannot drift.

  2. The agent cannot edit the oracle. Snapshot references and the tests that encode the acceptance criteria are protected, as described in protecting the oracle. Claude Code gets per-path deny rules (the Claude Code tab above). For Codex and Cursor this page relies on a tool-neutral guard instead, and it backs up Claude Code too: CODEOWNERS on the reference folders, plus a CI step that fails when they change without a label.

    # .github/CODEOWNERS
    **/__Snapshots__/ @your-org/mobile-leads
    **/src/test/snapshots/ @your-org/mobile-leads
    # CI job step (check out with fetch-depth: 0 so origin/main is available)
    - name: Block unlabelled snapshot changes
    if: ${{ !contains(github.event.pull_request.labels.*.name, 'snapshot-update') }}
    run: |
    if git diff --name-only origin/main... | grep -E '__Snapshots__/|src/test/snapshots/'; then
    echo "Snapshot references changed without the snapshot-update label"; exit 1
    fi
  3. A named person signs off on the evidence. The developer who assigned the task reviews the evidence bundle and the test diff. The code diff gets a full human read only where risk requires it: signing, entitlements, payments, and anything touching Info.plist permissions or AndroidManifest.xml.

The pull-request contract for this bundle is described in the evidence bundle, and the reasoning for reviewing evidence instead of diffs in reading evidence instead of code.

Start on your tool’s default model and raise effort before you switch model; current defaults and prices are on the models hub. Native tasks are more often limited by build feedback than by the model, so fix the loop first.

Where to go next with native iOS and Android agents

Section titled “Where to go next with native iOS and Android agents”