Mobile App Performance: The Five Metrics That Decide Retention and Store Ratings
Performance work in mobile development wastes more effort than almost any other kind of engineering. Teams burn sprints shaving milliseconds off benchmarks that no user will ever feel. Meanwhile the problems that actually make people uninstall, the freezes, the janky feeds, the phone that’s dead by 4pm, sit in the backlog because nobody put a number on them.
The gap between what’s easy to measure and what users care about is wide. Closing it starts with picking the right metrics.
The list that matters is short. Cold start time. ANR rate on Android (and hang rate on iOS). Crash rate. Scroll smoothness in list views. Battery drain. Almost everything else is either a stand-in for one of these or a number engineers like to look at.
Why most performance dashboards lie
Start with the numbers that look good in a slide deck and mean little. Average launch time on a flagship test phone. Synthetic benchmark scores. Bundle size tracked to the kilobyte with no link to install conversion. Memory usage on a device with 16 GB of RAM. Lighthouse-style scores borrowed from the web.
None of these are useless. The trouble is what they leave out. Averages hide the tail, and the tail is where users leave. A median cold start of 1.2 seconds can sit on top of a bottom quartile at 5 seconds, and that bottom quartile is full of real people on three-year-old phones with half-full storage. They’re the ones writing the one-star reviews.
So two rules apply to every metric below. Measure on physical devices that look like your users’ devices, not your team’s. And look at percentiles, p75 and p90 at least, never just the mean.
Cold start time
Cold start is the time from tapping the icon to the first moment the user can do something useful. For most apps it’s the single most important performance number. It’s also the one teams most often measure wrong.
The usual mistake is stopping the clock at first render. An app that draws a spinner in 200 milliseconds and then waits three seconds on an API call has a three-second cold start as far as the user is concerned. The metric has to run from tap to usable. On Android that’s the gap between time to initial display and time to full display, and the platform lets you report the second one yourself with reportFullyDrawn(). Use it. On iOS, MetricKit and the Xcode Organizer give you launch time from the field, and Instruments’ App Launch template shows you where the time goes on a device in your hand.
Where does the time go? Usually somewhere boring. SDK initialisation is the classic offender: analytics, ads, crash reporting, feature flags and A/B frameworks all want to start in Application.onCreate or didFinishLaunching, and each one costs a little. Add dependency injection graphs built eagerly, a database migration run on the main thread, and a blocking network call for config before the first screen, and you’ve got a slow app made entirely of reasonable decisions.
The fixes are well known:
Defer anything that isn’t needed for the first screen. Lazy-load SDKs, or initialise them after first frame.
Show cached content immediately and refresh in the background, so the first screen never waits on the network.
On Android, ship Baseline Profiles. They let the runtime precompile the code paths used at startup and on critical journeys, and they’re one of the cheapest wins available. Measure the result with the Macrobenchmark library instead of trusting a stopwatch.
On iOS, cut down dynamic frameworks, avoid heavy work in static initialisers, and watch for anything that runs before main.
As a budget, aim for under two seconds to usable on a mid-range device. Then check your slowest quartile of real devices. If it’s well over that, you have a problem, whatever the average says.
ANRs, hangs and crashes
On Android, an ANR (Application Not Responding) fires when the main thread stays blocked long enough that the system gives up on it. For input events that’s five seconds. The user sees a dialog asking whether to wait or close the app. Most close it.
ANRs matter twice. Users hate them, and Google Play tracks them. Android Vitals reports a user-perceived ANR rate, and Play Console sets “bad behavior” thresholds for it and for crash rate, both overall and per device model. Cross them and Play can reduce your app’s visibility and show a warning on your store listing. That’s a direct hit to organic installs, which makes ANR rate a growth metric as much as an engineering one.
The cause is almost always the same: blocking work on the main thread. Disk reads, SharedPreferences commits, database queries, network calls, JSON parsing of a big payload, image decoding. The fix is also the same. Move it off the main thread with coroutines, executors or WorkManager, and keep the UI thread for drawing and handling input. StrictMode in debug builds will flag a lot of this before it ships.
The harder part is diagnosis. Many ANRs never show up on a developer’s phone. They depend on slow storage, a busy CPU, a lock held by some other thread, or a broadcast receiver that fires at the wrong moment. You need production traces. Android Vitals, Firebase Crashlytics and Bugsnag all collect ANR stack traces from the field, and the ApplicationExitInfo API lets your own app find out why its last process died.
iOS doesn’t have an ANR dialog, but it has the same problem under a different name. Xcode Organizer and MetricKit report hang rate, time the main thread spent unresponsive in the field. Treat it the way Android teams treat ANRs. The watchdog will also kill an app that takes too long to launch or to respond to certain system events, and those terminations look like crashes to the user.
Crash rate belongs in the same bucket. It’s the most obvious failure there is, it feeds the same Play thresholds, and it’s the metric store reviewers mention by name. Track crash-free users, not just crash-free sessions. A crash that hits a small group of users every day is worse than its session rate suggests, because those users are the ones who leave.
Scroll performance
List views are where users spend most of their time. The news feed, the timeline, the product grid, the chat history. That’s where people decide whether an app feels good, and dropped frames there are instantly visible.
The frame budget is tight. At 60 frames per second each frame gets about 16.7 milliseconds. At 120 Hz, which is now normal on flagship phones and common in the mid-range, it’s about 8.3. Miss it and the scroll stutters.
The usual causes:
Expensive layout. Deeply nested view hierarchies, constraint layouts solving more than they need to, or Compose and SwiftUI views that recompose or re-render far more often than they should.
Overdraw. Layers of opaque backgrounds painted on top of each other, each costing GPU time for pixels nobody sees.
Image work on the main thread. Decoding a full-size JPEG to show it as a 120-pixel thumbnail is the textbook case. Use a proper image loader, decode off-thread, and request images at the size you’ll display them.
Synchronous data access while binding cells. A database read or a string format in onBindViewHolder or cellForRowAt adds up fast across a long list.
Unstable item identity. If the list can’t tell which item is which, it rebuilds rows it could have reused. Stable keys in Compose, DiffUtil in RecyclerView and diffable data sources on iOS all exist for this.
The fixes are well understood. The hard part is noticing the problem. Scroll jank rarely shows up in a synthetic benchmark with ten tidy items. It shows up with real data: a thousand items, mixed cell types, images of every size, text that wraps differently in German. Test with production-shaped data. On Android, JankStats and Macrobenchmark’s frame timing metrics give you numbers, and Perfetto shows you exactly which frame missed and why. On iOS, Instruments’ Animation Hitches template and MetricKit’s hitch rate do the same job.
Battery drain
Battery is the metric users notice least directly and complain about most. Nobody opens your app and thinks “this is using too much power.” They notice their phone dying earlier. Eventually some of them open the battery screen, see your app near the top, and uninstall it. That user may have been perfectly happy with everything else.
Both platforms now make this easier to spot and harder to get away with. iOS shows battery use per app and restricts background work aggressively. Android flags excessive wake locks and background activity in Vitals and limits apps that misbehave.
The main sources of drain:
Location. This is the one most often handled badly. Apps ask for continuous, high-accuracy location when they only need a single coarse fix, or keep location updates running in the background with no visible benefit to the user. Request the lowest accuracy that works, stop updates as soon as you have what you need, and use geofencing or significant-change APIs when you only care about big moves.
Network chatter. Lots of small requests keep the radio awake far longer than one batched request would. Batch, cache, and let the OS schedule non-urgent sync through WorkManager or BGTaskScheduler.
Background processing. Polling, long-running services and wake locks held “just in case.” Use push to trigger work, and let the system decide when deferred work runs.
CPU-heavy work. Video processing, ML inference, crypto, heavy animation. Some of it is the point of the app. Make sure the rest isn’t running when the screen is off.
Measure it on device. Android’s Battery Historian and Perfetto power traces, and the Energy Log in Xcode, show where the power goes. Field data from Android Vitals and MetricKit tells you whether real users are paying for it.
Where to put the effort
Pick the metrics above, put field numbers on them, and look at the slow end of the distribution. Set budgets: two seconds to usable, zero main-thread disk or network access, Play’s bad-behavior thresholds as the ceiling you never touch, a frame budget that matches the screens your users own. Then wire those budgets into CI with Macrobenchmark or XCTest performance tests so a regression fails the build instead of reaching the store.
Most of the fixes aren’t clever. They’re the same handful of moves, applied to the right problems, and checked against what real users experience.
Measure what users feel. Fix what costs them.
Project Articles
How Long a New HyperCrux Engine in Go Would Take: About Two Weeks, Task by Task
Matching Candidates to Jobs and Deleting Them Cleanly: A Small Recruiting Agency on HyperCrux
At 3 a.m., Find the Incident That Looks Like This One: On-Call Notes With Links and Similarity
Finding the Shot You Half Remember: A Photographer’s Archive Searched by Image Embeddings
An AI Assistant’s Memory That Forgets Properly: People, Facts and Embeddings in One Transaction
BareProxy Plugins Go WebAssembly, Starting With RenderCache and AI Crawler Control
BareProxy Runs on Linux, macOS and Windows, and One Plugin File Runs on All of Them
BareProxy Is an Operating System for Web Traffic: A Small Core and a Long Stream of Plugins
BareProxy Markdown Serving Plugin: Markdown Files Rendered to HTML With No Build Step
BareProxy 0.2.0 Runs WebAssembly Plugins Through the Proxy-Wasm Interface
AltSql Core and AltSql DB Are Now Open Source Under the Apache License 2.0
AltSql DB 0.2: SQL and Direct Calls Wrote Byte-Identical Files Over 100,000 Random Steps
AltSql DB Keeps a Whole Fleet in One File and Reads by Key 5 Times Faster Than SQLite Through SQL
AltSql Is a Native Hybrid of Key-Value and SQL, From the Sensor to the Gateway
Graph Key Builder Shows the Keys a Graph Takes in RocksDB, LMDB or etcd, and What Each Hop Costs
Tampered Log Case Study: 12 of 12 Edits to an AI Agent’s Sealed Record Caught at the Exact Line
DNS Leak Case Study: A Stand-In VPN Client Sends 81 Queries to the Home Router While It Reconnects
The Next VPN Works Prototypes: A 14-Week Beta for AI Agents and a Real Gateway Pilot for Scope
Stolen VPN Login Case Study: Rules Learned From Two Weeks of Traffic Stop 136 of 144 Attempts
Precomputing 0.1.1 Alpha Is Now Open Source Under the Apache License 2.0
Wikipedia Live: Every Change to Every Wiki, Counted in the Browser
MCP Case Study: An AI Agent Gets Its Answers in a Few Hundred Tokens Each
Incident Déjà Vu Case Study: Every Returning Incident Named Within Ten Seconds
Agent Traces Case Study: A Day of Coding-Agent Calls Kept 14 Times Smaller and Rebuilt Byte for Byte
Preconfig Doctor Reads a Failed Agent Setup, Names the Cause and Writes the Fix Into the Spec
What Sets Preconfiguration Apart: The Machine Layer, Checked and Proven, Across Every Agent Platform
The Next Preconfiguration Prototype: A 12-Week Beta on the Agent Platforms Themselves
A Raspberry Pi on a Patchy Connection: Readings Wait in a ukue File Until the Internet Is Back
What Is a Job Queue? Background Jobs, Retries and Dead Letters Explained
Photo Upload Thumbnails With ukue: Previews First, Print Sizes When the CPU Is Free
Signup Emails From a Python Web App, Sent in the Background With ukue
Checking 300 Websites Every Night With ukue, Cron and a Shell Script