On-Device AI Models for Mobile: What Actually Runs Offline
Which model sizes fit on a phone, what tasks they can handle offline, and where cloud fallback is still required.
Jump to 7 sections
Models in the 1-4 billion parameter range now run acceptably on recent flagship phones for tasks like summarization, basic classification, and simple assistants. Anything requiring broad world knowledge or complex reasoning still needs a cloud fallback.
On-device AI got a real capability jump over the past two years as phone chipsets added dedicated neural processing hardware and models shrank without losing as much capability as they used to. This review covers what actually runs offline on current hardware, not what a chipset vendor claims in a keynote.
It is written for mobile developers deciding whether a feature can run on-device or needs a cloud API, not for people optimizing model architectures themselves.
This is also a fast-moving area of mobile hardware, with each generation of chipset bringing meaningfully more dedicated AI processing capability than the last. A model that felt genuinely constrained on last year's hardware can run comfortably on the current generation, which means this evaluation is worth repeating periodically rather than treating a single hardware assessment as permanent.
In this article: What Size Model Actually Fits on a Phone · What On-Device Models Are Actually Good For · Battery and Thermal Reality vs. Benchmark Numbers · On-Device vs. Cloud: The Practical Tradeoffs · The Hybrid Pattern Most Apps Are Landing On · Testing an On-Device Model on Real Hardware
What Size Model Actually Fits on a Phone
Current flagship phones can run models in the 1-4 billion parameter range at usable speed, thanks to dedicated neural processing units and aggressive quantization that shrinks a model's memory footprint with a modest, often acceptable accuracy tradeoff. Below 1 billion parameters, models run fast but lose enough capability to feel noticeably limited for anything beyond simple classification. See MLCommons: MLCommons' benchmarking work.
Mid-range phones lag flagships meaningfully here — the same 3-billion-parameter model that runs smoothly on a top-tier device can be noticeably sluggish or thermally throttled on hardware two years older, which matters for any app targeting a broad device range rather than just the newest phones.
What On-Device Models Are Actually Good For
On-device models handle summarization, basic classification, simple rewriting, and narrow assistant tasks — anything that leans on pattern-matching over a bounded input rather than broad world knowledge. A keyboard suggesting a reply, a notes app summarizing a long entry, or a camera app tagging a photo's contents are all realistic on-device tasks today. See Stanford HAI's AI Index: Stanford HAI's AI Index report.
They fall short on anything requiring current events, obscure factual knowledge, or multi-step reasoning across a large context — the smaller a model gets, the more its knowledge is compressed and the more gaps show up on anything outside its core training distribution. For more on this, see Model Drop's small models that run on one GPU.
Battery and Thermal Reality vs. Benchmark Numbers
A benchmark score measured on a cool, idle device does not predict how a model performs during sustained real-world use, where thermal throttling kicks in after a few minutes of continuous inference and measurably slows response time. This is the gap between a vendor's demo numbers and what a user actually experiences during a long session.
We have found the more useful test is running a model for 10-15 minutes of continuous use, not a single quick benchmark pass, since that is closer to how a real feature — live transcription, for instance — actually gets used. For more on this, see Model Drop's self-hosting versus API inference.
On-Device vs. Cloud: The Practical Tradeoffs
| Factor | On-device | Cloud API |
|---|---|---|
| Privacy | Data never leaves the device | Data sent to a server |
| Works offline | Yes | No |
| Capability ceiling | Limited by phone hardware | Much higher |
| Cost per request | None after app download | Per-token or per-request |
| Battery/thermal impact | Real, sustained-use dependent | Minimal on-device |
The Hybrid Pattern Most Apps Are Landing On
The common production pattern now runs a small model on-device for fast, simple, privacy-sensitive tasks, and falls back to a cloud API for anything the on-device model flags as uncertain or outside its scope. This gets most of the latency and privacy benefit of on-device inference without giving up cloud-level capability for harder requests.
Model Drop's guide to self-hosting versus API inference covers the equivalent tradeoff on the server side, which is worth reading alongside this one if the cloud-fallback portion of a hybrid app is being built from scratch rather than bought from a vendor. For more on this, see Model Drop's mixture-of-experts versus dense models.
Testing an On-Device Model on Real Hardware
Test on the actual range of devices your user base runs, not just the newest flagship phone sitting on a developer's desk. The gap between a two-year-old mid-range device and a current flagship is large enough to change whether a feature is usable at all, not just how fast it feels.
Measure performance after ten to fifteen minutes of continuous use, not on the first request. Thermal throttling is the gap between a benchmark number and what a real user experiences during sustained use, and it rarely shows up in a quick initial test.
Track battery drain over a typical session length alongside speed, since a feature that is fast but noticeably drains battery will get uninstalled just as quickly as one that is simply too slow.
Instrument the feature in production with real device and OS version data from day one, not just app-store aggregate crash reports. On-device AI failure modes correlate strongly with specific chipset and OS version combinations in ways that are hard to predict from lab testing alone.
Consider a staged rollout for any new on-device model, starting with newer flagship devices before expanding to the full device range. This limits the blast radius of any unexpected performance or thermal issue to a smaller, generally more forgiving user segment while the rollout is still being monitored closely.
Conclusion
On-device AI on mobile has crossed a real usability threshold for narrow, bounded tasks — summarization, simple classification, basic assistants — but still needs a cloud fallback for anything requiring broad knowledge or complex reasoning. Model Drop's recommendation is to test any on-device model under sustained real use, not a quick benchmark, and to design the cloud-fallback path from day one rather than bolting it on after launch.
Start by mapping which of your app's AI features are genuinely privacy-sensitive or need offline support — those are the strongest candidates for on-device, everything else can likely stay cloud-first.
One last practical note: involve your mobile engineering team early in any on-device AI evaluation, since hardware and OS-specific quirks are often easier to catch from someone who works with the platform daily than from a model-evaluation process run in isolation from the actual app codebase.
Keep in mind that app store review guidelines occasionally place specific constraints on background AI processing and battery usage, so it is worth checking current platform policies before finalizing an on-device AI feature's technical design, not just after implementation is complete.
Model Drop covers AI launches, models, tools, and platforms for developers and builders tracking the frontier.