Models

On-Device AI Models for Mobile: What Actually Runs Offline

Which model sizes fit on a phone, what tasks they can handle offline, and where cloud fallback is still required.

Dana Kwon

Contributing Reviewer

Published 6 min read
Close-up of a smartphone showing a chat app interface on a wooden table.
Jump to 7 sections

Models in the 1-4 billion parameter range now run acceptably on recent flagship phones for tasks like summarization, basic classification, and simple assistants. Anything requiring broad world knowledge or complex reasoning still needs a cloud fallback.

On-device AI got a real capability jump over the past two years as phone chipsets added dedicated neural processing hardware and models shrank without losing as much capability as they used to. This review covers what actually runs offline on current hardware, not what a chipset vendor claims in a keynote.

It is written for mobile developers deciding whether a feature can run on-device or needs a cloud API, not for people optimizing model architectures themselves.

This is also a fast-moving area of mobile hardware, with each generation of chipset bringing meaningfully more dedicated AI processing capability than the last. A model that felt genuinely constrained on last year's hardware can run comfortably on the current generation, which means this evaluation is worth repeating periodically rather than treating a single hardware assessment as permanent.

In this article: What Size Model Actually Fits on a Phone · What On-Device Models Are Actually Good For · Battery and Thermal Reality vs. Benchmark Numbers · On-Device vs. Cloud: The Practical Tradeoffs · The Hybrid Pattern Most Apps Are Landing On · Testing an On-Device Model on Real Hardware

What Size Model Actually Fits on a Phone

Two smartphones charging side by side on a desk. Modern and technological setting.

Current flagship phones can run models in the 1-4 billion parameter range at usable speed, thanks to dedicated neural processing units and aggressive quantization that shrinks a model's memory footprint with a modest, often acceptable accuracy tradeoff. Below 1 billion parameters, models run fast but lose enough capability to feel noticeably limited for anything beyond simple classification. See MLCommons: MLCommons' benchmarking work.

Mid-range phones lag flagships meaningfully here — the same 3-billion-parameter model that runs smoothly on a top-tier device can be noticeably sluggish or thermally throttled on hardware two years older, which matters for any app targeting a broad device range rather than just the newest phones.

What On-Device Models Are Actually Good For

Close-up of a hand holding a smartphone with app icons and a nature wallpaper.

On-device models handle summarization, basic classification, simple rewriting, and narrow assistant tasks — anything that leans on pattern-matching over a bounded input rather than broad world knowledge. A keyboard suggesting a reply, a notes app summarizing a long entry, or a camera app tagging a photo's contents are all realistic on-device tasks today. See Stanford HAI's AI Index: Stanford HAI's AI Index report.

They fall short on anything requiring current events, obscure factual knowledge, or multi-step reasoning across a large context — the smaller a model gets, the more its knowledge is compressed and the more gaps show up on anything outside its core training distribution. For more on this, see Model Drop's small models that run on one GPU.

Battery and Thermal Reality vs. Benchmark Numbers

A smartphone charging from a power bank with an orange cable on a wooden table.

A benchmark score measured on a cool, idle device does not predict how a model performs during sustained real-world use, where thermal throttling kicks in after a few minutes of continuous inference and measurably slows response time. This is the gap between a vendor's demo numbers and what a user actually experiences during a long session.

We have found the more useful test is running a model for 10-15 minutes of continuous use, not a single quick benchmark pass, since that is closer to how a real feature — live transcription, for instance — actually gets used. For more on this, see Model Drop's self-hosting versus API inference.

On-Device vs. Cloud: The Practical Tradeoffs

A detailed view of a backlit laptop keyboard with glowing blue lights.
On-Device vs. Cloud: The Practical Tradeoffs
FactorOn-deviceCloud API
PrivacyData never leaves the deviceData sent to a server
Works offlineYesNo
Capability ceilingLimited by phone hardwareMuch higher
Cost per requestNone after app downloadPer-token or per-request
Battery/thermal impactReal, sustained-use dependentMinimal on-device

The Hybrid Pattern Most Apps Are Landing On

Detailed industrial model display showcasing machinery and systems indoors, emphasizing technology and design.

The common production pattern now runs a small model on-device for fast, simple, privacy-sensitive tasks, and falls back to a cloud API for anything the on-device model flags as uncertain or outside its scope. This gets most of the latency and privacy benefit of on-device inference without giving up cloud-level capability for harder requests.

Model Drop's guide to self-hosting versus API inference covers the equivalent tradeoff on the server side, which is worth reading alongside this one if the cloud-fallback portion of a hybrid app is being built from scratch rather than bought from a vendor. For more on this, see Model Drop's mixture-of-experts versus dense models.

Testing an On-Device Model on Real Hardware

Test on the actual range of devices your user base runs, not just the newest flagship phone sitting on a developer's desk. The gap between a two-year-old mid-range device and a current flagship is large enough to change whether a feature is usable at all, not just how fast it feels.

Measure performance after ten to fifteen minutes of continuous use, not on the first request. Thermal throttling is the gap between a benchmark number and what a real user experiences during sustained use, and it rarely shows up in a quick initial test.

Track battery drain over a typical session length alongside speed, since a feature that is fast but noticeably drains battery will get uninstalled just as quickly as one that is simply too slow.

Instrument the feature in production with real device and OS version data from day one, not just app-store aggregate crash reports. On-device AI failure modes correlate strongly with specific chipset and OS version combinations in ways that are hard to predict from lab testing alone.

Consider a staged rollout for any new on-device model, starting with newer flagship devices before expanding to the full device range. This limits the blast radius of any unexpected performance or thermal issue to a smaller, generally more forgiving user segment while the rollout is still being monitored closely.

Conclusion

On-device AI on mobile has crossed a real usability threshold for narrow, bounded tasks — summarization, simple classification, basic assistants — but still needs a cloud fallback for anything requiring broad knowledge or complex reasoning. Model Drop's recommendation is to test any on-device model under sustained real use, not a quick benchmark, and to design the cloud-fallback path from day one rather than bolting it on after launch.

Start by mapping which of your app's AI features are genuinely privacy-sensitive or need offline support — those are the strongest candidates for on-device, everything else can likely stay cloud-first.

One last practical note: involve your mobile engineering team early in any on-device AI evaluation, since hardware and OS-specific quirks are often easier to catch from someone who works with the platform daily than from a model-evaluation process run in isolation from the actual app codebase.

Keep in mind that app store review guidelines occasionally place specific constraints on background AI processing and battery usage, so it is worth checking current platform policies before finalizing an on-device AI feature's technical design, not just after implementation is complete.

Model Drop covers AI launches, models, tools, and platforms for developers and builders tracking the frontier.

What size AI model can run on a modern smartphone?
Flagship phones handle models in roughly the 1-4 billion parameter range at usable speed, thanks to dedicated neural processing hardware and quantization. Mid-range and older devices handle noticeably smaller models well.
Do on-device AI models drain phone battery quickly?
Sustained use does have a real battery and thermal cost, more than a quick benchmark suggests. Short, occasional requests have minimal impact; continuous use like live transcription is where the cost becomes noticeable.
Is on-device AI more private than cloud AI?
Yes, meaningfully — data processed on-device never leaves the phone, which matters for privacy-sensitive features. Cloud AI requires sending data to a server, even when that data is handled securely once there.
Can on-device models work without an internet connection?
Yes, that is one of their main advantages. A properly on-device model runs fully offline, unlike a cloud API which requires connectivity for every request.

Written by

Dana Kwon

Contributing Reviewer

Dana builds and re-runs a fixed evaluation harness against every model and coding tool Model Drop reviews, so a rating from March means the same thing as a rating in August.

Covers

  • model evaluation
  • inference infrastructure
  • cost & performance comparisons