On-device AI runs a model directly on your phone or laptop, using its own chip to work through a request without sending anything over the internet. Cloud AI does the opposite: it sends your request to a company's remote servers, where a far bigger model does the work and sends the answer back. Both get called "AI," and both can look identical on your screen, but they run in almost entirely different ways, and the difference matters for how fast a feature feels, how capable it is, and where your data actually ends up.
The confusion is understandable. Phone makers, model companies, and app reviewers all use "on-device" and "cloud" loosely, sometimes to describe an entire product and sometimes to describe one feature buried three menus deep in settings. The same phone can run one feature entirely locally and another entirely on a server, and the marketing copy for both often uses nearly identical language about being fast, smart, and safe. Knowing the actual mechanics, rather than either side's framing, is what lets you reason about speed, battery drain, and privacy for any specific AI feature you use, instead of taking a company's claim at face value.
How on-device and cloud AI actually work
On-device AI: your phone doing the work
On-device AI depends on the chip already inside the device. Apple's iPhones ship with a Neural Engine, a part of the processor built specifically to run small AI models quickly and using little power, and Apple Intelligence uses it by default whenever a task is simple enough. Apple has published that its on-device language model uses roughly 3 billion parameters, small enough to generate around 30 tokens per second locally on an iPhone 15 Pro. That is a genuinely small model. A leading cloud model runs at a scale no phone chip comes close to, which is the core tradeoff: on-device models trade raw capability for speed and for keeping data local.
When a request needs more reasoning than that small model can handle, such as pulling information together across several long documents, Apple's software routes it to Private Cloud Compute instead, a server system that runs on Apple's own hardware rather than a third party's. Apple states that personal data sent to Private Cloud Compute is not accessible to anyone other than the user, including Apple's own staff, that the data is deleted once the response is returned, and that independent researchers can inspect the actual software running on those servers to check the claim. Apple also says the device cryptographically verifies which server will handle a request before sending anything, so it can confirm that server is running the audited software rather than something else. It is still a cloud system, just one built with unusually strict constraints on what it can retain.
Google takes a similar hybrid approach on Pixel phones. Gemini Nano, Google's smallest model, runs locally through a system called AICore, which Google's own developer documentation describes as a way to deliver generative AI features "without needing a network connection or sending data to the cloud." AICore is built to isolate each request and keep no record of the input or output once it finishes. On the Pixel 10 series, that on-device model, paired with the phone's Tensor G5 chip, is what powers Magic Cue, Voice Translate, and Pro Res Zoom. Google says on-device AI capabilities overall now run "even more smoothly while using less power" on the new chip. A longer, open-ended conversation in the full Gemini app, by contrast, still goes to Google's servers.
Samsung splits things even more visibly, on the same phone. Its own support documentation lists which Galaxy AI features process data on the device, including translation and the AI weather wallpaper, and which ones require the cloud, including auto formatting, note summaries, and generative photo editing. Samsung even ships a toggle, under Settings > Advanced Intelligence, to process data only on the device, and turning it on switches off whichever features need a server to work. That single settings page is a good demonstration that a phone is rarely purely "on-device" or purely "cloud"; it is usually both, one feature at a time.
Cloud AI: a data center doing the work
The tools most people picture when they hear "AI," ChatGPT, Claude, and the main Gemini chat assistant, are cloud AI by default. Typing a question sends it over the internet to a data center, where a model with far more parameters than anything that fits on a phone generates a response and sends it back. Anthropic, for example, states that Claude conversations are encrypted both in transit and while stored at rest, and that employee access to conversation content is restricted rather than routine. That is a real safeguard, and it reduces several kinds of risk, but it does not change the basic fact: your prompt leaves your device and reaches a company's servers, even if only for a moment, before you see a reply.
A quick way to tell which one you're using
- If a feature keeps working in airplane mode or with Wi-Fi off, it is running on-device.
- If the app shows a loading state tied to your connection, or stalls when you lose signal, it is calling a server somewhere.
- Marketing phrases like "runs on your device" or "processed locally" describe on-device features; "powered by" a named large model almost always means cloud.
- A privacy page that names a specific chip, such as a Neural Engine or Tensor processor, is usually describing an on-device feature rather than a cloud one.
None of this makes one category better than the other in general; it makes them suited to different jobs. A quick reply drafted in a messaging app benefits from an on-device model that answers instantly and never needs a connection. A request to summarize a stack of unfamiliar documents benefits from a cloud model with the size and context window to actually do it well. The table below lays out that tradeoff plainly, without the marketing language either side tends to attach to its own approach.
On-device AI vs. cloud AI, side by side
| What matters | On-device AI | Cloud AI |
|---|---|---|
| Speed | Near-instant; no round trip to a server | Adds network latency, typically under a second to several seconds |
| Model size and capability | Small, usually a few billion parameters; strong at narrow, repeatable tasks | Far larger models; better at open-ended reasoning, long documents, and creative work |
| What leaves your device | Nothing, in the simplest cases; the prompt and the answer stay local | Your prompt, and anything attached to it, travels to the provider's servers to be processed |
| Needs an internet connection | No, for features built to run locally | Yes, every time |
| Typical uses today | Live translation, call scam detection, quick replies, voice-memo summaries | Open-ended chat, research, coding help, image generation, long-document analysis |
What this actually means for your privacy
On-device processing removes one real category of risk: your content never has to pass through a company's servers, so it cannot be caught in a breach of that company's cloud infrastructure, read by an employee with server access, or produced in response to a legal request directed at the company, because the company never had it. That is a genuine, structural advantage, and it is the reason phone makers highlight it in their marketing.
But "on-device" does not mean "automatically safe." A phone with no passcode set, or one that is lost or infected with spyware, exposes whatever is stored on it directly, with no server breach required. Most on-device AI features also inherit the security of the rest of the phone: its passcode strength, whether the owner bothered to set a screen lock at all, and how current its software updates are. A weak link anywhere in that chain undoes the advantage. Plenty of the source material behind "on-device" features, the photos, messages, and recordings themselves, still syncs to a cloud backup anyway, for reasons that have nothing to do with the AI feature at all.
Cloud AI carries the opposite reputation problem: because data technically leaves the device, people assume it is inherently less private, but that skips over what actually decides the outcome. Encryption at rest means that if someone gains access to the raw storage without the decryption key, the data is unreadable. Per-user isolation means one customer's data is walled off from another's, so a bug or a compromised account cannot spill into someone else's records. Access logging and multi-party approval mean that even employees inside the company leave a trail when they touch customer data. None of that makes cloud storage untouchable, but it means "cloud" and "insecure" are not the same word, just as "on-device" and "safe" are not the same word.
In practice, it helps to picture two different attackers rather than two abstract categories. One has brief physical access to a phone left on a table. The other is a remote attacker going after a company's servers at scale, or a government request directed at that company. The first threat makes on-device storage risky specifically when a phone has no passcode; the second makes cloud storage risky specifically when a provider's encryption, key management, or access controls are weak. Per-user key isolation and logged access exist to answer that second threat, the same way a passcode answers the first. Neither one, on its own, protects against the other.
The question worth asking is not which category sounds safer in the abstract, but what you are actually protecting against. Worried about a network eavesdropper, or a company mining your prompts for advertising data? On-device wins outright, because there is nothing for anyone to intercept or read. Worried about losing your phone or having it stolen? A cloud account protected by strong authentication and reachable only after re-verifying on a new device can end up safer than data sitting locally on a screen with no passcode. Worried about a breach or a subpoena aimed at the company itself? On-device data was never there to hand over in the first place; cloud data was. The right answer changes depending on which threat you are weighing, which is exactly why neither approach deserves to be treated as the automatic winner.
MemX, the memory app behind this blog, is a useful real example of where these tradeoffs land in practice. It is not an on-device app: photos, documents, voice notes, and messages you save are stored in the cloud so they stay searchable across every device you own, rather than depending on one phone's storage or battery. MemX describes its approach as private by architecture rather than end-to-end encrypted. Each user's data sits behind its own encryption key managed in a hardware security vault, is encrypted at rest and in transit, and is decrypted only briefly, in server memory, to answer a search you asked for, with that access logged. That is a different privacy model from Apple's or Google's on-device features, and MemX is explicit that it is neither end-to-end encrypted nor zero-knowledge; it is a cloud product built to narrow who can read your data and when, not a claim that your data never leaves your device. It is one honest example among several ways to handle the tradeoff described above, not the objectively correct one; a fully on-device memory tool would make different, also-defensible choices, and would give up cross-device search and larger-model reasoning to do it.
01Is on-device AI always more private than cloud AI?
Not automatically. On-device AI keeps data off a company's servers, which removes breach and subpoena risk, but a stolen phone with no passcode exposes local data directly. Cloud AI can be well protected with encryption and access controls. Privacy depends on the specific threat, not just where processing happens.
02Can on-device AI features work without internet?
Yes, that is usually the point. Features like Gemini Nano's on-phone summaries or Apple's basic on-device Siri commands run on the chip itself and need no connection. More demanding requests on the same phone often still route to the cloud.
03Why don't companies just make everything on-device if it's more private?
Phones do not have the memory, storage, or battery budget for the largest models. On-device models run a few billion parameters; cloud models run far more, which is why long, open-ended, or creative tasks still get sent to a data center.
04Does encryption at rest mean a company can't read my data at all?
It means stored data is unreadable without the decryption key. A provider holding both the encrypted data and the key, which is how most cloud AI works, can still decrypt it to serve a request. That differs from end-to-end encryption, where even the provider cannot read it.
05Is MemX on-device or cloud-based?
Cloud-based. MemX stores photos, documents, and notes in encrypted cloud storage so they stay searchable across devices, using per-user encryption keys and logged access. MemX calls this private by architecture, not end-to-end encrypted or zero-knowledge.
