Animated preview of the PocketMind chat: an offline question, a nearby places search using OpenStreetMap,
and a calculator tool call, each with live speed, CPU, memory and temperature readings.
9:41
LFM2.5 1.2B Instruct
13.9tok/s3%1.37GB35.8°C
Explain how vaccines work in two sentences.
LFM2.5 1.2B Instruct
Vaccines show your immune system a harmless piece or copy of a germ, so it learns to recognize it. If the
real germ shows up later, your body fights it off faster and harder.
41 tok · 13.9 tok/s · TTFT 541 ms
Message…
Runs entirely on-device. Responses may be inaccurate.
Powered by open models viallama.cpp
Qwen
MiniCPM
Gemma
Llama
LFM
SmolLM
Phi
DeepSeek-R1
Granite
Qwen
MiniCPM
Gemma
Llama
LFM
SmolLM
Phi
DeepSeek-R1
Granite
Features
An assistant, not a model playground.
Everything you expect from a modern AI chat, plus the one thing cloud assistants can't offer: it works with your phone's data without that data ever leaving it.
Knows your phone
"Am I free at 4?" "Pharmacy near me." "Add Rahul's new number." "Summarise this chat." PocketMind uses
typed tools for contacts, calendar, location and notifications, and you can share any text or chat
to it from other apps.
Contacts
Calendar
Nearby places
Notifications
Share to PocketMind
Web search
Calculator
Update Rahul's number to 98100 12345
Edit contact
Rahul Sharma · mobile → 98100 12345
Save this change?
CancelOK
Every edit, delete or new event asks first.
Private by default
No account, no server, no analytics. What you ask stays on your phone.
Chats & history on device
Memories on device
Model files on device
Voice audio on device
Contacts & calendar on device
Real-time metrics & a benchmark lab
Honest performance: live tokens/s, CPU, RAM and temperature while you chat, plus a benchmark lab to find
the model that suits your phone. Export results as CSV or JSON.
speed
14.1 tok/s
TTFT
552 ms
RAM
1.40 GB
temp
36.4 °C
Works offline
On a flight, in the mountains, or on a patchy connection. Once a model is downloaded, chat, local voice
input and on-phone tools like contacts and calendar work in airplane mode. No per-query cost, ever.
Airplane mode
still answering
Voice input
Speak instead of typing. Use Google's on-device recognizer, or a fully local Moonshine or Whisper model.
4 s of audio → text in 83 ms
Thinking that doesn't overthink
Auto thinking reasons only when it helps (math, code, planning). A budget and loop detection stop small
models from going in circles.
OffAutoAlways
Thought for 21s
Stopped overthinking
Memory you control
Ask it to remember something and it saves a short, editable note on your phone. View, edit or delete any
memory at any time.
Prefers short answers
Vegetarian
Writes Kotlin at work
Gym on weekdays
Your models, your choice
Pick from 20 open-weight models sized for your phone, from 292 MB to 3.1 GB, with a device-fit badge on
each. Paste any GGUF link from Hugging Face. Or connect an OpenAI-compatible server you trust, such as
Ollama on your laptop; API keys are encrypted on the device.
On this phone
MiniCPM5 1B688 MB
LFM2.5 1.2B731 MB
Qwen3.5 2B1.3 GB
Phi-4 mini2.5 GB
Remote · optional
Ollama on your laptop
Any OpenAI-compatible API
Keys encrypted on device
Clearly labelled when messages leave your phone
How it works
Up and running in a minute.
Step 01
Download a model
One tap. PocketMind reads your phone's RAM and recommends models that will run well, with a fit badge on every option.
from 292 MB · resumable download
Step 02
Ask anything
Type or speak. Answers stream token by token with Markdown, code, tables and LaTeX, and your history stays on the phone.
loads in 1.5–3 s · works offline
Step 03
Let it use your phone, with permission
Switch on only the sources you want: contacts, calendar, location, notifications. Every change it proposes waits for your OK.
opt-in per source · OK / Cancel
Privacy
Private by architecture, not by promise.
The model runs on your phone's CPU, so your questions, chats and personal data are processed in your pocket. Data only crosses the line for things that need the internet, and you control each one.
No account, no server
There is nothing to sign up for and no PocketMind backend that sees your chats.
No ads, analytics or trackers
No tracking SDKs in the app. Your data is never the business model.
Only the minimum, under your control
Web search sends just the query and can be switched off. Nearby places sends an area rounded to about 100 m.
Delete anything
Clear chats, memories and models in the app. Uninstalling removes everything.
Query only → DuckDuckGo, when online. Switch it off any time.
Remote model
Opt-in. Only to a server you add, labelled in chat.
Nearby places (OpenStreetMap) and model downloads (Hugging Face) are also user-initiated.
Models & performance
Real numbers from a real mid-range phone.
No cherry-picked flagship. Every figure below was measured in PocketMind's benchmark lab on an iQOO Z10 (Snapdragon 7s Gen 3, 8 GB RAM).
Decode speed, higher is betterPrompt processing (pp512), higher is betterTime to first token, lower is better
MiniCPM5 1B Q4_K_M14.1115552
LFM2.5 1.2B Instruct Q4_K_M14113580
Qwen3.5 0.8B Q8_010.5131422
MiniCPM5 2B Q4_K_M6.7451128
Qwen3.5 2B Q4_K_M3.2n/an/a
4 threads, Q4_K_M / Q8_0 GGUF, measured 4 Oct 2026. *Qwen3.5 2B measured in chat with a 692-token prompt.
Decode speed is limited by memory bandwidth, so smaller files are faster.
1.5–3 s
to load a model
83 ms
to transcribe 4 s of speech (Moonshine)
Need more power?
Point PocketMind at your own GPU. Qwen3.5 4B via Ollama on a GTX 1650 laptop streams at
~38–44 tok/s over Wi-Fi.
20 models in the catalog292 MB – 3.1 GB
Qwen3 / Qwen3.50.6B – 4B
MiniCPM51B, 2B
Gemma 3 / Gemma 4270M – E2B
LFM2.5350M – 2.6B
Llama 3.21B, 3B
SmolLM33B
Phi-4 mini3.8B
DeepSeek-R1 Distill1.5B
Granite 4.01B
Full benchmark table
Benchmark results on iQOO Z10, Snapdragon 7s Gen 3, 8 GB RAM
Model
Quant
File
Generation tok/s
Prompt tok/s
First token
Peak RAM
MiniCPM5 1B
Q4_K_M
688 MB
13.8–14.4
112–118
552 ms
1.40 GB
LFM2.5 1.2B Instruct
Q4_K_M
731 MB
13.5–14.4
101–125
541–622 ms
1.72 GB
Qwen3.5 0.8B
Q8_0
812 MB
10.5
131
422 ms
—
MiniCPM5 2B
Q4_K_M
1.6 GB
6.4–7.0
39–50
1128 ms
—
Qwen3.5 2B
Q4_K_M
1.3 GB
3.2*
—
—
1.2 GB
The app
Built for everyday use.
Real screenshots from PocketMind running on an iQOO Z10.
Chat
A clean start, with live metrics up top.
Code & Markdown
Streaming answers with syntax-highlighted code.
Web search
Searches DuckDuckGo when online and cites sources.
Yes. After you download a model, chat runs entirely on your phone's processor, and so do local voice input and on-phone tools like contacts and calendar. Try it in airplane mode. Only the features that need the internet (web search, nearby places, remote models and downloading models) need a connection.
Which phones does it support?
Android 9 or later on a 64-bit (arm64) phone. We recommend 8 GB of RAM; 6 GB works well with small models (about 1.5 GB or less). The app shows a fit badge on every model so you know what will run well. Our reference device is an iQOO Z10 (Snapdragon 7s Gen 3, 8 GB).
What data does it access?
Only what you switch on in Settings → Permissions: contacts, calendar, location and notifications. PocketMind never asks for your SMS inbox or call history. Data is read on your phone only when you ask a question that needs it, and it is never uploaded. Any edit, delete or new event needs your OK first. See the privacy policy for details.
Is it free?
Yes, early access is free, and there is no account or subscription to sign up for. We may add optional paid extras later, but there will be no ads and your data will never be the business model.
Why isn't it on the Play Store yet?
PocketMind is in early access while we test it on more phones and finish Google Play's review. Want to try it now? Request early access and we'll send you the APK.
Can I use my own server or models?
Yes. Paste any GGUF model link from Hugging Face to run it on-device, or add any OpenAI-compatible endpoint (for example Ollama or LM Studio on your computer, or a hosted API). API keys are encrypted with the Android Keystore, and chats that use a remote model are clearly labelled because those messages go to that server.
An AI that lives on your phone, not in someone's cloud.
PocketMind is in early access on Android. Tell us your phone model and we'll send you a build.