Skip to content
Google Says Gemma 4 Fits on a Phone in 1.1 to 2.5 GB of Memory, but the Independent Benchmark Launched in August Published Phone Timings for Just One iPhone
Artificial Intelligence

Google Says Gemma 4 Fits on a Phone in 1.1 to 2.5 GB of Memory, but the Independent Benchmark Launched in August Published Phone Timings for Just One iPhone

Illustration by tuput

English

Apple's on-device model has about 3 billion parameters and a 4,096-token window. Google lists 1.1 GB and 2.5 GB of phone memory for Gemma 4 E2B and E4B. Nearly every accuracy score was run by the company that built the model, and the independent phone benchmark launched in August published timings for a single iPhone.

· · 7 min read

Google’s Gemma 4 model card lists 2.3 billion effective parameters for its smaller phone model, E2B, and Google’s size table gives a mobile memory figure of 1.1 GB for it and 2.5 GB for the larger E4B. Apple’s on-device model is a dense 3-billion-parameter network that works inside a 4,096-token window per session. These are the vendors’ own numbers as of 8 October 2026, and most of the scores attached to them were also run by the vendors.

A parameter is one learned number inside a model, and a token is a word fragment of roughly three to four English characters, according to Apple.

What Apple publishes

Apple announced its third-generation foundation models on 8 June 2026. Two of them run on the phone. AFM 3 Core is a dense model of 3 billion parameters. AFM 3 Core Advanced holds 20 billion parameters in flash storage, activates 1 to 4 billion per request, and pulls the chosen parts into working memory. Apple says the larger one is meant for its most capable chips and names no device. It gives no memory figure for either model.

The limit developers hit first is the window. Apple’s technical note TN3193, updated on 31 March 2026, says the on-device model has a context window of 4,096 tokens per session. Instructions, prompts, tool definitions and the model’s own replies all count against it. English text runs about three to four characters per token, and Chinese, Japanese or Korean about one.

What Google publishes

The Gemma 4 card gives 2.3 billion effective parameters and 5.1 billion with embeddings for E2B, and 4.5 billion and 8 billion for E4B. Effective means the count the model actively uses, which is lower because large lookup tables sit outside it. Both sizes support a 128,000-token context.

The memory table lists 2.9 GB and 4.5 GB for 4-bit versions, and 1.1 GB and 2.5 GB under a mobile row for its LiteRT-LM runtime. Text-only mobile use is 0.84 GB and 2.2 GB. The table covers loading the weights and, according to its notes, leaves out the KV cache, the working memory that grows with the conversation.

Gemini Nano, the version Android phones run through the AICore system service, is built from this family. Google’s 2 April 2026 post says Gemma 4 is the base for Gemini Nano 4, which it calls up to 4 times faster and up to 60% lighter on battery than the previous version. The post states no test conditions for either claim. The ML Kit page lists Nano 4 on the Pixel 11 series and three Galaxy Z Flip8 and Fold8 models, while Pixel 9 and 10 phones appear under version 3. That page also warns that different Nano versions can answer the same prompt differently.

Meta and Microsoft

Meta’s Llama 3.2 1B card lists 1.23 billion parameters, released on 25 September 2024. On a OnePlus 12 using Meta’s ExecuTorch runtime, the quantized version takes 1,083 MB on disk and 1,921 MB of working memory and decodes 50.2 tokens a second. The full-precision version takes 2,358 MB and 3,185 MB and manages 19.2. Meta ran those tests. tuput found no newer small Llama model from Meta.

Microsoft’s Phi-4-mini-instruct has 3.8 billion parameters and a 128,000-token window, with a February 2025 release. Its card reports 67.3 on MMLU, a general-knowledge exam, and 88.6 on GSM8K, a set of school maths problems. The card gives no phone memory or speed figure.

The chips, and claims no one else has checked

MediaTek announced the Dimensity 9600 Pro on 15 September 2026. The release says the NPU 1090 (a neural processing unit, the part of the chip built for AI) has 51% higher prefill performance, meaning faster reading of a prompt, and 55% more tokens per watt. It also says the chip supports models up to 30 billion parameters. The only footnote says the metrics came from “demo device testing in MediaTek labs”, with no baseline chip or model named. The first phones are due this quarter.

Qualcomm launched the Snapdragon 8 Elite Extreme Gen 6 in September 2026. HotHardware’s report of 22 September gives Qualcomm’s figures as an NPU up to 35% faster, a shared memory 50% larger than the previous generation and up to 50% higher INT4 prefill. Qualcomm’s example is a mixture-of-experts model of more than 30 billion parameters with about 3 billion active for each token, which HotHardware cautions is not the same as running a dense 30-billion model. It names the Motorola Signature 27 as one of the first phones.

The independent test

Artificial Analysis, working with Liquid AI, launched phone benchmarks on 24 August 2026. It scored 41 quantized builds and ran 33 on an iPhone 17 Pro, and the launch article publishes phone timings for that one device. Builds had to fit in 8 GB at 4-bit or smaller precision and ran on the llama.cpp engine.

Two 2.6 to 3 billion parameter models tied at 63 on its five-test average. LFM2.5-2.6B took 8.0 seconds and 2.3 GB, and Nanbeige4.2-3B took 21.4 seconds and 4.0 GB. The two 9-billion models took over 25 seconds and 6.9 GB. Gemma 4 E4B is among the models it names. Apple’s model and Gemini Nano are not.

On MMLU Pro, a harder version of that general-knowledge exam, Google’s Gemma 4 card reports 69.4% for E4B and 60.0% for E2B, and says only that the models were evaluated against many datasets. Apple’s third-generation post reports no standard benchmarks. Its AFM 3 Core was preferred on 45.6% of prompts against 23.3% for the 2025 model in Apple’s own side-by-side human grading, and Apple said a technical report would follow. tuput did not find one by 8 October.

What the models cannot do

Apple’s WWDC25 session says to avoid using the on-device model for math or code and not to rely on it for facts. It will not know events after its training date, and it can invent answers. Phi-4-mini’s size limits the facts it can store, Microsoft’s card says, and the model is weaker outside English. Gemma 4 can state incorrect or outdated facts, according to Google’s card. Android’s Gemini Nano page lists token limits, hardware limits and “model capability boundaries” without giving the limits. Artificial Analysis left out agentic benchmarks because small models score close to zero on them.

Indian languages on the device

Apple’s support page, published 14 September 2026, lists 16 language entries for Apple Intelligence, among them English, Japanese and Turkish, and names no Indian language. The ML Kit summarization API on Android supports English, Japanese and Korean. Gemma 4 is pretrained on over 140 languages, with out-of-the-box support for more than 35, though the card does not list them. Meta lists Hindi among eight officially supported languages for Llama 3.2 1B, and its card shows the base model scoring 33.5 on multilingual MMLU in Hindi against 40.5 in French.

The Indian-built option is Sarvam Edge, announced on 14 February 2026. Sarvam’s post lists a 74-million-parameter speech recogniser of about 294 MB, a 24-million-parameter speech synthesiser of about 60 MB and a translation model of about 150 million parameters and 334 MB. Speech covers 10 Indian languages and translation those 10 plus English, giving 110 pairs. Speech recognition starts in under 300 milliseconds on a Snapdragon 8 Gen 3, Sarvam reports, though the post publishes no word error rates for its comparison with Google Cloud. The product page says 22 languages, which does not match the post’s 10. Sarvam’s scores for the cloud versions are covered in speech and translation models, and the wider push is in India’s work on AI in its own languages.

Intel India and the government’s Bhashini division announced BHASHINI Vidyalekha on 2 March 2026, an offline tool that transcribes English lectures and translates them for students. It runs on Intel Core Ultra laptops, not phones, and the announcement does not name its languages. Sarvam’s product page says Edge is a commercial offering, available by contacting the company, and not open source.

Share
Copied!

Sources & further reading

  1. Apple Machine Learning Research: Introducing the Third Generation of Apple's Foundation Models (8 June 2026)
  2. Apple Developer: TN3193, Managing the on-device foundation model's context window (updated 31 March 2026)
  3. Apple Developer: Explore prompt design and safety for on-device foundation models (WWDC25, session 248)
  4. Apple Support: Apple Intelligence supported languages and iPhone models (published 14 September 2026)
  5. Google AI for Developers: Gemma 4 model card
  6. Google AI for Developers: Gemma 4 model overview and memory requirements
  7. Android Developers Blog: Gemma 4, the new standard for local agentic intelligence on Android (2 April 2026)
  8. Google for Developers: Overview of the ML Kit GenAI APIs (updated 7 October 2026)
  9. Google for Developers: ML Kit GenAI Summarization API (updated 7 October 2026)
  10. Android Developers: Gemini Nano and AICore (updated 8 October 2026)
  11. Hugging Face: Meta Llama-3.2-1B model card
  12. Hugging Face: Microsoft Phi-4-mini-instruct model card
  13. MediaTek: Dimensity 9600 Pro press release (15 September 2026)
  14. HotHardware: Snapdragon 8 Elite Extreme Gen 6 release coverage (22 September 2026)
  15. Artificial Analysis: Mobile phone intelligence and inference benchmarks (24 August 2026)
  16. Sarvam AI: Announcing Sarvam Edge (14 February 2026)
  17. Sarvam AI: Sarvam Edge product page
  18. IANS: Intel and Digital India BHASHINI bring offline multilingual capabilities to AI PCs (2 March 2026)

Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.

#on-device ai#small language models#gemma 4#gemini nano#apple foundation models#sarvam edge#npu#indian languages

Enjoyed this? Get the next one.

One good read at a time, straight to your inbox. No spam, unsubscribe anytime.

More in Artificial Intelligence
Scale's SWE-Bench Pro Board Puts the Best Coding Agent at 61.5%, Maintainers Reject Many Fixes That Pass the Tests, and METR Says Its Latest Trial Cannot Show How Much Time Agents Save
Benchmarks say agents are strong. Maintainers, a randomised trial and a pile of incident reports say the picture is narrower. Here is what each number actually measures.
Two Supreme Court Rulings in 2026 Set Aside Orders Built on Fake AI Citations, While Indian Courts Use AI for Translation and Transcription and the Court's Draft AI Rules Are Still Not Final
Parliamentary replies and court records show which AI tools are live, which are still pilots, and what judges are barred from doing with them. Three Supreme Court orders show what happens when invented case law gets through.
Sarvam Raised $234 Million at a $1.5 Billion Value and Released Two Open Models, but the Government's August List of Launched Models Names Neither Soket AI nor Gan.ai, Which Share Rs 287 Crore of Approved Support
Two of the four startups the government picked in 2025 have put models in public view. The rest of the field is a mix of previews, plans and scores that the builders ran themselves.
← all articles