Skip to content

A 0.6B Model on a 2017 Phone Just Drove a Desktop Browser

#on-device-llm #llama-cpp #ollama #open-weights #edge-inference #small-models

A 0.6B Model on a 2017 Phone Just Drove a Desktop Browser ​

A 0.6B parameter model, quantized with Q4_K_M, running inside Termux on a 2017 Samsung Galaxy Note 8, drove a real desktop Chrome browser. Ten runs per task, every result checked against a fixed expected value, logs and replay scripts published in the repo. The test author is upfront about the caveats: three fixed tasks, not a benchmark. Even so, the results change what's credible for sub-1B models.

The Note 8 test that makes small models look competent ​

The setup is deliberately minimal. The phone runs llama.cpp in Termux with Qwen3-0.6B at Q4_K_M, a model file of about 400 MB. A separate laptop runs a Chrome session, not headless. The phone drives that browser through a relay built by the testing team.

The model's job is narrow. It receives a structured representation of the page, about 10 named links or fields, roughly 200 tokens. It picks one by name, then copies the facts it was given into JSON. Everything else, capturing the page as structure, candidate selection, the click, reading the facts, verifying the result, is done by the stack around it. The model never sees HTML, a screenshot, or a URL.

Three tasks:

  1. A sandbox bookstore: find a category, find a book, extract price, rating, and stock.
  2. Live Wikipedia: from an unrelated site, navigate to the Galaxy Note series page, pick "Note 8" among decoys called "Note 8.0", "Samsung Galaxy Note 8.0", "Galaxy Note 8.0", and "Note FE", on a page with roughly 760 interactive nodes. Return the release date from the infobox.
  3. Extract five fields, including a UPC, from a table.

Each task ran 10 times against a fixed expected value. Qwen3-0.6B scored 10/10 on all three.

Key numbers: 0.6B parameters at Q4_K_M fits a 6 GB phone from 2017 with no GPU, and the whole file is about 400 MB. The model's entire view of a page is about 200 tokens. It hit 10/10 on every task. The same work took 500 tokens and 80 seconds with structured input, versus 12,000 tokens and 22 minutes with raw HTML.

Twelve models, one task: size isn't the story ​

The test runner put 12 small models through the same sandbox task, same script, same prompt. The spread is wide.

ModelParamsTask 1 (out of 10)Note
Qwen3-0.6B0.6B10clean run
Qwen2.5-1.5B1.5B10clean run
GLM-Edge-1.5B1.5B10rating as digit
Gemma-2-2B2.6B10clean run
Llama-3.2-3B3B10clean run
MiniCPM5-2B2B9"£" written as "$" once
Qwen2.5-0.5B0.5B6inconsistent
LFM2.5-1.2B1.2B0placeholder output
Llama-3.2-1B1B0pseudo-code instead of answer
Gemma-3-1B1B0placeholder output
LFM2-350M0.35B0random click
Gemma-3-270M0.27B0placeholder output

Every model at 1.5B parameters or above nails it. Below 1B, only Qwen3-0.6B works at all, and the rest collapse in instructive ways. LFM2.5-1.2B doesn't output a selection, it writes a placeholder. Llama-3.2-1B answers with pseudo-code. The two smallest Gemma models emit placeholders, and LFM2-350M clicks randomly.

The pattern matters: at this scale, training data and architecture dominate parameter count. Qwen3-0.6B beats models two to three times its size on identical hardware. MiniCPM5-2B, a brand-new 2B release, made exactly one mistake in ten runs, a currency symbol slip.

If you're choosing a model for edge deployment, reproduce this kind of result instead of reading benchmark cards. The repo publishes every JSONL log, model hash, and environment detail, and replay.py rebuilds the prompts from the logs and runs them on any OpenAI-compatible local server. Checking whether your model picks "Note 8" among the decoys takes about ten minutes.

Structured input beats raw HTML, every time ​

The control test is the most useful part of the whole exercise. Same pipeline, same model, but raw HTML instead of structured browser perception. On the sandbox task, the model gets there 4 times out of 5, at 12,000 tokens and 22 minutes, instead of about 500 tokens and 80 seconds. On Wikipedia, the page HTML is 467,000 characters. Only 9% fits into a 16k context, and the model can't find the link in that 9%. Zero for three.

The loop is simple once you see it.

The approach has roots in prior work. AgentOccam showed the same effect for GPT-4-class models; WebLINX and MindAct for fine-tuned small models. What changed is the extreme end: no fine-tuning, below 1B parameters, 2017 phone hardware. If you're building on-device agents, the input representation is the lever, not the model card.

Quick Take: A 0.6B model given clean structure outperforms 2B models fed raw documents. The bottleneck is input design, not parameter count.

The agent gap no benchmark measures ​

Ask people running local agentic models what frustrates them, and a different failure mode surfaces. A model that continues on its own too eagerly. If a requirement is ambiguous, I'd rather it stop and ask "do you mean A or B?" than spend ten minutes reasoning, making an assumption, calling tools, and confidently building the wrong thing.

I've run agentic coding models locally for months, and the pattern is consistent. The model doesn't lack capability. It lacks the trigger to stop. It treats missing information as something to guess past.

Nothing in the standard benchmark suite measures this. We test coding, reasoning, tool use, context length. Nobody tests whether a model knows when it doesn't have enough information to continue.

The Note 8 test runs into the same wall. The tasks are name matching and copying. Where judgment about the page is required, the 1.5B class breaks: it can't pick "next" among topical decoys. Pagination wasn't tested because the small models simply couldn't handle it.

The thread asking which model knows when to stop has a common theme in the replies: the answer is mostly the agent framework, not the model. Cap tool calls, detect low-confidence states, and build an explicit escalation path that returns to the user with a question. Small models will reliably say "I don't know" if you give them a sanctioned way to do it.

What's shipping in the small-model corridor ​

Three releases frame where this space is heading.

MiniCPM5-2B landed on Hugging Face as a 2B generalist. In the Note 8 test it scored 9/10 with a single currency slip. That's a usable result for a model that runs on a phone.

Ling-3.0-flash-VL takes the opposite path: a 124B total model with only 5.5B parameters activated per token, which keeps inference cost near a 5.5B dense model while holding 124B of capacity. It also brings a 1M-token context along with native image and video understanding, so an entire codebase or a long agent session fits without chunking. A ViT encoder extracts visual features, a two-layer MLP projector aligns them with text, and VideoRoPE encodes both spatial position and temporal order. The 42-layer hybrid backbone alternates KDA and Gated MLA layers at a 5:1 ratio, which is what lets it process long histories across text, images, and video.

DeepSeek is running the same cost play on the API side. Flash 4.1, an intermediate version, is in internal beta with native multimodal support, faster speeds, and pricing identical to deepseek-v4-flash. You call it by setting the model name to deepseek-v4.1-flash-expires-on-0910, with 20 concurrent requests per account. The pattern across the board: flash-tier models are becoming the default for cost-constrained deployment, and the line between "phone model" and "API model" is blurring fast.

One community counterpoint is worth hearing. A running wish on r/LocalLLaMA is that Google keeps Gemma chat-first and doesn't turn it into another benchmark-chasing Qwen clone. The fear is that "benchmark maxxing" flattens the personality that made models like Gemma fun to use. It's a reminder that capability measures and taste are different axes.

Buying hardware: GB per dollar, bandwidth per dollar ​

The GPU shopping conversation on the local LLM subs keeps circling back to two specs. VRAM sets the ceiling on what you can load. Memory bandwidth sets the speed at which tokens come out. Compute cores barely matter for llama.cpp-style decoding.

Community guides plotting GB per dollar and bandwidth per dollar make the point visually. A card that is "on paper" faster than a 3090 can lose badly on price per gigabyte, and bandwidth, not FLOPs, is what predicts tokens per second. Secondhand server cards keep showing up in these comparisons because they win the GB per dollar line, and that's the line that matters.

If you're buying, decide what you're actually running first. A 0.6B model doesn't need a GPU at all, a laptop CPU handles it. A 7B model at Q4 lands around 4-5 GB of VRAM, so an 8 GB card works. MoE models with small active parameter counts, like Ling at 5.5B active, are starting to make the same capacity available on far more modest budgets.

The llama.cpp and Ollama split ​

The most heated argument in the local LLM world right now is tooling. Credit and control are the stakes.

The background: llama.cpp is the C++ inference engine that made local LLMs possible on consumer hardware. Ollama is the wrapper that made it accessible, and for that it deserves credit. But the history is messy. For over a year, Ollama's README and binaries shipped without the MIT license notice that llama.cpp's license requires, and the GitHub issue asking for it sat for over 400 days without a maintainer response. A community PR adding a single credit line was eventually merged, with the team noting they planned to move to "more systematically built engines."

Then the divorce. In mid-2025, Ollama built its own inference backend directly on the lower-level ggml library, citing stability needs. The result, by community account, was a series of regressions: broken structured output, vision model failures, GGML assertion crashes, and models like GPT-OSS 20B that wouldn't run because of missing tensor type support.

Community measurements on identical hardware tell the same story: around 161 tokens per second for llama.cpp versus 89 for Ollama. I've seen the same gap on my own machine, roughly 1.8x. On CPU-only runs it's 30-50%. The overhead comes from the daemon layer, the offloading heuristics, and a vendored backend that trails upstream.

The model naming problem is harder to forgive. When DeepSeek released R1, Ollama's library listed the distilled Qwen and Llama variants as "DeepSeek-R1". Running ollama run deepseek-r1 pulls an 8B Qwen distillate, not the 671B model. It took me a while to realize why my "R1" behaved so differently from everyone's screenshots. DeepSeek named these models "R1-Distill" for a reason, and Hugging Face lists them correctly.

The workflow complaints are where Ollama's design shows its age. GGUF files embed everything needed to run a model: chat template, stop tokens, metadata. Ollama's Modelfile re-implements this in a separate Go-template syntax, and if the GGUF's embedded Jinja template isn't in Ollama's hardcoded list, it silently falls back to a bare prompt template and breaks the model's instruction format. Change one sampling parameter and Ollama can end up copying the entire model file, which people report at 30-60 GB for bigger models.

Then there's availability. When a new model drops, community quantizations land on Hugging Face within hours, and llama.cpp runs them directly: llama-server -hf unsloth/whatever:Q4_K_M. Ollama users wait for someone to package it into their registry, with a limited quantization menu that typically stops at Q4_K_M and Q8_0. No Q5, Q6, or IQ quants.

The other side deserves a fair hearing. Ollama built the on-ramp. Thousands of people ran their first local model because of it, and for casual use it still works. But the costs show up in throughput, quant choice, model availability, and recent moves like the closed-source desktop app and cloud routing that sends prompts to third-party providers. When a tool that built its brand on local privacy ships CVE-2025-51471, a token exfiltration vulnerability where a malicious registry server can steal your auth token during a model pull, it's fair to ask where the priorities sit.

Across the subs, the working advice has converged: if you want to test new models, use llama.cpp, transformers, vLLM, or SGLang directly. If you want a GUI, LM Studio accepts Jinja templates and reads GGUF metadata without translation. Ollama remains the easiest start. It's no longer the best default.

Common pitfalls ​

  1. Picking a model by parameter count. Qwen3-0.6B outworked models twice its size in the Note 8 test. At this scale, training data and architecture decide. Run the actual task before you commit.

  2. Feeding raw documents to small models. 467,000 characters of HTML collapsed a 16k context, and the model found nothing. Structure, summarize, or chunk first. The token budget is the constraint.

  3. Buying hardware on compute instead of bandwidth. Local decoding is memory-bound. VRAM sets what you can load, memory bandwidth sets how fast tokens come out. Compute cores are a minor factor.

  4. Trusting wrapper model names. ollama run deepseek-r1 gives you an 8B distill, not the 671B model. Verify what you're actually running: check the GGUF metadata, the quant, and the template.

  5. Expecting judgment from small models without an escape hatch. Name matching works at 0.6B. Choosing "next" among decoys broke the 1.5B class. Build a "stop and ask" path into your agent framework, because the model won't invent one.

One thing to remember: the small-model world has passed the point where parameter count decides what's possible. A 2017 phone running a 0.6B model completed real web tasks with perfect reliability because the input representation was designed for the model. The releases landing on Hugging Face right now, from 2B generalists to 124B multimodal MoE models, are all pushing the same direction: more capability per token, per watt, per dollar.

The bottom line ​

If you're building an on-device agent, adopt the structured-perception pattern from the Note 8 test. A 0.6B model with a clean 200-token input beats a 2B model fed raw documents on cost, speed, and reliability, and that gap only widens as tasks get harder