AI That Runs in Your Browser, With No Server and No API Key
Published October 8, 2026
A 360 million parameter model can now run inside a browser tab. What made this possible, and where the limit sits.

In early October 2026, MicroLLM Lab and a handful of similar playgrounds could run language models sized between 26 million and 360 million parameters entirely inside a browser tab, with no server, no API key, and no local CUDA install.
Status note: unverified. The MicroLLM Lab claim and the playground figures here come from company-published material that has not been retested by a third party. The download sizes below are our own estimates rather than official numbers. Treat them as a direction, not as exact figures. The models were quantised to 4 bits and executed through WebGPU, the browser API that exposes GPU compute to JavaScript.
That is worth understanding properly, because the interesting question is not "can it run" but "what does it run, and what does that limit". Both answers are decided by three mechanisms that had to arrive together.
Three things have to be true at once
A small model alone is not enough. WebGPU alone is not enough. What makes it work is the three meeting:
| Layer | Its job | What breaks without it |
|---|---|---|
| 4-bit quantisation | Weights 16 bits down to 4 | 135M parameters needs 270 MB, not 54 MB |
| WebGPU | Direct GPU access, no plugin | Runs on CPU at 1-2 tokens per second |
| Browser cache | Download happens once | Every reload pulls hundreds of MB again |
The middle row is the one people forget. WebGPU is the API that unifies several GPU interfaces under one JavaScript API: Metal on Apple, DirectX 12 on Windows, Vulkan on Linux. The same code runs on all three without changes.
Why 4-bit does not damage models as much as it should
The intuitive argument says a quarter of the precision should cost a quarter of the model. In practice it does not, and the reason is how weights are actually used.
Inside a transformer layer, every token goes through matrix multiplication. The weights in those matrices are almost all small values clustered near zero. Storing most of them as 4-bit integers with a per-group scale factor produces a very small error, because the values being rounded were already close to zero.
What genuinely degrades is the ability that depends on fine representation, which is exactly what small models do badly. Small models have no redundancy; every weight carries real information. Large models have many interchangeable paths, and rounding on one is often covered by another.
That shows up in the numbers people quote: the 135 million parameter model is described as "often incoherent", while 360 million is already usable for simple work. The limit is not the memory footprint. It is how many parameters the model genuinely needs to store knowledge.
The limit you cannot engineer around: coherence
Documentation for in-browser runtimes carries the same warning, and it is routinely skipped: models under 2 billion parameters hallucinate heavily, and their output should be read as a draft rather than a reference.
The reason is that a small model cannot store the world. Its parameters do hold language patterns, compressed well enough to write a plausible sentence, but not enough to hold retrievable facts. A plausible wrong answer is not a bug. It is the most likely result when you squeeze more information through less capacity.
The uses that do make sense for browser models are therefore narrow tasks with checkable answers:
- Sorting text into categories
- Pulling names, dates, and numbers out of text
- Rewriting tone and length
- Translating technical terms against a list you supply
- Running retrieval, where the small model writes the search query rather than composing the answer
Those tasks share one property: the output can be verified without a larger model.
What this means for readers in Indonesia
This is not a small thing for a tools site, because the way it works is the way a browser tool already works: the work happens on the device, and no file leaves the machine.
Three practical implications. First, sensitive data can be handled without a privacy review, because it never leaves the device. Second, the marginal cost is zero, because there is no per-token charge when no tokens are sent anywhere. Third, the hardware limit is real: most phones have 4 GB of RAM, and a 360 million parameter model already takes a meaningful slice of that.
If you are building a tool that currently calls a paid API, there is an intermediate step rarely tried. Move the part that runs most often -- classification, extraction, short summarisation -- onto a small local model, and keep the API call for the cases that genuinely need broad knowledge.
Three things to check before expecting much:
- WebGPU support. Chrome and Edge 113+, Safari 18+, and Firefox still behind a flag. Without it, the fallback to WebAssembly is dramatically slower.
- Memory, not only the GPU. Integrated machines run out of memory before they run out of compute.
- Output quality in your own language. Every measurement above comes from English testing.
Four numbers make it easier to judge. Fifteen to ninety seconds is the first download of a model on a middling connection. After that, subsequent loads take seconds because the weights are already in the browser cache.
Fifty to eighty-four megabytes is the memory footprint for a 100 million parameter model at 4-bit quantisation with group-32 packing. That is not marketing copy, it is the file size after packing.
One hundred fifteen tokens per second was recorded for a 124.6 million parameter model on an Apple M4 through Safari and Metal. The 135 million parameter model on the same hardware reached 66 tokens per second. That difference is not chance: a larger architecture needs more operations per token.
On the other side, one frontier model on a server usually answers at 30 to 60 tokens per second. That means a small browser model is not always slower than a large one in a data centre. For models of 100 to 400 million parameters, on a machine with a GPU, it can be faster, because there is no round trip.
That is the point rarely mentioned: much of the disadvantage of a local model comes from the network and from queueing, not from computation. Removing the trip removes most of the disadvantage.
A short history: from plugins to a standard API
Running AI in a browser has a long history full of failures. The first was WebGL, an API that had existed for decades and worked everywhere, but could only compute over shaders. Shaders were too limited for attention: one attention layer needs to hold matrices for the whole context, and WebGL offers no way to read and write those buffers at the speed a transformer requires.
The next attempt used WebAssembly plus a separate worker. WebAssembly compiles to native machine instructions, so it is far faster than JavaScript. But WebAssembly has no GPU access. The model did run, on the CPU, which was too slow for anything interactive.
WebGPU closes both gaps at once. A browser can now write compute shaders and, more importantly, has explicit control over memory. Separate shaders can do the attention while the context buffer stays in video memory. That is what lets a 1.5 billion parameter model run on a laptop with no server at all.
The change from WebGL to WebGPU is not only speed, it is paradigm. WebGL was treated as a graphics application that happens to compute. WebGPU was designed as a general compute API with explicit GPU access, closer to CUDA opened up to browsers.
Analysis: will this replace APIs
Not for everything, and the reason is structural rather than technical.
| Factor | Local model in browser | Model over API |
|---|---|---|
| Stored knowledge | hundreds of MB | the whole scale of the internet |
| Cost per token | zero | not zero |
| Privacy of input | full | depends on the provider |
| Older devices | cannot run | always works |
| Quality on broad answers | low | high |
The dividing line is clear. A local model wins where privacy and marginal cost decide; an API model wins where broad knowledge decides. Classification, extraction and short summaries sit on the first side. Long document analysis and creative writing sit on the second.
What has genuinely changed in the last few years is not the models but the handoff. Because local models became cheap, the pattern of retrieving locally and generating remotely became sensible: search and filtering run locally, and only the final step calls a large model.
What to watch
Three things. First, whether WebGPU spreads to browsers that lack it. While Firefox sits behind a flag, a lot of readers here on Firefox are left out.
Second, whether 4-bit quantisation holds or gets replaced by 3-bit and 2-bit. Every further reduction in precision saves less memory and costs more quality, and there is a point where the trade stops working.
Third, whether browser memory is enough. The limit right now is not compute but available RAM. A 3 billion parameter model needs roughly 2 GB of download, which is already past what many devices accept without the page feeling slow.
Related tools
Free browser tools that apply to this topic.
- AI Writing TellsFind generic patterns in your own writing. Not an AI detector.
- AI Text HumanizerRewrite text into more natural, readable language while preserving its meaning.
- AI Text SummarizerCreate an abstractive summary locally in your browser with AI.
- Android API Level TimelineLook up the API level, release date, and security patch status of any Android version.
Share this article
Share to
Related articles

October 9, 2026
Why Non-English Costs More: What a Tokenizer Actually Decides
One tokenizer can turn the same word into 5 tokens or 15. That difference decides your API bill and how much context you actually have.

October 8, 2026
One Model or Many: What Multimodal Architecture Actually Changes
Unified models run text, image and audio through one network. Pipelines chain separate models. The difference is not only technical.

October 8, 2026
Why AI Benchmark Scores Are So Often Misleading
The numbers carrying a vendor's biggest claims are often measured on tests users never run. Four reasons, and how to read them.



