I started thinking about the Web Models API while experimenting with building AI-native GUIs. I was exploring the idea of UI as a function of AI: the model would generate the content and choose the component best suited to display it.
One experiment was a news site which used LLMs to identify key topics (people, places, or things) in an article and turn those topics into links. A reader could click on a topic to get more context, presented in a component suited to that content.
While working on demos I kept running into the same problem. Running an LLM in the browser with WebAssembly or WebGPU meant having the app load and run the model itself, and the performance wasn’t good enough for what I was building. Using a cloud provider meant paying for usage, even when I just wanted to share a demo.
That got me thinking about what this means for indie developers. Putting an experiment or small product online can mean an ongoing bill before you know whether anyone finds it useful or have a way to fund it.
I wanted web apps to use on-device models with the user’s permission, much like asking to use a camera or microphone. This led me to build my first attempt at solving this. Browser.AI was a prototype that exposed an API on the window object to access models running on-device. The learnings from that became the foundation for the Web Models API proposal.
How it works
The Web Models API exposes access to open-weight on-device models through a simple verb oriented API on navigator.models. The app requests an open-weight model using a universal ID so developers can target the exact model release they built and tested against.
The user decides if a site can access that model, and at any point can revoke access. The permission is tied to the model and site pairing, so permission to use one model does not give the site access to every model on the device.
The browser sets up and runs models, manages resource use, and shares model weights across sites. Each app doesn’t need their own copy of the model, but the data and permissions stay separate.
The API also supports tool calls, structured output, and multimedia. An app could provide functions for filtering a list, choosing a component, or updating a form; the model requests a function call, and the app checks it before running it. Structured output means the app gets back JSON it can render directly, not text it has to parse. Multimedia input means a model can take an image directly, useful for something like a receipt reader.
So for my news site experiment, the app would still decide what context to send, check the model’s response, and render the selected component. The browser handles running the model.
Cost and control
When a suitable model can run on the user’s device, you don’t need a paid cloud inference service. For me, this is also about control: over where our money goes, and over where our users’ data goes. Developers should be able to decide when paying a cloud provider adds value and when a local model is enough.
In many cases there’s no good reason to send that data to a server when the same task can run securely on the user’s machine. Someone filling out a tax form or discussing a medical scan shouldn’t have to trust that data to a cloud provider when the device they’re using can do the job.
There are also reasons to run locally beyond cost and privacy. A notes app could search personal documents without sending their contents to a model provider, and with the model, app, and documents available locally, that search works offline. Local processing also skips the network round trip.
Why model choice matters
I believe competitive open-weight models are an important part of that future. Different tasks need different models. A small model may be enough to choose a UI component. Document search needs an embedding model. A receipt reader needs a model that can handle images. Developers should be able to choose based on what works for their app, independently of which browser someone uses.
The Prompt API provides access to a browser specified language model. But apps often end up tuning their prompts to one model’s quirks. Those prompts may work poorly in another browser or after a model update.
In my UI experiments, the model’s ability to identify content and select the right component was part of how the app worked. A different model could make different choices. You’ll run into the same issue if you want to use embeddings for document search: an app must use the same embedding model to index documents and later search them. Developers need a defined model target they can test, and a way to choose when to move to a new release.
Being forced to use the model a big tech company shipped with their browser makes us developers dependent on their internal business strategies. For example, if they decide that on-device inference is a challenge to their cloud business, they might be less motivated to improve the model. I want developers and users to have real alternatives, with less dependence on a few companies’ pricing, model choices, and decisions about how their services work.
Help shape it
I care about this because I want to keep building and sharing useful and interesting applications of AI, and I want other developers to be able to do the same. Funding an expensive cloud inference bill and sending your users’ data into a black box shouldn’t be requirements to access AI, especially when the task can run on hardware the user already has.
The Web Models API is still early, and I’d like to hear from developers.
If you could run a specific open-weight model in any browser, for free, what would you build? Which model would you use? If you’ve already shipped AI features with Transformers.js, the Prompt API, or a cloud provider, what got in your way?
Let me know in the comments. The explainer and API reference are on Github if you want the details. In the next post, I’ll walk through the API with code.