Private AI · a new tab for Chrome

A private AI in every tab.
No account. No key.

Type a question in Chrome’s new tab and a Qwen3.5 model answers from inside your browser, on your GPU through WebGPU or on your CPU. It downloads once and then it is a file on your disk. Your questions go nowhere. Turn off the Wi-Fi and it keeps answering.

Free. Open source. One model download, 533 MB to 2.7 GB, cached for good.

chrome://newtab
Ready · Qwen3.5 4B · WebGPUFollow-ups stay in this conversation.
your question stayed heremodel cached on this devicekey none
What happens to a question

The model is a file. The answer is made where you are.

Every step, and where it happens. In local mode none of them involves a server, because there is none.

You doFogar doesIt happens
Open a tabNothing, yet. A new tab does not load the model on its own.Opening tabs to type an address stays free. A tab nobody has looked at for five minutes gives the memory back.This tab
AskWakes the cached model from your browser’s private file storage and streams the answer.Arithmetic is computed, not guessed. Addresses open. “remind me…” and “todo:” go where they belong.This device
Follow upSends the conversation so far along with the new question.“Shorter”, “now as an email”, “and if I already pushed it?” all work. The oldest turns drop off to fit the 4,096-token window.This device
ReadRenders the answer as Markdown: lists, bold, code, links. One click copies it.The system prompt tells the model to say when it does not know, and never to invent names, dates, quotes, or numbers.This tab
CheckCounts every request the page made, by host, in Settings under Network, this tab.Read from the browser’s own timeline. Nothing inferred.You can see it

No account. No analytics. The conversation lives in the tab and leaves with it.

How it works

One download, sized to your machine.

Pick a size

The first screen looks at your GPU and memory and recommends a tier. You can pick another any time in Settings.

recommended, never forced

Fetch it once

Qwen3.5, quantized, from Hugging Face. It lands in your browser’s private file storage (OPFS) and never downloads again.

533 MB · 1.3 GB · 2.7 GB

Ask

The model runs through wllama, llama.cpp compiled to WebAssembly, on your GPU through WebGPU or across your CPU cores. If WebGPU fails, it falls back to the CPU on its own.

warm load 1.3 to 3.2 s · 44 to 92 tok/s · Apple silicon, WebGPU
0.8B · 533 MB

Fast

Quick rewrites and summaries. Loads in about a second once cached. The pick for machines without WebGPU.

92 tok/s
2B · 1.3 GB

Balanced

Noticeably smarter. Recommended on WebGPU machines with 8 GB of memory or more.

71 tok/s
4B · 2.7 GB

Best

The first tier that behaves like an assistant: it got the fact, the arithmetic, and the sorting right in our tests. Recommended on Apple silicon.

44 tok/s
When you want more

Four ways to get more out of it.

Sources

Search the web first

Tick one box and Fogar looks the question up before answering, then cites what it found. Free with no key, or full web results with your own key. Answers with sources.

Cloud

Your key, your endpoint

Point Fogar at any OpenAI-compatible endpoint: OpenAI, OpenRouter, Groq, or Ollama on localhost. The key stays in this browser and goes only to the address you typed.

Files

Ask about a document

Drop a PDF, Word, or text file on the page. It is read in the tab, never uploaded, and the answer comes from it. About attachments.

Writing

What small models do best

Rewrite, reply, summarize, explain. Forms for each, and ones you make yourself. About recipes.

Straight talk

A small model, not a big one.

Every mode, and what it sends. See the ledger.

  • It is not ChatGPT. A 4B model on your laptop writes well and knows less. The 0.8B and 2B get facts wrong more often, which is why Fogar says so and offers web search for facts.
  • It is only as fast as your machine. The speeds above are Apple silicon with WebGPU. On the CPU, reading a 900-word letter took 10 seconds on the 0.8B before the first word.
  • The window is 4,096 tokens. Long conversations forget their start, and a long file is read in part, with a chip that says how much.
  • The ask box does not keep history. Close the tab and the conversation is gone. The LLM Chat widget keeps conversations if you want that.
  • When it loads the model online, Fogar can still make a small check with Hugging Face. Your question is never in it, and the ledger shows it.
Questions

Fair questions, short answers.

Is this a private ChatGPT alternative?

For everyday questions, rewriting, and explaining, yes: the model runs in your browser and nothing you type is sent anywhere. It is much smaller than ChatGPT, so it knows less. When you need more, switch to Cloud with your own key, or tick “Search the web first” for facts.

Does it really work offline?

Yes, in local mode, once the model is cached. Turn off Wi-Fi and it keeps answering. Web search and cloud mode need a connection, by definition.

Do I need an account or an API key?

No. There is no Fogar account and no server. A key is needed only if you choose to use a cloud model or a paid search provider, and it stays in this browser.

Which model does it use?

Qwen3.5 in three sizes: 0.8B (533 MB), 2B (1.3 GB), and 4B (2.7 GB), quantized to 4 bits. Fogar recommends one from your GPU and memory. A Qwen3 0.6B is there for machines with very little memory.

Will it slow down my browser?

The model loads only when you ask something, and a tab nobody has looked at for five minutes releases it. Loading from disk takes one to three seconds. Without WebGPU it runs on the CPU, slower but working.

How do I know nothing is sent?

Open Settings and look at Network, this tab. It counts every request the page made, by host, from the browser’s own timeline. The code is open source too, so you can read what the page does.

Install

Add to Chrome

Fogar is on the Chrome Web Store. One click to add it, and it updates itself from there. Then open a new tab.

Free, from the Chrome Web Store. Brave, Arc, and Edge install from it too.

On your phone, or not on Chrome? The same Fogar runs as a web app at app.fogar.ai. Open it once with a connection, add it to your home screen (Safari: Share, then Add to Home Screen; Chrome: the menu, then Add to Home screen), and set it up from there; after that it works offline. The widgets that need Chrome’s own APIs, Sessions, bookmarks, Inbox, Feed, Jira, and Agenda, stay on the desktop.

Loaded the zip before the listing went live? The store copy has a different extension ID, so your data will not follow on its own. Three steps move it over.

  1. In your current Fogar, open Settings and click Download a backup.

    Todos, notes, recipes, widgets, sessions, and settings go into one file.

  2. Add Fogar from the store.

    It opens on your next new tab, empty.

  3. Open its Settings and Restore from a backup.

    Once it looks right, remove the old copy in chrome://extensions.

Can't use the store? The same build is here as a zip: unzip it, turn on Developer mode in chrome://extensions, and Load unpacked. It will not update itself.