Pick a folder
Any project on your disk. Nook works on a private copy, so the agent can try, fail and try again without touching your files.
Point Nook at a folder and say what you want changed. A coding model on your GPU reads the project, edits the files and runs your tests on a private copy, then hands you the diff to keep or throw away.
Windows 11 · NVIDIA 8 GB or more recommended · No account, no cloud, no upload
Three steps, and your folder is only touched when you say so.
Any project on your disk. Nook works on a private copy, so the agent can try, fail and try again without touching your files.
In plain words: add a field, fix a failing test, rename a class across the project. The model reads the code it needs and edits the files.
Nook runs your tests and hands you the changes. Keep what you like, throw away the rest.
Cloud coding agents send your repository to someone else's servers on every request. A local one doesn't.
Client work under NDA, proprietary code, a codebase you are not allowed to upload: the model runs on your card and nothing is sent. Everything Nook serves listens on 127.0.0.1 only.
Open models answer as often as you ask. No usage cap that resets at the end of the week, no metered API key.
Once a model is downloaded, the agent works on a train, on a plane, or with the network cable pulled.
Qwen3-Coder and gpt-oss to start, or any GGUF on Hugging Face, sized against your card before you download it.
Nook is also an MCP server. Claude Code, Claude Desktop, Codex and Cursor can hand it the work that should stay on your machine: search a repository by meaning, pull the lines of a long log that answer a question, transcribe a recording, make a picture. One click on Home writes the entry into each tool's configuration, with a backup beside the file. What the agent then does with a result is up to the agent and its provider.
| Requirement | |
|---|---|
| System | Windows 11. Per-user install, no admin rights, no WSL, no account. |
| Graphics card | An NVIDIA card with 8 GB runs it; more memory runs bigger coding models. AMD and Intel cards run through Vulkan. |
| Engine | llama.cpp, placed on your GPU by the Nook Runtime, which sizes each model to your VRAM. |
| Models | Qwen3-Coder, gpt-oss, or any GGUF from Hugging Face. |
No. The coding model runs on your GPU, the agent works on a private copy of your folder, and nothing you type or change is sent anywhere. The only downloads are the models themselves, from Hugging Face.
Qwen3-Coder and gpt-oss to start. You can also browse Hugging Face for any GGUF model; Nook checks it against your card's memory before you download it.
Open models have closed much of the gap. For a change in a codebase they are more than enough, and they answer with no rate limit and no bill. The very largest cloud models still lead on the hardest reasoning; Nook lets you run the biggest model your card can hold.
Yes. Nook is an MCP server over Streamable HTTP at 127.0.0.1:41434/mcp, or Nook.exe --mcp over stdio. Connect a tool with one click on Home.
Yes. An NVIDIA card with 8 GB runs Nook, and more memory lets you run a larger coding model. AMD and Intel cards work through Vulkan.
One app for the AI you run at home: a coding agent, pictures, video, and speech to text with Whisper. All of it on your GPU.
A Windows desktop app: one download, no account, no admin rights, and it keeps itself up to date.