Get Nook
  1. Nook
  2. Labs
  3. Note 000

Note 000About LabsInference

Introducing Nook Labs

Nook Labs is where we test AI inference on the kind of computers people actually own, and publish what we find: the setup, the numbers, the runs that went wrong, and how to repeat it on your own machine.

Why a lab

Nook is a desktop app that runs AI on your own computer: a coding agent, speech, documents and more, on your graphics card instead of someone else's servers. Building it means answering questions that no spec sheet answers. Will this model fit on an 8 GB card? Is the CUDA build of an engine worth a download of more than a gigabyte? How small can a model be and still do the job? What actually happens when a card runs out of memory?

We have been answering questions like these for months, one test at a time, because the app depends on the answers. Until now the answers ended up in commit messages and team chats, where nobody else could use them. Nook Labs is where they go from now on.

What we test

Inference, on the hardware people actually have: a mid-range graphics card with 8 GB of memory, a laptop, a computer that is a few years old. In particular:

  • Speed and memory. Tokens a second, seconds per job and peak video memory, on cards we name.
  • Engines and backends. llama.cpp, whisper.cpp, stable-diffusion.cpp and audio.cpp, on CUDA, Vulkan, Metal and the processor alone.
  • Model choices. Quantisation levels, small models against big ones, and what a model that fits on any machine can still do.
  • Quality, not only speed. A faster setting that gets the answer wrong is not faster. Where we can, we score the output against a fixed set of tasks.
  • Failure. What breaks, how it breaks, and how to tell that it has.

How we run an experiment

Six rules for every note:

  1. Name the hardware. Every note lists the card, the driver, the operating system, the engine and its version, and the model file down to its quantisation. A number without its setup is a rumour.
  2. One question per note. Asked at the top, answered in the first few lines, then shown.
  3. Publish the boring and the broken. If an experiment fails, or the answer is “no difference”, that goes up too.
  4. Say how sure we are. One run is called one run. A small sample is called small.
  5. Make it repeatable. Commands, settings and test files wherever we can share them, so you can run the same test on your machine.
  6. Correct in the open. If we get something wrong, the note changes and says what changed and when.

We test what you can run yourself: open models, on engines you can download. Nook's own source code is public on GitHub, so when a note measures something Nook does, you can read the code that does it.

Three things we have already learned

Here are three findings from building Nook, written up the way Labs notes will be. They come from our build notes: one machine, mostly single runs, so read them as first looks rather than verdicts. Each will get a full note of its own.

1. The CUDA build of our voice engine is more than twice as fast

Nook's speech translator can speak a translation in the speaker's own voice, using the audio.cpp engine. It first ran on audio.cpp's Vulkan build, which works on any graphics card. We built the CUDA version for NVIDIA cards and timed the two back to back, cloning a voice to speak the same three lines.

Setup
Graphics card
NVIDIA GeForce RTX 4060, 8 GB
System
Windows 11
Engine
audio.cpp, CUDA 12.4 build against the Vulkan build
Task
Clone a voice and speak three lines
Runs
One of each, back to back
Time to speak three lines in a cloned voiceSeconds, lower is better. RTX 4060, one run each.
  1. CUDA build8.2 s
  2. Vulkan build18.3 s

The CUDA build took less than half the time, and its peak video memory was half a gigabyte lower: 3.4 GB against 3.9 GB. Since Nook 0.5.6, NVIDIA cards get the CUDA build; other cards keep Vulkan, and computers without a usable graphics card run on the processor. The cost is the download: with NVIDIA's libraries included, the CUDA build comes to about a gigabyte.

2. On Windows, running out of video memory does not fail. It crawls.

A translation we were testing sat at “Speaking · 0 of 1” for minutes, with no error. The cause: Whisper, the speech-to-text model, was still loaded when the voice model started, and together they did not fit in 8 GB. On Windows, a program that runs short of video memory can spill into the computer's ordinary memory instead of stopping, and then everything slows to a crawl. With the Vulkan build, Windows does this on its own. With CUDA it depends on a driver setting, NVIDIA's “CUDA – Sysmem Fallback Policy”, which allows it by default.

Since Nook 0.5.5, the app unloads every idle model before the voice speaks. The lesson for anyone running AI locally on Windows: if a model is suddenly far slower than it should be, check whether it still fits. Task Manager's Performance tab shows the card's dedicated and shared memory side by side; if shared memory climbs while the model runs, the model does not fit.

3. A 132 MB model can pick the right tool

Nook's Nooklets are small ready-made jobs: transcribe a recording, edit a PDF, translate speech, convert a document. You ask for one in plain words and a model picks it. We wanted that to work on any computer, so we tried a small model: multilingual-e5-small, in 8 bits, 132 MB, running on the processor alone. It turns a request into a list of numbers and compares it with examples of each Nooklet.

It passed all 30 cases in our test set. Requests that matched no Nooklet scored between 0.83 and 0.85, just under the 0.88 we treat as a sure match: a narrower margin than we would like. Thirty cases is a small test, and the full note will use many more. The point stands, though: not every AI job needs a big model. Picking from a short menu is a job for something tiny.

What is next

These questions are on our list. Tell us which one you want answered first.

  • What quantisation costs. The same coding model at 8, 5, 4 and 3 bits: speed, memory, and how many tasks it still gets right.
  • What really fits in 8 GB. The biggest chat, speech and coding models an 8 GB card runs at a usable speed, and where each one falls over.
  • Whisper, size by size. From tiny to large: speed against accuracy on real recordings, on a graphics card and on the processor.
  • CUDA, Vulkan or the processor. The voice test above done properly, across engines and cards, with many runs.
  • Apple silicon against a PC. A Mac and a Windows PC at similar prices, on the same jobs.

Follow along

New notes go up here, on the RSS feed and on @usenook. If you want something tested, or you ran one of our tests and got a different number, write to hello@usenook.ai. Send your setup with it, and we will add readers' results to the note.

Got a different number on your machine, or think we got something wrong? Tell us at hello@usenook.ai or @usenook, with your setup. Corrections are made in the open, dated, at the end of the note.

Run it on your own GPU.

Nook runs open models on your own computer: no account, no upload, and it keeps itself up to date.