The smallest useful amount of intelligence
On the Cardputer, once it has booted, about 18 KB of memory are left. A language model does not fit. That is not why the assistant isn't one.
by Nicola Sabaini 3 min read Project: NucleoOS Cardputer
The Cardputer fits in one hand. Inside is an ESP32-S3 with 512 KB of memory and no expansion, and once NucleoOS has finished booting about 18 of them are free: the space of a two page text file, for everything that happens next.
A language model does not fit in there. That is not why ANIMA, the system assistant, is not a language model.
Grounded by construction
ANIMA is a natural language engine written in C: it searches on three levels, reasons over a knowledge graph, does arithmetic and unit conversion, learns, and keeps a profile of whoever is talking to it.
What it cannot do is say something it does not have. Not because a check stops it, but because there is no step in which an answer would be composed out of nothing. It is a property of the shape, not a rule stuck on top.
Inside an operating system that difference is everything. Ask how much battery is left, or which application has network permission, and a plausible answer is worse than no answer: nobody has a way to notice it is merely plausible. An assistant that invents inside the OS gets found out late, and you find out because you did something you should not have done.
A voice that does not generate
The speech follows the same road. The synthesis is concatenative: a finite vocabulary is recorded once on the PC, and on the device the right clips are glued in the right order.
Natural voice, almost no memory, no network. It reads the time to the minute, confirms launches, says the state of the system.
In exchange it pronounces only what is in the vocabulary. Which is the same property as before, applied to sound: the perimeter of what can come out is known in advance, not afterwards. And for things that should not be said out loud, it answers that they are on the screen.
On a machine with no limits, the same choices
Aria is the assistant inside ExaWar, which runs on a PC: as much memory as it wants, no ceiling. It is built the same way.
It understands by meaning in five languages thanks to a small model that turns sentences into comparable numbers, and that model is optional: if it is missing, Aria keeps working with less finesse instead of shutting down.
When it has to write code in the game’s internal language it does not compose free words. It retrieves relevant examples, then generates from a grammar, so what comes out is valid by construction. Before running it, it runs it dry, watches what happens and corrects. There is no step that evaluates arbitrary text, and the same request always gives the same result.
On a half megabyte chip those choices look forced by the constraint. On a PC nothing forces them, and I made them the same.
When a real model is needed
For the rest there are the frontier models, configurable together, with automatic fallback if one does not answer. Two details matter more than the choice of provider.
A microcontroller cannot hold an encrypted session to the outside while doing everything else, and more to the point it should not. It keeps the state, serves the data, and the heavy work happens on the machine of whoever is looking.
The smallest system that does the job is not a fallback. It is the only one I can describe, before switching it on, in terms of what it will not do.