your journal and your computer's memory: local-first, and your cloud options
your journal gets processed. what you've seen and heard goes in as text and images, and audio becomes text too. all of it runs locally, by default, on every device, everywhere. the bundled model runs right in your journal, so your thinking never leaves the machine. a cloud lane is there only as an option: for a machine that can't run a local model, or if you'd rather not spend your device's own power on it.
local is the default
out of the box, your journal is processed on your own device:
- thinking and screen-analysis: the reasoning over your journal, and describing what's on screen, run through a bundled local model in your journal. your thinking never leaves the machine.
- transcription, turning audio into text, runs locally by default too, with a small on-device model: roughly ~0.9 GB. the one exception is a choice you make: with confidential processing on, transcription can run on the service too — see the transcribe audio on the service setting in the journal's thinking app. see the lane notes below.
the journal's thinking app never quietly reaches for a cloud model on its own. if a machine can't run a local model and you haven't chosen a cloud lane, there simply is no model to think with yet. it tells you, rather than picking for you.
what it takes to run locally
the bundled thinking model is sized to the machine it runs on: a current consumer machine runs it. the practical floor is roughly 6–8 GB of GPU memory on a supported GPU, or 16 GB of unified memory on apple silicon (the local model itself needs about 13 GB free; the journal's thinking app checks before it loads anything, so a busy 16 GB machine can still fall short in the moment). on disk that's about 3.4 GB on linux, and about 10.5 GB on apple silicon, which runs a larger model, plus roughly 0.9 GB for the transcription model. your journal checks before it loads anything:
- apple silicon: it checks your free memory first and won't turn on the local model if it won't fit, pointing you to a cloud lane instead.
- linux: the local model runs on a supported hardware GPU (AMD, NVIDIA, or Intel via Vulkan). without one, your journal won't run a local model and points you to a cloud lane instead.
most modern machines clear this bar and run local without you doing anything. if yours is below it, that's what the cloud lanes are for.
choosing in the thinking app
open your journal (by default at http://localhost:5015) and go to its thinking app. it lays out how your journal gets processed as a few lanes, one active at a time:
- local: the bundled model runs right in your journal, so your thinking never leaves. this is the default. the lane tells you whether this computer can run one, and flags when it's short on memory or missing a supported GPU before anything loads.
- byo: your own key from Claude, Gemini, or GPT, your own ChatGPT account, or your own endpoint URL. your billing, your control; a key stays in your journal. with ChatGPT, your computer talks to OpenAI directly, and OpenAI sees the parts of your journal that processing uses. this is the cloud option if your machine can't run local.
- confidential processing — available to approved scouts. processing off your own device, without using your device's own power: your thinking leaves your device and goes to confidential hardware sol pbc runs that keeps nothing — no content retained, no human review, nothing used to train. your journal itself stays on your device either way.
(scouting for solstone? the scout benefit is early access to confidential processing.)
if your journal is using too much memory
- open the thinking app and confirm the local lane is active and your machine has the memory and a supported GPU for it — that's the intended footprint on a capable machine.
- if your machine is below the local bar, switch to a cloud lane: your own byo key.
- make sure you're on a current version — recent releases refined these memory checks. see keeping solstone up to date.
- still stuck? see getting help with solstone, or file a request and include your
journal doctoroutput.