02.1agents / llm

A self-hosted assistant on a 64K budget

One container, a chat front end, a knowledge-graph memory, and a provider that answered with nothing.

layer agents / llmtrack ai-agentsread 3 minpublished 2026-10-09status live

What was built

A personal assistant that runs in a single container on a laptop. A chat app is the front end. An open-source agent framework runs the loop: plan, call a tool, read the result, repeat. Memory is a folder of linked Markdown notes the agent can read and write, a small knowledge graph on disk.

The model is swappable. Some days it runs on a local model runtime. Some days it calls a hosted provider. That switch turned out to be where everything broke.

What broke

① The context window is a floor, not a setting

The framework refuses to start below a 64,000-token context window, and it checks the window the model actually loads with, not the number in the config. With the local runtime you have to set it twice: once to tell the runtime to load the model with a 64K window, and once to tell the agent what to expect. Set only one and the agent either refuses to start or quietly truncates.

On a machine with 16 GB of memory, a mandatory 64K window rules out the 20B and 30B models outright. Roughly 8B parameters is the practical ceiling. A 1.7B model fits easily and fails at the job: it can’t hold a multi-step tool loop together.

② An empty reply with no error

Pointed at the hosted provider with an older model ID, the agent returned nothing. No error, no stack trace, no partial answer: just “no final response was produced”.

The key worked. The raw API worked. The framework only recognised a curated list of newer models, and for anything else it couldn’t resolve the model’s context metadata. Without that it gave up silently. Switching to a newer model fixed it instantly.

Two things made this slow to find:

③ The GPU crash that looks like a model bug

On the laptop, the local runtime occasionally crashes compiling a GPU shader. The agent sees it as a dead model. The fix is a restart of the runtime, not a change to the agent. Knowing which layer to restart saves an evening.

The number

64,000 tokens. Every decision downstream (which models fit, which machine is enough, which provider works) fell out of that one floor.

What I’d do again