I published Codendum 0.1.0 today under the Apache 2.0 licence. It is a service for sharing one machine that runs inference with local open-weight models for generative coding. Prompts and source code stay on the organization’s network, and the rest of the work stays on the workstations.
The repository holds the scripts, example configuration and documentation to install the service, lock it down, test it, measure it and operate it. The documentation is a separate site, built with Sphinx and MyST.
What it does
A single NVIDIA GB10 system, the DGX Spark or an equivalent OEM machine, runs vLLM with an open-weight coding model. The workstations run OpenCode, configured as an OpenAI-compatible provider, and talk to that machine over HTTPS on the local network or a VPN. The served model is called coder, it is Qwen3-Coder-30B-A3B-Instruct-FP8, and tool calling is enabled explicitly.
Which languages and platforms are available is decided by the model you put in. The service forwards requests to an OpenAI-compatible endpoint and stays indifferent to the language passing through it, so the coverage is the underlying model’s and it moves when the model changes.
The shape of the service is this:
- vLLM listens on
127.0.0.1only. It cannot be reached from outside. - An nginx container is the only way in. It accepts HTTPS on a dedicated port, checks that the request comes from an allowed network, verifies a per-user API key and forwards only
/v1/chat/completionsand/v1/models. Requests with non-normalized paths are rejected. - The proxy container has a read-only file system, minimal capabilities and
no-new-privileges. - Images and model are pinned by digest and by revision, and the CI actions by commit SHA with Dependabot updates.
The scripts live on the host, drive Docker and act as test clients, without installing packages or services on the system.
Why the host only serves inference
The choice comes from a property of the hardware. The GB10 has 128 GB of unified memory shared between CPU and GPU, not 128 GB of VRAM added to system memory. The operating system, Docker, nginx, the vLLM processes, the model weights and the KV cache all draw on the same pool. It is the same characteristic that made memory bandwidth the number to watch when I wrote about medical inference on this machine.
So users’ code, Git, the toolchains, the compilers and the builds run on the workstations or in isolated development environments. A whole class’s builds, run on the model host, would compete for memory with the KV cache, which is why the documentation says not to do it.
Sizing the memory
The FP8 weights take about 31.2 GB on disk and roughly the same in memory. What is left goes to the KV cache, and that part can be computed.
The model has 48 layers, 4 KV heads and a head size of 128. With an FP8 KV cache, one byte per value, every token held in context costs:
2 (K and V) × 48 layers × 4 KV heads × 128 × 1 byte = 49,152 bytes = 48 KiB per token
From there the theoretical cost of a configuration follows: 16 requests of 65,536 tokens want 48 GiB, 24 requests want 72, and it scales from there with the number of concurrent requests. That is a lower bound, before weights and runtime overhead.
On the test machine CUDA saw 121.6 GiB. The profiles run with --gpu-memory-utilization at 0.80, so vLLM gets about 97, and once weights and overhead are taken out the default profile got 62.4 GiB of KV cache, that is 1,362,640 tokens, or 20.8 full 64K contexts at the same time. vLLM writes these figures to the log at startup and scripts/metrics.sh reads them back, so it is worth reading them on your own host rather than trusting these.
Three numbers I keep apart in the documentation, because they are easy to conflate: connected users are everyone with the client open, and they send nothing most of the time, active requests are the ones being processed and are capped by --max-num-seqs, and waiting requests sit in a queue inside vLLM in order of arrival. When the KV cache runs out, vLLM preempts a running request and recomputes it later, which shows up as a latency spike and as preemptions in the metrics.
Measured on a GB10
On 29 September 2026 I ran the service on a Lenovo ThinkStation PGX with DGX OS 7.2.3 and driver 580.178.04, vLLM 0.30.0, with OpenCode 2.0.19 as the reference client. The simulated students ran in Docker on a remote workstation over WireGuard, with a round trip of about 100 ms, so the latencies include the network.
First the calibration against a real session. One exercise done with OpenCode produced 33 model calls in 63 seconds: one for the title, 20 steps of the main agent and 12 of an exploration subagent. Prompts grew from 6.8K to 11K tokens. With one or two users, time to first token was about 0.4 seconds, each received about 31 output tokens per second, and a coding request finished in 1.7 to 2.2 minutes.
Then the class. A full class of simulated students, four requests each and pauses of 20 to 90 seconds, using the system prompt and tools captured from a real OpenCode session. The window for new requests was fifteen minutes, plus the time needed to finish the ones in flight. Two profiles side by side:
classroom-64k (16 active) | classroom-64k-high-concurrency (24 active) | |
|---|---|---|
| Server output, steady state | ~110 tok/s | ~140–150 tok/s |
| Time to first token, P50 / P95 | 17.5 / 60 s | 10.6 / 34 s |
| Output speed per user, P50 | 7.2 tok/s | 6.1 tok/s |
| Time to complete a coding request, P50 / P95 | 11 / 22 min | 8.3 / 18 min |
| Coding requests completed | 45 | 64 |
| Peak waiting requests | 24 | 15 |
| Peak KV cache usage | 8.5% | 11.5% |
| Preemptions / server errors | 0 / 0 | 0 / 0 |
| Prefix-cache hit rate | 98.2% | 98.2% |
In my runs the limit was generation, not memory. Agent contexts stayed between 7K and 19K tokens and the KV cache never went above 12%, while admitting more requests at once raised the server’s total output. The memory would hold more than 24 active requests, but I did not measure that point.
Prefix caching counts for a lot: 98% of seven million prompt tokens came from the cache instead of from computation. In a class the system prompt and the tool definitions are the same for everyone, and they are paid for once.
Capacity in requests per hour
To size a lab I needed to know how many requests the machine carries in an hour, more than its peak speed. The answer comes from putting two measurements together.
A coding request produces about 2,700 output tokens spread across some fifteen model calls. At 110 to 150 tokens per second, in my measurements one GB10 completed roughly 150 to 200 such requests per hour for the whole class. How many each person gets depends on how many are connected.
The simulation, with pauses under 90 seconds, asks for more than that, and the requests duly queue. A class that asks less waits less.
What stays uncovered
The security model is written out in full in the documentation, with a table of threats and mitigations and the part that stays with whoever runs the service. I summarize here the points I think are worth stating.
There are no per-user token quotas and no fair scheduling. nginx limits requests, not tokens, and vLLM serves first come, first served: a 60K-token request costs far more than a 2K one, and neither layer knows it. Real quotas and priorities need a dedicated LLM gateway in front of vLLM, which Codendum neither includes nor tests.
This is where I would like to bring in Admina, the open source governance framework I work on, so that quotas, policies and tracing sit in a declared layer instead of being scattered through the nginx configuration. For now it is a working idea: there is no code and no date.
The prefix cache is shared between users. It is what makes the service viable, and it is also a timing side channel: in principle response times could reveal that someone else recently sent the same prefix. It can be turned off with one extra argument, at a cost in throughput, and the documentation says how.
Agents run commands on the workstations with the user’s permissions. The boundary sits there, in the client configuration, not in the server — the same point at which a terminal agent inherits the environment you launch it in. Isolating repositories, credentials and build environments is the workstations’ job.
Limits
One machine and no high availability: when the GB10 or vLLM goes down, everyone loses the agent.
The measurements come from one host, one model and one client version. Different hosts, images, models or a different OpenCode release change the numbers, which is why bench.sh and bench-classroom.sh are in the repository, to be run again after every upgrade. About 1% of requests failed on the VPN path without ever reaching the proxy, so that figure belongs to the test network.
Generated code has to be compiled, tested and read by people, and a project is built in steps with review after each one. Short, focused requests are also what keeps a shared service responsive.
Finally, prompts contain users’ code and the nginx logs contain user ids and timestamps. Whoever runs the service answers for the applicable policies, including data protection rules when the users are students or employees, and for log retention. The Apache 2.0 licence covers the files in the repository, while container images, model weights and client software are downloaded separately and keep their own licences.
- Codendum, the repository on GitHub — https://github.com/stefanoferi/codendum
- The 0.1.0 release — https://github.com/stefanoferi/codendum/releases/tag/v0.1.0
- The project’s full documentation — https://stefanoferi.github.io/codendum/
- The benchmarks measured on a GB10 — https://stefanoferi.github.io/codendum/benchmark.html#measured-on-a-gb10
- Memory and KV cache sizing — https://stefanoferi.github.io/codendum/sizing.html
- The security model and responsibilities — https://stefanoferi.github.io/codendum/security.html
- Pinned versions, licences and compatibility — https://stefanoferi.github.io/codendum/reference.html
- The project CHANGELOG — https://github.com/stefanoferi/codendum/blob/main/CHANGELOG.md
- NVIDIA, the DGX Spark product page — https://www.nvidia.com/en-us/products/workstations/dgx-spark/
- vLLM, the server documentation — https://docs.vllm.ai/
- OpenCode, the client used on the workstations — https://opencode.ai/
- Admina, the open source governance framework — https://admina.org/
- Qwen, the FP8 checkpoint of the served model — https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8