In one month, my coding agents logged 20.8 billion logical tokens, including repeated conversation history. That made me ask: what changes when I own the machine that keeps that history?
Buy the memory once. Reuse cached context without a per-token bill. A metered API can charge for the same cached history on every turn. Locally, I pay for the machine and its operation, not another token fee each time I reuse that context. Electricity, processing and hardware wear still cost money; the opportunity is to spread those costs across useful work, day and night.
I explore that opportunity with Z.ai’s GLM-5.3-Flash—model and provider pricing on OpenRouter—and quantized local GLM. The question is how much cheaper suitable agent work could become over the next year if faster workstations and serving software deliver. The savings below are conditional scenarios, not results I have measured or a claim of equal quality across models.
First, look inside one agent turn.
Each turn has two jobs: prefill processes the input; decode generates new tokens. A reusable prefix cache lets the next turn skip processing unchanged history from scratch. That is the work I want to avoid repeating while my agents move between prompts, tools and tests.
Read. Remember. Generate.
Process the prompt
The first request processes its input and builds reusable context state.
Retain the state
The cache holds intermediate state, not a saved answer.
Generate new tokens
Each next token still needs model computation and access to context state.
Give every agent a place to remember.
Keeping context is useful only if it fits. GLM-5.3-Flash’s compact cache can leave room for many agents beside one shared model. The memory map below compares three separate single-machine scenarios: one M3 Ultra Mac with 512GB, one M5 Ultra Mac with 512GB, or a hypothetical AMD workstation with 400GB. These are memory-budget estimates, not tests of connected machines; I compare speed and price later.
One model. Many resident contexts.
Each square is a place for one agent’s context—not a promise that all agents can generate at once.
Reuse the same prefix. Keep the same memory.
At 128K tokens, one GLM-5.3-Flash cache read costs $0.0039 at the standard API rate. Locally, that read has no separate token fee.
Adjust the memory budget
These are sizing assumptions, not measured allocations. The 176GB allowance assumes a suitable quantized model, not the native FP8 checkpoint. The 400GB AMD pool is a hypothetical lower-end scenario, not an extra 400GB of system RAM. All pools must be accessible to the inference engine.
Why the footprint is small—and how the counts are calculated
The model configuration specifies 45 layers, including 34 linear-attention and 11 sparse-MLA layers with a 512-wide latent and no extra rotary dimension. The latent payload is 11 × 512 × 2 = 11,264 bytes/token in BF16 (11 KiB), or 5.5 KiB in FP8. At 131,072 tokens that is about 1.48GB or 0.74GB respectively, before other state. Cache precision is separate from weight precision.
I budget 14.5 KiB/token for growing cache data and 0.15GB per resident context for recurrent and other per-agent state. That extra allowance is a sizing input, not an engine guarantee. For scale, 34 × 64 × 128 × 128 FP32 values alone occupy about 0.143GB. A published deployment report shows that recurrent-state allocation and serving-engine limits can constrain concurrency independently of available KV tokens; speculative state can multiply the requirement.
128K = 131,072 tokens; 1M = 1,048,576 tokens 1 KiB = 1,024 bytes; 1 GB = 1,000,000,000 bytes Per-agent GB = tokens × KiB/token × 1,024 / 1e9 + extra state Default 128K slot = 1.946157056 + 0.15 = 2.096157056 GB Resident slots = floor(max(0, RAM − model − reserve) / slot) 512GB Mac: floor((512 − 176 − 50) / 2.096157056) = 136 400GB hypothetical AMD: floor((400 − 176 − 50) / 2.096157056) = 83
Ignoring additional per-agent state gives an idealized 146-context Mac budget. Including the default allowance, 1M contexts give 18 Mac slots or 11 AMD slots. Both Macs have the same modeled memory capacity; faster bandwidth may improve serving speed, not the number of bytes available. No prefix sharing between different agents is assumed. Actual model allocation, cache representation, state snapshots and allocator limits must be checked on the intended engine.
For a separate billing example, reusing a fixed 128K prefix 480 times a day for 30 days would cost $56.62 in GLM-5.3-Flash API cache reads alone at $0.03/M, or $28.31 at the half-price provider rate. Locally, the token fee is zero—not the whole cost of serving those turns. This reuse schedule is illustrative, not a throughput promise.
Now give those agents useful work.
Fitting contexts in memory answers one question: how many agents can stay resident? The next is how much useful work I can get from the machine. An agent waiting for a tool can keep its context while another takes a turn. Enough ready work can turn memory capacity into a well-used system.
Reuse the context.
Retain compatible attention state so a returning agent can avoid processing its unchanged prefix again.
Fill the idle hours.
While one agent waits for tests or tools, another can run. Overnight work spreads ownership cost across more useful output.
Share the work.
A serving engine can process multiple sequences together and reuse weight reads. More agents can raise aggregate throughput—not just utilization.
Busy is not the same as efficient. Going from 90% to 100% utilization adds only 11% more time. A 2× gain in aggregate decode throughput is a different lever. It must be measured; two agents do not automatically produce a 2× gain.
What evidence supports batching?
vllm-mlx reports Qwen3-30B-A3B increasing from 98.1 to 233.3 aggregate tokens/s with five requests on an M4 Max. A first-hand M5 Ultra review found only a 23% aggregate gain from three concurrent Flash-Next requests. Neither establishes a GLM batching multiplier on a 512GB M5.
A GLM serving test on four MI355X virtual functions rose from about 262 to 2,331 aggregate output tokens/s between concurrency 1 and 48. That demonstrates the opportunity on a different stack, not single-MI455X performance. Larger batches can also slow individual agents. Long contexts, expert routing and cache pressure limit the gain.
Today’s reference. Tomorrow’s possibility.
That gives me two hardware questions: how much context can stay resident, and how quickly can I serve it? These three scenarios put a price on both: an M3 reference, an M5 projection, and a hypothetical AMD workstation that I would like to see within the next twelve months, through September 2027.
M3 Ultra

- Unified memory
- 512GB
- Memory bandwidth
- 819GB/s
- Complete system · assumed budget
- $9,499USD
A concrete starting point. Published single-stream GLM rates, combined with assumed ownership costs.
27% modeled saving90% utilization · no batching gainM5 Ultra

- Unified memory
- 512GB
- Memory bandwidth
- 1.2TB/s
- Complete system · assumed budget
- $17,000USD
Faster memory, faster prefill, and enough agent demand to use an assumed 2× decode gain.
61% modeled saving95% utilization · projected speedsMI450-class workstation

- HBM4 · hypothetical
- 400+GB
- Memory bandwidth
- ~23TB/s
- Complete system · assumed budget
- $40,000USD
What if AMD brought roughly 23TB/s memory to an accessible workstation? Assume it sustains 400 effective output tokens/s.
83% modeled saving95% utilization · $40k complete systemHardware facts, projections and the hypothetical AMD workstation
The M3 benchmark reports 437.8 prompt tokens/s and 21.8 generated tokens/s for 4-bit GLM-5.3-Flash at 65,536 tokens on an 80-core GPU M3 Ultra with 512GB. My $9,499 budget and 180W power are assumptions.
Apple specifies 1.2TB/s bandwidth for M5 Ultra, versus 819GB/s on M3 Ultra, and advertises up to 4× faster prompt processing. The 512GB version is scheduled for late October. I project 31.94 decode tokens/s by bandwidth scaling, and offer unchanged or 4× prefill. These are not GLM measurements. The $17,000 purchase price and 200W draw are placeholders.
The actual MI455X has 432GB HBM4 and 23.3TB/s bandwidth, but is a liquid-cooled accelerator module for rack-scale infrastructure. The workstation imagined here is not an announced AMD product. I assume a complete $40,000 machine, 2kW average wall draw, $200/month additional overhead, and user-selected effective throughput. Neither its availability nor these inputs are established facts.
400 effective output tokens/s is a target to test, not a bandwidth-derived prediction. It includes time spent processing input. If achieving it needs multiple accelerators, their full system cost and power must replace the single-workstation assumptions.
When does owning the machine cost less?
The cards summarize the opportunity. Now I want to see whether the same useful work costs less after paying for the whole machine. The calculator scales the estimated input/output mix from my agent month and compares only the volume served locally—not every task my agents might attempt.
Start with the M5 projection and an assumed 2× decode gain, then lower utilization, reduce API prices or limit demand. The question is how much of the advantage survives.
Same served volume. Different cost curves.
M5 Ultra · projected speeds + assumed 2× aggregate decode throughput
API baseline: GLM-5.3-Flash. Per million tokens: $0.03 cached input · $0.15 fresh input · $0.50 output. These are Z.ai’s standard rates, not the cheapest available offer. OpenRouter model & provider prices ↗
What about cheaper OpenRouter providers?
Checked September 22, 2026: OpenRouter lists GMICloud and DeepInfra offers at $0.015 cached input, $0.075 fresh input and $0.25 output per million tokens—half this baseline. Set “API price sensitivity” to 50% to compare those token rates: the default M5 saving falls from 61% to 22%, and the hypothetical AMD saving from 83% to 67%. Provider availability, cache support and performance vary. These comparisons exclude platform funding fees and tax; the linked listing is not a live feed into this calculator.
Cost, cache and demand assumptions
Full-month local budget versus the accumulating API bill.
Open the math: token mix, equations and reference scenarios
My agents logged 20.8B logical tokens in 30 days across two Claude Code Max 20× subscriptions costing $400. The estimated split is 20.4B cached input, 326M fresh input and 74M output—not an exact category export. Logical tokens include repeated history; they are not unique text or a compute measure.
GLM-5.3-Flash API prices, checked September 22, 2026, give $612 cached input + $48.90 fresh input + $37 output = $697.90 per reference mix. This reprices token volume; it does not establish equal quality or tokenization across Claude, hosted GLM and quantized local GLM. The subscription is context, not the scalable comparator.
Local monthly cost = purchase / ownership months + wall kW × 720 × electricity price + overhead Mac processing seconds per reference mix = (326M + local cache-miss fraction × 20.4B) / prefill rate + 74M / (single-stream decode rate × assumed gain) AMD processing seconds = 74M / effective output rate (effective rate already includes input-processing time) Mixes served = 2,592,000 × utilization / processing seconds (capped at 1 if the demand limit is enabled) API bill = mixes served × $697.90 × API price factor Saving = API bill − full local monthly cost
At 95% utilization, the faster M5 projection gives 28%, 61% and 78% savings for 1×, 2× and 4× decode throughput. The corresponding output volumes are 72.8M, 135.5M and 238.1M tokens/month, each with its proportional input. The 2× scenario needs 1.83 reference months of demand; the 4× scenario needs 3.22. If I need only the original month, faster service does not create extra avoided spending.
The hypothetical AMD system costs $40,000 / 36 + 2 × 720 × $0.17 + $200 = $1,555.91/month. At 95% utilization, targets of 100, 200 and 400 effective output tokens/s yield 33%, 67% and 83% savings. Break-even is about 67 tokens/s; 75% savings needs about 268. These are requirements for the imagined product, not predictions.
Simple cash payback = purchase / (monthly avoided API spend − electricity − other operating costs), when positive. At the default M5 assumptions it is about 14 months, not an immediate cash saving. The model excludes financing, tax, resale and unentered operating costs. Batching power and cache behavior require measurement.
What has to become true?
The cost model points to an opportunity. Before I rely on it, I need three things to hold in a real deployment.
Price the whole box.
The hypothetical AMD workstation must actually exist, retain fast memory and meet its total system budget—not just an attractive chip price.
Deliver the throughput.
Useful batching, working caches, acceptable latency and good answers. A busy GPU producing retries is not a successful deployment.
Keep useful work ready.
Savings need real demand. API prices may fall, and local cache misses add work. Change those assumptions in the calculator.
Own the steady work.
Rent the peaks.
The steady work stays home.
Keep suitable recurring tasks and reusable context on my machine.
LOCAL / SERVINGThe peaks have somewhere to go.
Rent extra capacity when demand exceeds what I can serve locally.
API / READY FOR A BURSTI would keep suitable recurring work and reusable context local, then use APIs for overflow or tasks that need a stronger model. The opportunity is not to replace every API call. It is to stop renting the steady work when owning it makes more sense.
