Local AI Is Back · the cache changes the mathSeptember 2026

Local AI is back. The cache changes the math.

Reuse the context. Keep the machine busy. Let agents share the work. A visual case for why tomorrow’s local hardware could beat a metered API—by a lot.

Sep 2, 202610 min read

24/7agent loops are the opportunity, not a utilization measurement
61%M5 scenario saving · projected speeds + assumed 2× decode gain
83%hypothetical AMD workstation saving · assumed 400 effective output tokens/s

In one month, my coding agents logged 20.8 billion logical tokens, including repeated conversation history. That made me ask: what changes when I own the machine that keeps that history?

Buy the memory once. Reuse cached context without a per-token bill. A metered API can charge for the same cached history on every turn. Locally, I pay for the machine and its operation, not another token fee each time I reuse that context. Electricity, processing and hardware wear still cost money; the opportunity is to spread those costs across useful work, day and night.

I explore that opportunity with Z.ai’s GLM-5.3-Flash—model and provider pricing on OpenRouter—and quantized local GLM. The question is how much cheaper suitable agent work could become over the next year if faster workstations and serving software deliver. The savings below are conditional scenarios, not results I have measured or a claim of equal quality across models.

First, look inside one agent turn.

Each turn has two jobs: prefill processes the input; decode generates new tokens. A reusable prefix cache lets the next turn skip processing unchanged history from scratch. That is the work I want to avoid repeating while my agents move between prompts, tools and tests.

01 / INSIDE AN AGENT TURN

Read. Remember. Generate.

Decode · produce the next token
01

Process the prompt

The first request processes its input and builds reusable context state.

02

Retain the state

The cache holds intermediate state, not a saved answer.

03

Generate new tokens

Each next token still needs model computation and access to context state.

Illustrative, not a benchmark. Caching skips repeated prefill—not attention or decoding. Both local models and APIs do this work.

Give every agent a place to remember.

Keeping context is useful only if it fits. GLM-5.3-Flash’s compact cache can leave room for many agents beside one shared model. The memory map below compares three separate single-machine scenarios: one M3 Ultra Mac with 512GB, one M5 Ultra Mac with 512GB, or a hypothetical AMD workstation with 400GB. These are memory-budget estimates, not tests of connected machines; I compare speed and price later.

02 / THE CONTEXT NEIGHBORHOOD

One model. Many resident contexts.

Memory budget · not a benchmark
136resident agent contexts
2.10GBbudget per context + state
$0local per-token cache-read fee
Model · 176GBReserve · 50GBContext pool · 286GB

Each square is a place for one agent’s context—not a promise that all agents can generate at once.

Reuse the same prefix. Keep the same memory.

At 128K tokens, one GLM-5.3-Flash cache read costs $0.0039 at the standard API rate. Locally, that read has no separate token fee.

0 illustrative reusesAPI cache reads $0.0000Local token fees $0
Adjust the memory budget

These are sizing assumptions, not measured allocations. The 176GB allowance assumes a suitable quantized model, not the native FP8 checkpoint. The 400GB AMD pool is a hypothetical lower-end scenario, not an extra 400GB of system RAM. All pools must be accessible to the inference engine.

Resident does not mean simultaneously generating. A compatible serving engine must retain and reuse each agent’s hybrid attention state; eviction, context growth, speculative decoding and extra state can reduce capacity. Local cache reuse still consumes computation, bandwidth and power. This memory illustration is independent of the throughput calculator below.
Why the footprint is small—and how the counts are calculated

The model configuration specifies 45 layers, including 34 linear-attention and 11 sparse-MLA layers with a 512-wide latent and no extra rotary dimension. The latent payload is 11 × 512 × 2 = 11,264 bytes/token in BF16 (11 KiB), or 5.5 KiB in FP8. At 131,072 tokens that is about 1.48GB or 0.74GB respectively, before other state. Cache precision is separate from weight precision.

I budget 14.5 KiB/token for growing cache data and 0.15GB per resident context for recurrent and other per-agent state. That extra allowance is a sizing input, not an engine guarantee. For scale, 34 × 64 × 128 × 128 FP32 values alone occupy about 0.143GB. A published deployment report shows that recurrent-state allocation and serving-engine limits can constrain concurrency independently of available KV tokens; speculative state can multiply the requirement.

128K = 131,072 tokens; 1M = 1,048,576 tokens
1 KiB = 1,024 bytes; 1 GB = 1,000,000,000 bytes

Per-agent GB = tokens × KiB/token × 1,024 / 1e9 + extra state
Default 128K slot = 1.946157056 + 0.15 = 2.096157056 GB
Resident slots = floor(max(0, RAM − model − reserve) / slot)

512GB Mac: floor((512 − 176 − 50) / 2.096157056) = 136
400GB hypothetical AMD: floor((400 − 176 − 50) / 2.096157056) = 83

Ignoring additional per-agent state gives an idealized 146-context Mac budget. Including the default allowance, 1M contexts give 18 Mac slots or 11 AMD slots. Both Macs have the same modeled memory capacity; faster bandwidth may improve serving speed, not the number of bytes available. No prefix sharing between different agents is assumed. Actual model allocation, cache representation, state snapshots and allocator limits must be checked on the intended engine.

For a separate billing example, reusing a fixed 128K prefix 480 times a day for 30 days would cost $56.62 in GLM-5.3-Flash API cache reads alone at $0.03/M, or $28.31 at the half-price provider rate. Locally, the token fee is zero—not the whole cost of serving those turns. This reuse schedule is illustrative, not a throughput promise.

Now give those agents useful work.

Fitting contexts in memory answers one question: how many agents can stay resident? The next is how much useful work I can get from the machine. An agent waiting for a tool can keep its context while another takes a turn. Enough ready work can turn memory capacity into a well-used system.

01 / REUSE

Reuse the context.

Retain compatible attention state so a returning agent can avoid processing its unchanged prefix again.

02 / OCCUPY

Fill the idle hours.

While one agent waits for tests or tools, another can run. Overnight work spreads ownership cost across more useful output.

03 / BATCH

Share the work.

A serving engine can process multiple sequences together and reuse weight reads. More agents can raise aggregate throughput—not just utilization.

Busy is not the same as efficient. Going from 90% to 100% utilization adds only 11% more time. A 2× gain in aggregate decode throughput is a different lever. It must be measured; two agents do not automatically produce a 2× gain.

What evidence supports batching?

vllm-mlx reports Qwen3-30B-A3B increasing from 98.1 to 233.3 aggregate tokens/s with five requests on an M4 Max. A first-hand M5 Ultra review found only a 23% aggregate gain from three concurrent Flash-Next requests. Neither establishes a GLM batching multiplier on a 512GB M5.

A GLM serving test on four MI355X virtual functions rose from about 262 to 2,331 aggregate output tokens/s between concurrency 1 and 48. That demonstrates the opportunity on a different stack, not single-MI455X performance. Larger batches can also slow individual agents. Long contexts, expert routing and cache pressure limit the gain.

Today’s reference. Tomorrow’s possibility.

That gives me two hardware questions: how much context can stay resident, and how quickly can I serve it? These three scenarios put a price on both: an M3 reference, an M5 projection, and a hypothetical AMD workstation that I would like to see within the next twelve months, through September 2027.

PUBLISHED SPEED REFERENCE

M3 Ultra

AI-generated illustration of a silver Mac Studio-style desktop for the M3 Ultra reference
AI illustration · M3 reference
Unified memory
512GB
Memory bandwidth
819GB/s
Complete system · assumed budget
$9,499USD

A concrete starting point. Published single-stream GLM rates, combined with assumed ownership costs.

27% modeled saving90% utilization · no batching gain
M5 + BATCHING SCENARIO

M5 Ultra

AI-generated Mac Studio-style illustration for the projected M5 Ultra scenario, not an official product photo
AI illustration · M5 scenario
Unified memory
512GB
Memory bandwidth
1.2TB/s
Complete system · assumed budget
$17,000USD

Faster memory, faster prefill, and enough agent demand to use an assumed 2× decode gain.

61% modeled saving95% utilization · projected speeds
HYPOTHETICAL PRODUCT

MI450-class workstation

AI-generated concept of a graphite workstation tower for the hypothetical AMD scenario, not a real product
AI concept · not an announced product
HBM4 · hypothetical
400+GB
Memory bandwidth
~23TB/s
Complete system · assumed budget
$40,000USD

What if AMD brought roughly 23TB/s memory to an accessible workstation? Assume it sustains 400 effective output tokens/s.

83% modeled saving95% utilization · $40k complete system
Hardware facts, projections and the hypothetical AMD workstation

The M3 benchmark reports 437.8 prompt tokens/s and 21.8 generated tokens/s for 4-bit GLM-5.3-Flash at 65,536 tokens on an 80-core GPU M3 Ultra with 512GB. My $9,499 budget and 180W power are assumptions.

Apple specifies 1.2TB/s bandwidth for M5 Ultra, versus 819GB/s on M3 Ultra, and advertises up to 4× faster prompt processing. The 512GB version is scheduled for late October. I project 31.94 decode tokens/s by bandwidth scaling, and offer unchanged or 4× prefill. These are not GLM measurements. The $17,000 purchase price and 200W draw are placeholders.

The actual MI455X has 432GB HBM4 and 23.3TB/s bandwidth, but is a liquid-cooled accelerator module for rack-scale infrastructure. The workstation imagined here is not an announced AMD product. I assume a complete $40,000 machine, 2kW average wall draw, $200/month additional overhead, and user-selected effective throughput. Neither its availability nor these inputs are established facts.

400 effective output tokens/s is a target to test, not a bandwidth-derived prediction. It includes time spent processing input. If achieving it needs multiple accelerators, their full system cost and power must replace the single-workstation assumptions.

When does owning the machine cost less?

The cards summarize the opportunity. Now I want to see whether the same useful work costs less after paying for the whole machine. The calculator scales the estimated input/output mix from my agent month and compares only the volume served locally—not every task my agents might attempt.

Start with the M5 projection and an assumed 2× decode gain, then lower utilization, reduce API prices or limit demand. The question is how much of the advantage survives.

03 / THE OWNERSHIP CROSSOVER

Same served volume. Different cost curves.

Scenario model · USD

M5 Ultra · projected speeds + assumed 2× aggregate decode throughput

API baseline: GLM-5.3-Flash. Per million tokens: $0.03 cached input · $0.15 fresh input · $0.50 output. These are Z.ai’s standard rates, not the cheapest available offer. OpenRouter model & provider prices ↗

What about cheaper OpenRouter providers?

Checked September 22, 2026: OpenRouter lists GMICloud and DeepInfra offers at $0.015 cached input, $0.075 fresh input and $0.25 output per million tokens—half this baseline. Set “API price sensitivity” to 50% to compare those token rates: the default M5 saving falls from 61% to 22%, and the hypothetical AMD saving from 83% to 67%. Provider availability, cache support and performance vary. These comparisons exclude platform funding fees and tax; the linked listing is not a live feed into this calculator.

Assumed aggregate gain, not agent count.
Cost, cache and demand assumptions
61%lower modeled monthly cost
$497local ownership / month
$1,278GLM-5.3-Flash API / month
$781saving / month (− = extra cost)

A MONTH OF THE SELECTED WORKLOAD
API · accrued
$1,278
Local · full budget
$497

Full-month local budget versus the accumulating API bill.

Local monthly costGLM-5.3-Flash API, same volume

The bars animate calculated bills, not measurements. The full monthly ownership budget includes amortized capital; it is not cash paid on day one. Choosing hardware resets its budget, power, utilization and throughput assumptions; advanced cost and demand settings stay yours.
Open the math: token mix, equations and reference scenarios

My agents logged 20.8B logical tokens in 30 days across two Claude Code Max 20× subscriptions costing $400. The estimated split is 20.4B cached input, 326M fresh input and 74M output—not an exact category export. Logical tokens include repeated history; they are not unique text or a compute measure.

GLM-5.3-Flash API prices, checked September 22, 2026, give $612 cached input + $48.90 fresh input + $37 output = $697.90 per reference mix. This reprices token volume; it does not establish equal quality or tokenization across Claude, hosted GLM and quantized local GLM. The subscription is context, not the scalable comparator.

Local monthly cost = purchase / ownership months
  + wall kW × 720 × electricity price + overhead

Mac processing seconds per reference mix =
  (326M + local cache-miss fraction × 20.4B) / prefill rate
  + 74M / (single-stream decode rate × assumed gain)

AMD processing seconds = 74M / effective output rate
  (effective rate already includes input-processing time)

Mixes served = 2,592,000 × utilization / processing seconds
  (capped at 1 if the demand limit is enabled)

API bill = mixes served × $697.90 × API price factor
Saving = API bill − full local monthly cost

At 95% utilization, the faster M5 projection gives 28%, 61% and 78% savings for 1×, 2× and 4× decode throughput. The corresponding output volumes are 72.8M, 135.5M and 238.1M tokens/month, each with its proportional input. The 2× scenario needs 1.83 reference months of demand; the 4× scenario needs 3.22. If I need only the original month, faster service does not create extra avoided spending.

The hypothetical AMD system costs $40,000 / 36 + 2 × 720 × $0.17 + $200 = $1,555.91/month. At 95% utilization, targets of 100, 200 and 400 effective output tokens/s yield 33%, 67% and 83% savings. Break-even is about 67 tokens/s; 75% savings needs about 268. These are requirements for the imagined product, not predictions.

Simple cash payback = purchase / (monthly avoided API spend − electricity − other operating costs), when positive. At the default M5 assumptions it is about 14 months, not an immediate cash saving. The model excludes financing, tax, resale and unentered operating costs. Batching power and cache behavior require measurement.

What has to become true?

The cost model points to an opportunity. Before I rely on it, I need three things to hold in a real deployment.

THE HARDWARE

Price the whole box.

The hypothetical AMD workstation must actually exist, retain fast memory and meet its total system budget—not just an attractive chip price.

THE ENGINE

Deliver the throughput.

Useful batching, working caches, acceptable latency and good answers. A busy GPU producing retries is not a successful deployment.

THE WORKLOAD

Keep useful work ready.

Savings need real demand. API prices may fall, and local cache misses add work. Change those assumptions in the calculator.

Own the steady work.
Rent the peaks.

04 / A HOME FOR EVERY TURN
Local capacity handles steady demand; the API handles overflow An orange band represents work served locally. Mint-colored peaks above the local capacity line represent extra work sent to an API. Demand changes over time; the allocation is illustrative, not a throughput measurement. Local capacity Work over time →
Local · owned capacityAPI · rented overflow

The steady work stays home.

Keep suitable recurring tasks and reusable context on my machine.

LOCAL / SERVING

The peaks have somewhere to go.

Rent extra capacity when demand exceeds what I can serve locally.

API / READY FOR A BURST
Illustrative routing, not a measured deployment. API overflow adds to the bill and is not included in the local-only savings above. Moving context can add cost; tasks needing a stronger model may also go to the API.

I would keep suitable recurring work and reusable context local, then use APIs for overflow or tasks that need a stronger model. The opportunity is not to replace every API call. It is to stop renting the steady work when owning it makes more sense.