Field brief · August 28, 2026

Grok Build is metering you. GLM-5.3-Flash on 4× RTX PRO 6000 is not.

Two looping projects ate 30% of a weekly Grok Build pool. That is the product working as designed — agent loops are compute-weighted, not “messages.” This brief sizes the downgrade to GLM-5.3-Flash, the quantization to run on a 4× RTX PRO 6000 box and on M5 Ultra 256GB (one box, four, and six), how many concurrent agents actually feel good, whether pegging those farms can print $10,000 of Grok 4.6 API-equivalent work in a month, and what that same volume invoices on OpenRouter Flash — plus what $300 of Flash credits actually buy.

Aug 28, 2026 Grok 4.6 GLM-5.3-Flash RTX PRO 6000 ×4 M5 Ultra 256GB OpenRouter

BLUF — Bottom line up front

The intelligence gap is four Artificial Analysis Index points (Grok 4.6 high = 61, Flash = 57). The money gap is ~10× per finished eval task. For looping coding agents, Flash is a mild quality cut and a large capacity unlock.

1. Why 30% already feels like a lot

Grok subscriptions since June 2026 share one compute-weighted weekly pool across Chat, Imagine, Voice, and Build. xAI does not publish a token number. The Usage page shows a percentage. Agent and Build tasks draw more than chat by design — Grok said so on X when Heavy users burned 25–30% in a day of backend work.

That matches the lived experience: two projects, a couple of looping tasks, 30% gone. You are not “light.” Looping agents re-send growing context every turn. The meter is weighting that, not counting prompts.

The only public, instrumented SuperGrok week I trust is Ben Bishop’s $30 plan hitting 100% after 31 Grok Build sessions (Mission Control extract, week of 10–17 Aug 2026):

Sessions

31

one $30 SuperGrok week to 100%

Input

9.66M

fresh prompt tokens

Output

1.23M

completions

Cache read

306M

reused agent prefix

Price that mix at Grok 4.6 API rates ($2 / $0.50 cached / $6 per 1M, prompts <200K):

9.658M × $2  +  305.717M × $0.50  +  1.231M × $6  =  $179.56

So 100% of a $30 SuperGrok week of real Build work is about $180 of Grok 4.6 API. Thirty percent of that week is about $54 of API. Bishop is explicit that this is not “the official quota in tokens” — it is the workload recorded while the meter hit 100%. It is still the best ruler we have.

SuperGrok Heavy is $300/month and xAI still does not publish a token cap. Community reports (30% of Heavy in <24h of Build) mean Heavy is larger than SuperGrok, not that it is 10× linear, and not that it is unlimited. Model both below.

2. The two models

Grok 4.6 (high) · xAI

AA Index 61 $0.94 / AA task 500K ctx
  • API: $2 / $6 per 1M in/out; cached $0.50. Doubles at ≥200K prompt.
  • Blended (7:2:1 cache/in/out): $1.35 / 1M.
  • Hosted decode ~58 tok/s; TTFT ~43s at high reasoning.
  • 72M output tokens to run the AA Intelligence Index — relatively thrifty.
  • Default in Grok Build. Proprietary. No local weights.

GLM-5.3-Flash · Z.ai

AA Index 57 $0.09 / AA task 1M ctx
  • 320B MoE / 18B active, MIT, multimodal. Was Ox Alpha on OpenRouter.
  • List: $0.15 / $0.50; cached $0.03. Promo through 9 Sep 2026: $0.075 / $0.25 (OpenRouter matches promo).
  • Blended list ~$0.10 / 1M. Hosted decode ~50 tok/s; TTFT ~1.5s.
  • 150M output tokens on the same Index — about 2.1× more verbose than Grok 4.6 high.
  • Thinking on by default (low / high / max). Tool parser glm47.

AA cost per Intelligence Index task

GLM-5.3-Flash
$0.09
Grok 4.6 (high)
$0.94

Artificial Analysis, captured 28 Aug 2026. Flash ~10.4× cheaper per Index task. VentureBeat’s same-day chart used ~$0.09 vs ~$0.94 and called it “about 10× for a four-point gain.”

3. How big is the downgrade?

Short version: small on coding/agent benches, real on token thrift and product glue, not a cliff.

Axis Grok 4.6 (high) GLM-5.3-Flash Read as
AA Intelligence Index 61 57 Four points. Not a generation.
AA $/Index task $0.94 $0.09 ~10×
Index output tokens 72M 150M Flash talks ~2.1× more
Hosted tok/s 57.8 49.8 Similar API speed
TTFT (high / max think) 43s 1.5s Flash feels snappier at the socket
Context 500K 1M Flash wins on window
Weights Closed MIT, 320B/18B Local is the whole point

AA model pages, 28 Aug 2026. Coding sub-scores on AA are charts, not a single published Flash-vs-4.6 Terminal-Bench cell I could copy without the interactive UI. Same-harness Rails/OpenCode numbers below are for GLM-5.3 (flagship), not Flash — treat as an upper bound on the family.

What the same-harness coding runs actually said

“GLM 5.3 re-runs the test suite about 20 times per run, where most models run it 3 to 6 times. It takes its time and spends your tokens.” Agents on Rails · 17 Aug 2026 · about GLM-5.3 flagship, not Flash

Where you will feel it in an agent loop

Grok 4.6 still wins

  • Ceiling on messy, multi-hour debugging.
  • Token thrift — fewer output tokens per finished task.
  • Grok Build as a product: tools, compaction, cache routing, Imagine, X search.
  • Less “finish early / needs extra review” than the cheaper open-weight class, in the July r/cursor pattern for 4.5. 4.6 is the same family, smarter.

Flash is good enough when

  • The loop is mechanical: implement, test, fix, repeat.
  • You already have tests. Flash will burn tokens on them; that is cheaper than Grok doing the same.
  • You want 1M context or local data that should not leave the box.
  • The weekly meter is the actual bottleneck, not model IQ.

This is the same emerging stack as the Grok 4.5 vs Opus 5 brief, with the executor slot swapped: plan/review on Grok 4.6 (or Opus), execute volume on Flash. If you drop Grok entirely, expect more human review, not a collapsed product.

4. Quant on 4× RTX PRO 6000

Assumption, stated so it can be wrong: 4× RTX PRO 6000 Blackwell, 96 GB each, 384 GB total, SM120. That is the card every current GLM-5.3-Flash serving recipe targets. If these are Ada 48 GB cards (192 GB total), skip to the callout — NVFP4 does not have a comfortable home there.

Build Size Fits 384 GB? Quality / speed note
BF16 ~642–650 GB No Don’t.
Native FP8 (official HF) ~306 GiB Tight ~78 GB left for KV + runtime. Few long agents.
NVFP4 (LibertAIDAI) ~182 GiB Yes, headroom SM120 W4A4. 0xSero 4× recipe. Pin this.
EXL3 4bpw (brandonmusic) ~165 GiB Yes Fits 2× 96 GB. Use for two replicas.
Unsloth UD-Q4_K_XL 200 GB Yes 93% of BF16 top-1% accuracy. llama.cpp, not the fast SGLang path.
UD-IQ3_XXS 120 GB Easy 82% accuracy. For 128 GB boxes, not this one.
UD-IQ1_S 93 GB Easy 71% accuracy. A demo, not an agent.

Sizes: Unsloth GLM-5.3-Flash docs (28 Aug 2026); 0xSero NVFP4 182 GiB lock; tpurtell EXL3 “roughly 165 GiB.” Accuracy column is Unsloth’s top-1% retention vs BF16, not a coding-agent eval.

Run NVFP4. It is the quantization that keeps both speed and user experience on this silicon. Unsloth’s 4-bit dynamic GGUF is the quality floor you should refuse to go under (93%). NVFP4 is 4-bit weights consumed as W4A4 on Blackwell Tensor Cores, with FlashInfer CUTLASS MoE that actually has an SM120 module. CuteDSL MoE is compiled for datacenter Blackwell (sm100/103) and dies on SM120. That is why the 4× recipe pins flashinfer_cutlass.

Do not run Q3/Q2 on a 384 GB box to “go faster.” 0xBakeer’s Qwen 3.8 work (different model, same lesson) showed the 4-bit advantage vanishing under concurrency: +27% at c=1, +0.2% at c=16. Once a batch is sharing a weight read, extra quant is mostly quality you threw away.

If the cards are Ada 48 GB (4×48 = 192 GB): NVFP4 at 182 GiB leaves almost nothing for KV. EXL3 4bpw is theoretically loadable and miserable. Practical path is Unsloth UD-Q3 / IQ3 with llama.cpp or a 2-bit and a lot of disappointment. This brief’s concurrency and $10k math assume 96 GB Blackwell.

5. How many concurrent agents

Two measured recipes, not a vibe:

4× RTX PRO 6000 · NVFP4 · SGLang

0xSero lock, TP/EP 4/4, MTP5, CUDA graphs for batch ≤ 8.

  • C1 decode: ~143 tok/s low effort, ~208 tok/s default (NEXTN).
  • KV: fp8_e4m3, 3.79M tokens/rank (~28 GB).
  • Context advertised: 1,048,576.

2× RTX PRO 6000 · EXL3 4bpw · vLLM

tpurtell / brandonmusic, MAX_NUM_SEQS=16, adaptive MTP.

  • C1 decode: 107 adaptive / 98 MTP5 / 71 off.
  • C16 aggregate: 336 adaptive / 277 MTP5 / 543 off.
  • KV pool: 525K tokens at 0.95 (500K model limit).

What those numbers mean for agents, not benches

Mode Live agents Why that number UX
Interactive (pin this) 8 CUDA-graph max on the 4× NVFP4 recipe. Per-stream still fast. Good. Chat/agent feels like a hosted API.
Throughput 16 Qualified C16 on 2× EXL3. 4× NVFP4 has more KV; graph may go eager above 8. Fine for overnight loops. Per-stream drops (2× box: ~21 tok/s/stream adaptive at C16).
KV-limited 32K ~50–100 3.79M pool / 32–64K. Scheduler and decode, not VRAM, stop you first. Bad as interactive. Only a batch farm.
Long-context 200K ~8–15 Same pool / 200K. Plus activation headroom. Repo-in-context agents. Stay at 8.
500K–1M dumps 1–3 DGX Spark FP8 lanes were 15@200K / 5@500K / 3@1M on weaker memory. Not a farm. A document job.

User-experience number: 8 concurrent agents on one NVFP4 replica. That is the batch size the 4× recipe actually captured graphs for. Above 8 you trade latency for aggregate tokens. For looping coding agents overnight, 16 is the honest throughput cap I would set before measuring.

Max-concurrency play if the goal is “peg the box”: two EXL3 replicas, TP2 each (the 165 GiB quant fits on 2×96 GB). Then you have 8+8 interactive or 16+16 scheduler slots, and a dead replica does not take down the farm. 0xSero’s single TP4 replica is simpler and faster per stream (208 vs 107 tok/s C1). Pick TP4 if you want one fat endpoint; pick 2× TP2 if you want agent count.

MAX_NUM_SEQS=16 is scheduler concurrency, not a promise that sixteen 500K prompts fit. The 2× recipe’s measured pool is 525K logical tokens — a bit more than one 500K request. Agent sessions at 32–128K are a different sport, and that is the sport you actually play.

6. OpenRouter vs Grok vs a $300 subscription

Same Bishop mix (9.66M in / 306M cache / 1.23M out). This is one week of Grok Build that filled a $30 SuperGrok meter. Reprice it three ways.

Route That week vs Grok API Notes
Grok 4.6 API $179.56 1.0× $2 / $0.50 cache / $6
SuperGrok $30 (included) $7.50 included $30/mo ÷ 4 weeks. 24× under API if you max it.
GLM Flash OpenRouter promo $5.62 32× cheaper $0.075 / $0.015 / $0.25 through 9 Sep 2026
GLM Flash list (budget this) $11.24 16× cheaper $0.15 / $0.03 / $0.50 after promo
Flash list, 2.1× verbosity ~$15 12× cheaper Scale output 2.1× and loops 1.3× from AA verbosity
Local NVFP4 ~$6–10 power only See §8. One week of 24/7 is more than this mix.

Promo matches OpenRouter’s z-ai/glm-5.3-flash as of 26–28 Aug 2026. Budget against list. Cache is the column that matters in agent loops; Flash list cache is $0.03 vs Grok 4.6’s $0.50 (17×).

What a $300 SuperGrok Heavy month is worth

xAI will not say. Two honest brackets:

You will not hit 100% every week on Heavy without turning Grok Build into a job. Bishop had to work at it on the $30 plan. Heavy’s 16-agent mode also burns the pool faster per wall-clock hour. Treat $5k–$8k of Grok 4.6 API per fully used Heavy month as the range, not $10k.

Task counts, using Bishop’s 31 sessions / $30 week as “a SuperGrok-week of looping agent work”:

$30 SuperGrok

~31

sessions / week if you fill the meter

$300 Heavy (6–10×)

~190–310

sessions / week, if the pool even lets you

AA tasks @ $0.94

~5.3k–8.2k

Index tasks per maxed Heavy month

Same on Flash API

~$290–$480

Bishop-mix list for that Heavy month. AA Index tasks: $460–$770. Sheet in §7.

So: OpenRouter Flash does the same Bishop-shaped week for about $11 list instead of $180 of Grok API or a slice of a $30–$300 subscription. The catch is quality and verbosity, not the invoice.

Unit price vs the cap. SuperGrok / Heavy is a better deal per Grok-dollar than Flash list if you fit in the weekly pool — $30 buys ~$773 of API if you max four weeks ($0.039 per Grok-dollar vs Flash list $0.063). Promo Flash ($0.031) undercuts that until 9 Sep. The reason to leave is not list price. It is that two looping projects already ate 30% of a week, and the 4× box / OpenRouter have no weekly meter.

7. The sheet — 100 tasks, $300, $10k, $61k

One ruler, three invoices. The mix is Bishop’s SuperGrok week (9.66M in / 306M cache / 1.23M out) priced at Grok 4.6 API, then repriced at OpenRouter GLM-5.3-Flash. “Same tasks” inflates Flash output 2.1× and loops 1.3× from the AA verbosity gap, so you are comparing finished work, not token count.

A “Bishop-session” is one of those 31 Grok Build sessions (~40k output tokens with that week’s cache ratio). The original “100 tasks = a weekly Grok limit” hypothetical is in the table even though the measured SuperGrok week was 31, not 100.

What $300 of Flash credits buy, vs emptying Heavy

$300 on Flash list

~$4.8k

Grok 4.6 API-eq. ~830 Bishop-sessions. No weekly cap.

$300 on Flash promo

~$9.6k

Through 9 Sep 2026. About a maxed Heavy month of API value.

Empty a $300 Heavy month

$5–8k

Grok API-eq. Flash list invoice for that mix: $290–$480.

100 Bishop-sessions

$36

Flash list. Same 100 on Grok 4.6 API: $579. SuperGrok week is 31, not 100.

Workload Grok 4.6 API Flash list Flash promo Flash, same tasks How to read it
1 Bishop-session $5.79 $0.36 $0.18 $0.49 1/31 of a measured SuperGrok week
SuperGrok week (31 sessions) $180 $11.24 $5.62 $15 Filled a $30 meter. Included cost: $7.50/week
100 Bishop-sessions $579 $36 $18 $49 The “100 tasks = a weekly limit” hypothetical. 3.2 SuperGrok weeks
Maxed SuperGrok month ($30) $773 $48 $24 $65 4.3 weeks × $180. Subscription still $30
$300 Heavy month, 6× burn $4,600 $288 $144 $387 If Heavy is ~6× SuperGrok and you hit 100% every week
$300 Heavy month, 10× price-ratio $7,700 $482 $241 $647 If Heavy is 10× SuperGrok. Still under $10k of API
Spend $300 on Flash list $4,794 eq. $300 $3,568 eq. ~830 sessions. About the 6× Heavy month, uncapped
Spend $300 on Flash promo $9,586 eq. $300 $7,136 eq. Promo only through 9 Sep 2026. Budget list after that
$10k of Grok-eq (16% of 4× farm) $10,000 $626 $313 $841 Workday looping, or ~4 pegged hours/day at C8
Pegged 4× farm, 24/7 C8 $61,410 $3,843 $1,921 $5,164 ~8–13 emptied Heavy months of API value in one calendar month
2× EXL3 replicas, 24/7 $80k–$100k $5.0k–$6.3k $2.5k–$3.1k $6.7k–$8.4k Max concurrency. Worse tok/s per stream. Same four cards

Bishop mix at list: 9.658M × $0.15 + 305.717M × $0.03 + 1.231M × $0.50 = $11.24, vs Grok $179.56 (16.0×). Promo is half. “Same tasks” is $15.10 for that week (11.9×). $10k and $61k rows reprice the Grok token mix, not equal merged PRs. Local power for the pegged farm is $237–$394/month (§8) — that is the $61k row’s real bill, not $3,843.

Before you cancel Heavy

The $60k / $10k question, invoiced on OpenRouter. $10,000 of Grok 4.6 API-equivalent Bishop mix costs $626 on GLM-5.3-Flash list ($313 on the current promo). The 4× RTX PRO 6000 pegged farm’s ~$61k/month of Grok-eq costs $3,843 on Flash list ($1,921 promo) — or $237–$394 of electricity if you run it locally. A $300 SuperGrok Heavy subscription, emptied, is not $10k of API; it is $5k–$8k, and that mix is $290–$480 on Flash list.

8. Can the 4× box print $10,000 of Grok-equivalent?

Yes. $10k is not the ceiling. It is a light-duty month.

Pegged-farm math (one NVFP4 replica)

Conservative decode aggregate at C8 on 4×: 400 tok/s (the 2× EXL3 box already did 336 aggregate at C16 adaptive; 0xSero’s 4× C1 is 208). Agent loops are not 100% decode — tools, prefill, and sampler waits eat the rest. Use 40% decode duty for a farm that is actually looping without a human in the chair:

400 tok/s × 0.40 × 86,400 s/day × 30.44 days = 421 million output tokens / month

Bishop’s mix costs $145.87 of Grok 4.6 API per million output tokens (because each million of output rides ~7.9M input and ~248M cache). Scale that:

24/7 C8 farm

~$61k

Grok 4.6 API equivalent / month

To hit $10k

16%

of that farm, or ~4 h/day pegged

Bishop-sessions

~10.6k

421M out ÷ 39.7k out/session

vs SuperGrok

~85×

maxed $30 months packed into one calendar month

Utilization Grok 4.6 API-eq Flash list Flash promo Bishop-sessions What it looks like
Workday 8h, 4 agents, 25% duty ~$8k $501 $250 ~1.4k Just-you, looping while you work. Near $10k already.
16% of 24/7 C8 farm $10k $626 $313 ~1.7k The question. Not a stretch. OpenRouter invoice: six hundred dollars, not ten thousand.
Workday 8h, 8 agents, 40% duty ~$20k $1,252 $626 ~3.5k Serious personal farm, nights off.
24/7 C8, 40% duty ~$61k $3,843 $1,921 ~10.6k Pegged. OpenRouter is ~$3.8k list. Local bill is power (~$240–$390).
2× EXL3 replicas, 24/7 ~$80k–$100k $5.0k–$6.3k $2.5k–$3.1k more, slower each Max concurrency. Worse tok/s/stream.

“Bishop-session” = 1.23M output / 31 ≈ 39.7k output tokens, with that week’s cache/input ratio. Your loops will differ. If Flash is 2.1× more verbose, you finish fewer tasks per GPU-second, not fewer Grok-equivalent dollars — the Grok-eq column is priced at Grok rates for the same token mix. Flash list is 16.0× cheaper on that mix; “same tasks” is 11.9× (§7). Locally the tokens cost power.

Against the $300 subscription

A fully used Heavy month is maybe $5k–$8k of Grok 4.6 API. The 4× box at 24/7 is ~8–12 Heavy months of that value in one calendar month, or ~200 SuperGrok ($30) months. Electricity for 4× 600W + ~300W host ≈ 2.7 kW:

2.7 kW × 24 × 30.44 × $0.12/kWh ≈ $237/month  ·  at $0.20/kWh ≈ $394/month

So the local “subscription” is power plus amortization, not $300. $10k of Grok-equivalent tokens on the box costs a few hundred dollars of electricity, not $300 of xAI. Buying that same $10k of Grok-eq on OpenRouter Flash is $626 list / $313 promo — cheap vs Grok API, still ~10–16× the electricity bill. That is the whole argument for leaving the meter, and for not pretending OpenRouter is as cheap as owning the GPUs.

Could you generate $10,000 of Grok-equivalent usage? Yes. A pegged 4× NVFP4 farm does ~$61k/month of the Bishop mix at Grok 4.6 list ($3,843 on Flash list, $237–$394 of power). $10k is roughly four pegged hours per day at C8, or a full workday of 4–8 looping agents ($626 on Flash list). A $300 SuperGrok Heavy subscription, even if you empty it every week, does not get you to $10k of API-equivalent — the local box does, and OpenRouter Flash does it for hundreds, not thousands, of dollars.

9. M5 Ultra 256GB — 1×, 4×, 6× vs the 4× 6000

Same Bishop ruler. Same 40% decode duty. Different silicon: Mac Studio M5 Ultra, 256 GB unified memory, 1.2 TB/s (Apple, 25 Aug 2026). 256GB is the SKU you can order today; 512GB follows in late October. Flash at 4-bit is a 200GB Unsloth GGUF / ~178GB MLX 4-bit. It fits. KV is the squeeze.

What actually fits on 256GB

Build Size Fits 256 GB? Agent note
MLX 8-bit 334 GB No 512GB SKU. Not this one.
MLX 6-bit 256 GB No (weights = the box) Zero KV. Skip.
Unsloth UD-Q4_K_XL 200 GB Yes, tight 93% of BF16 top-1. 1–2 interactive agents. Pin this for quality.
MLX mixed 4/8-bit 182 GB Yes Better PPL than uniform 4-bit. More KV than UD-Q4.
UD-Q3_K_XL / IQ3_XXS 148 / 120 GB Yes, headroom 4 concurrent on one box. 82–87% accuracy. Use when you need slots, not when you have four Macs.
UD-Q2_K_XL 109 GB Easy 78% accuracy. Spark trick, not a 256GB trick.

Unsloth GGUF sizes 28 Aug 2026; PipeNetwork MLX PPL table (8-bit 334.1 GB / mixed-4/8 181.9 GB / 4-bit 177.6 GB). Unsloth’s 4-bit “162–210 GB total memory” band is why Q4 on 256GB is a 1–2 agent machine, not an 8-agent machine.

Decode: measured M3 Ultra, scaled to 1.2 TB/s

Nobody has published a warmed GLM-5.3-Flash decode curve on M5 Ultra yet — machines start arriving 22 Sep 2026. Two M3 Ultra 512GB numbers exist: omlx Q4e 24.1 tok/s (4k prompt), antirez ds4 Q4 35.5 tok/s short / 26.6 tok/s at ~12k. M5 Ultra bandwidth is 1.2 TB/s vs M3 Ultra 819 GB/s = 1.46×. Neural Accelerators (Apple’s 4.3× peak AI compute vs M3 Ultra) help prefill, not decode. Decode stays bandwidth-bound.

Agent C1 pin: 45 tok/s  ·  26.6 × 1.46 ≈ 39, plus a little Metal/MTP. High-concurrency aggregate per box: 70 tok/s (1.55× C1). NVIDIA SGLang got 1.9× C1→C8; Mac saturates the bus sooner.

Do not tensor-parallel Flash across Macs. Flash already fits on one 256GB box. Exo RDMA / Apple’s “4 Macs = 3× one system” is the trick for models that don’t fit (Kimi K3, GLM-5.3 753B). Splitting Flash with TP buys one faster stream and kills agent count. For looping agents: one replica per Mac. All-to-all TB5 maxes at 7 Macs (6 ports); 4 nodes need 6 cables / 3 ports each; 6 nodes need 15 cables / 5 ports each — it physically fits.

Concurrency that still feels fast

Box Interactive Overnight Per-stream at interactive
1× M5 Ultra 256GB, Q4 2 4 ~22–35 tok/s. KV is the wall, not the scheduler.
1× M5 Ultra 256GB, Q3 4 8 More slots, dumber weights. Fine for mechanical loops.
4× Ultra, replica Q4 8 16 Same per-stream as one box. Matches the 4× 6000 count.
6× Ultra, replica Q4 12 24 More agents than the 4× 6000. Each stream still ~25–35 tok/s.
4× RTX PRO 6000 NVFP4 8 16 ~50 tok/s/stream at C8 (400 aggregate). Faster each, same count as 4× Ultra.
4× Ultra, TP (wrong play) 1–2 2–3 Apple 3× one system ≈ 135 tok/s one stream. Fat Kimi, not Flash agents.

Interactive pin on one 256GB Q4 replica: 2 agents. That keeps each stream in the 20–35 tok/s band — slower than hosted Flash (~50 tok/s) and slower than the 4× 6000 at C8 (~50 tok/s/stream), still usable. Four agents on one Q4 box drops you under ~18 tok/s/stream. That is overnight, not chat.

High-duty Grok-equivalent — the farm sheet

Same formula as §8: aggregate decode tok/s × 40% duty × 30.44 days, priced at $145.87 of Grok 4.6 API per million output tokens of the Bishop mix. That is $153.50 of Grok-eq per tok/s of aggregate decode. 1× high-concurrency uses 70 tok/s; replica clusters multiply the boxes.

Farm (high-duty looping) Agg. tok/s Grok 4.6 API-eq vs $300 Heavy Flash list Power / mo
Empty SuperGrok Heavy month $5k–$8k 1.0× $290–$480 $0
$300 of OpenRouter Flash list $4,794 eq. ~0.6–1.0× $300 $0
1× M5 Ultra 256GB 70 $10.7k 1.4–2.3× $672 $16
1× Ultra, C1 only (no batching) 45 $6.9k 0.9–1.5× $432 $16
4× 6000, 16% duty ($10k question) 64 $10k 1.3–2.2× $626 ~$38
4× Ultra, replica 280 $43k 5.6–9.3× $2,690 $63
4× RTX PRO 6000 NVFP4 400 $61k 8.0–13.3× $3,843 $237
6× Ultra, replica 420 $64k 8.4–14.0× $4,035 $95

Heavy column is emptied-month multiples using the $4.6k–$7.7k API-eq bracket from §6. Power is 180 W/Studio under inference (M5 Ultra community ~180–214 W peak; Apple’s last Ultra max was 270 W) at $0.12/kWh, vs 2.7 kW for the 4× 6000. Promo Flash is half the list column. 1× Ultra C1-only is the honest floor if Metal batching does not show up.

1× Ultra vs Heavy

~2×

One 256GB Studio at high duty ≈ two emptied $300 months, every calendar month

6× Ultra vs 4× 6000

~$64k / $61k

Same Grok-eq band. Mac uses ~40% of the watts and more agent slots.

1× Ultra on OpenRouter

$672

Flash list for that $10.7k mix. Promo $336. Power $16.

4× Ultra on OpenRouter

$2.7k

vs $3.8k to match the 4× 6000 farm on Flash list

The X chart — $300 vs one Ultra vs 2×/4× 6000

Same ruler, six rows, the version made to post. Token-max Heavy and $300 of Flash buy about the same Grok-eq. $300 of flagship GLM-5.3 does not — it is $1.40/$4.40, 1.8× under Grok, not 16×. Pegged silicon is a different category: the box costs about one month of the Grok 4.6 API bill it replaces, then you pay watts.

PNG for X (1600×900, 2×)

Farm Agents Grok 4.6 to match vs $300 Heavy Flash list You pay
Token-max SuperGrok Heavy $5–8k 1.0× $290–480 $300
$300 GLM-5.3 Flash · OpenRouter $4.8k ~1× $300 $300
$300 GLM-5.3 · OpenRouter $550 0.07–0.11× $34 $300
1× M5 Ultra 256GB 2 $10.7k 1.4–2.3× $672 $16
2× RTX PRO 6000 4–8 $52k 6.7–11× $3.2k $127
4× RTX PRO 6000 8 $61k 8–13× $3.8k $237

$300 Flash = Bishop mix at list ($11.24/week → $4,794 Grok-eq). $300 GLM-5.3 flagship at $1.40/$0.26/$4.40 = $98.42/week → $547 Grok-eq. 2× 6000 is EXL3 C16 measured 336 tok/s × $153.46/tok/s-month. Power at $0.12/kWh. Current RTX PRO 6000 street ~$14–16k/card: 2× ≈ $30–37k vs $52k of Grok-eq/month; 4× ≈ $60–72k vs $61k; Ultra 256GB ~$10k vs $10.7k. One pegged month ≈ the hardware sticker in Grok 4.6 API dollars.

How to read the three invoices

Per-agent speed still favors the 4× 6000. NVFP4 + CUDA graphs at C8 is ~50 tok/s/stream. One Ultra replica at 2-wide is ~25–35 tok/s/stream. If the test is “eight agents that feel like the API,” buy Blackwell or pay OpenRouter. If the test is “quiet desk, no weekly meter, enough Grok-eq to embarrass Heavy,” one 256GB Ultra already does it. Cluster only when you want more slots, not when you want Flash to go faster — TP Flash is the wrong topology.

10. What I would actually run

  1. Serve LibertAIDAI/GLM-5.3-Flash-NVFP4 with the 0xSero SGLang SM120 image, TP4, flashinfer_cutlass, MTP5, FP8 KV, graphs ≤ 8. Parsers: reasoning glm45, tools glm47.
  2. Cap interactive agents at 8. Overnight loops at 16. Do not pass tiny max_tokens — max thinking will burn the budget inside <think> and return empty content.
  3. Effort: low for mechanical edits, high for implement-and-test, max only when a Grok-class problem would have justified Grok. Default is max; that will inflate verbosity toward the AA 2.1× figure.
  4. OpenRouter Flash as the overflow while the local image is compiling, and after 9 Sep 2026 budget list $0.15/$0.50 not the promo. Plan ~$626/month if you are going to do $10k of Grok-eq on the API, ~$3,840/month if you want the pegged-farm volume without the GPUs yet.
  5. Keep one Grok 4.6 (or Opus) seat for plan, review, and the 10% of tasks where four AA points and token thrift matter. Spend the weekly pool on that, not on loops.
  6. On M5 Ultra 256GB: Unsloth UD-Q4_K_XL or MLX mixed-4/8, llama.cpp / Unsloth Desktop / MLX. Cap interactive agents at 2 per box. Cluster with replicas (exo or just four API ports), not tensor-parallel Flash. Q3 only if you insist on 4 agents on one 256GB machine.

If the goal is maximum concurrent loops rather than nicest single-agent UX: two EXL3 TP2 replicas instead of one NVFP4 TP4 on the 4× 6000, or N Ultra replicas. Same idea either side — more endpoints, slower each stream.

11. Caveats

12. Sources