Field brief · August 28, 2026
Two looping projects ate 30% of a weekly Grok Build pool. That is the product working as designed — agent loops are compute-weighted, not “messages.” This brief sizes the downgrade to GLM-5.3-Flash, the quantization to run on a 4× RTX PRO 6000 box and on M5 Ultra 256GB (one box, four, and six), how many concurrent agents actually feel good, whether pegging those farms can print $10,000 of Grok 4.6 API-equivalent work in a month, and what that same volume invoices on OpenRouter Flash — plus what $300 of Flash credits actually buy.
The intelligence gap is four Artificial Analysis Index points (Grok 4.6 high = 61, Flash = 57). The money gap is ~10× per finished eval task. For looping coding agents, Flash is a mild quality cut and a large capacity unlock.
Grok subscriptions since June 2026 share one compute-weighted weekly pool across Chat, Imagine, Voice, and Build. xAI does not publish a token number. The Usage page shows a percentage. Agent and Build tasks draw more than chat by design — Grok said so on X when Heavy users burned 25–30% in a day of backend work.
That matches the lived experience: two projects, a couple of looping tasks, 30% gone. You are not “light.” Looping agents re-send growing context every turn. The meter is weighting that, not counting prompts.
The only public, instrumented SuperGrok week I trust is Ben Bishop’s $30 plan hitting 100% after 31 Grok Build sessions (Mission Control extract, week of 10–17 Aug 2026):
Sessions
31
one $30 SuperGrok week to 100%
Input
9.66M
fresh prompt tokens
Output
1.23M
completions
Cache read
306M
reused agent prefix
Price that mix at Grok 4.6 API rates ($2 / $0.50 cached / $6 per 1M, prompts <200K):
9.658M × $2 + 305.717M × $0.50 + 1.231M × $6 = $179.56
So 100% of a $30 SuperGrok week of real Build work is about $180 of Grok 4.6 API. Thirty percent of that week is about $54 of API. Bishop is explicit that this is not “the official quota in tokens” — it is the workload recorded while the meter hit 100%. It is still the best ruler we have.
glm47.Artificial Analysis, captured 28 Aug 2026. Flash ~10.4× cheaper per Index task. VentureBeat’s same-day chart used ~$0.09 vs ~$0.94 and called it “about 10× for a four-point gain.”
Short version: small on coding/agent benches, real on token thrift and product glue, not a cliff.
| Axis | Grok 4.6 (high) | GLM-5.3-Flash | Read as |
|---|---|---|---|
| AA Intelligence Index | 61 | 57 | Four points. Not a generation. |
| AA $/Index task | $0.94 | $0.09 | ~10× |
| Index output tokens | 72M | 150M | Flash talks ~2.1× more |
| Hosted tok/s | 57.8 | 49.8 | Similar API speed |
| TTFT (high / max think) | 43s | 1.5s | Flash feels snappier at the socket |
| Context | 500K | 1M | Flash wins on window |
| Weights | Closed | MIT, 320B/18B | Local is the whole point |
AA model pages, 28 Aug 2026. Coding sub-scores on AA are charts, not a single published Flash-vs-4.6 Terminal-Bench cell I could copy without the interactive UI. Same-harness Rails/OpenCode numbers below are for GLM-5.3 (flagship), not Flash — treat as an upper bound on the family.
This is the same emerging stack as the Grok 4.5 vs Opus 5 brief, with the executor slot swapped: plan/review on Grok 4.6 (or Opus), execute volume on Flash. If you drop Grok entirely, expect more human review, not a collapsed product.
Assumption, stated so it can be wrong: 4× RTX PRO 6000 Blackwell, 96 GB each, 384 GB total, SM120. That is the card every current GLM-5.3-Flash serving recipe targets. If these are Ada 48 GB cards (192 GB total), skip to the callout — NVFP4 does not have a comfortable home there.
| Build | Size | Fits 384 GB? | Quality / speed note |
|---|---|---|---|
| BF16 | ~642–650 GB | No | Don’t. |
| Native FP8 (official HF) | ~306 GiB | Tight | ~78 GB left for KV + runtime. Few long agents. |
| NVFP4 (LibertAIDAI) | ~182 GiB | Yes, headroom | SM120 W4A4. 0xSero 4× recipe. Pin this. |
| EXL3 4bpw (brandonmusic) | ~165 GiB | Yes | Fits 2× 96 GB. Use for two replicas. |
| Unsloth UD-Q4_K_XL | 200 GB | Yes | 93% of BF16 top-1% accuracy. llama.cpp, not the fast SGLang path. |
| UD-IQ3_XXS | 120 GB | Easy | 82% accuracy. For 128 GB boxes, not this one. |
| UD-IQ1_S | 93 GB | Easy | 71% accuracy. A demo, not an agent. |
Sizes: Unsloth GLM-5.3-Flash docs (28 Aug 2026); 0xSero NVFP4 182 GiB lock; tpurtell EXL3 “roughly 165 GiB.” Accuracy column is Unsloth’s top-1% retention vs BF16, not a coding-agent eval.
Run NVFP4. It is the quantization that keeps both speed and user experience on this silicon. Unsloth’s 4-bit dynamic GGUF is the quality floor you should refuse to go under (93%). NVFP4 is 4-bit weights consumed as W4A4 on Blackwell Tensor Cores, with FlashInfer CUTLASS MoE that actually has an SM120 module. CuteDSL MoE is compiled for datacenter Blackwell (sm100/103) and dies on SM120. That is why the 4× recipe pins flashinfer_cutlass.
Do not run Q3/Q2 on a 384 GB box to “go faster.” 0xBakeer’s Qwen 3.8 work (different model, same lesson) showed the 4-bit advantage vanishing under concurrency: +27% at c=1, +0.2% at c=16. Once a batch is sharing a weight read, extra quant is mostly quality you threw away.
Two measured recipes, not a vibe:
0xSero lock, TP/EP 4/4, MTP5, CUDA graphs for batch ≤ 8.
tpurtell / brandonmusic, MAX_NUM_SEQS=16, adaptive MTP.
| Mode | Live agents | Why that number | UX |
|---|---|---|---|
| Interactive (pin this) | 8 | CUDA-graph max on the 4× NVFP4 recipe. Per-stream still fast. | Good. Chat/agent feels like a hosted API. |
| Throughput | 16 | Qualified C16 on 2× EXL3. 4× NVFP4 has more KV; graph may go eager above 8. | Fine for overnight loops. Per-stream drops (2× box: ~21 tok/s/stream adaptive at C16). |
| KV-limited 32K | ~50–100 | 3.79M pool / 32–64K. Scheduler and decode, not VRAM, stop you first. | Bad as interactive. Only a batch farm. |
| Long-context 200K | ~8–15 | Same pool / 200K. Plus activation headroom. | Repo-in-context agents. Stay at 8. |
| 500K–1M dumps | 1–3 | DGX Spark FP8 lanes were 15@200K / 5@500K / 3@1M on weaker memory. | Not a farm. A document job. |
User-experience number: 8 concurrent agents on one NVFP4 replica. That is the batch size the 4× recipe actually captured graphs for. Above 8 you trade latency for aggregate tokens. For looping coding agents overnight, 16 is the honest throughput cap I would set before measuring.
Max-concurrency play if the goal is “peg the box”: two EXL3 replicas, TP2 each (the 165 GiB quant fits on 2×96 GB). Then you have 8+8 interactive or 16+16 scheduler slots, and a dead replica does not take down the farm. 0xSero’s single TP4 replica is simpler and faster per stream (208 vs 107 tok/s C1). Pick TP4 if you want one fat endpoint; pick 2× TP2 if you want agent count.
MAX_NUM_SEQS=16 is scheduler concurrency, not a promise that sixteen 500K prompts fit. The 2× recipe’s measured pool is 525K logical tokens — a bit more than one 500K request. Agent sessions at 32–128K are a different sport, and that is the sport you actually play.
Same Bishop mix (9.66M in / 306M cache / 1.23M out). This is one week of Grok Build that filled a $30 SuperGrok meter. Reprice it three ways.
| Route | That week | vs Grok API | Notes |
|---|---|---|---|
| Grok 4.6 API | $179.56 | 1.0× | $2 / $0.50 cache / $6 |
| SuperGrok $30 (included) | $7.50 | included | $30/mo ÷ 4 weeks. 24× under API if you max it. |
| GLM Flash OpenRouter promo | $5.62 | 32× cheaper | $0.075 / $0.015 / $0.25 through 9 Sep 2026 |
| GLM Flash list (budget this) | $11.24 | 16× cheaper | $0.15 / $0.03 / $0.50 after promo |
| Flash list, 2.1× verbosity | ~$15 | 12× cheaper | Scale output 2.1× and loops 1.3× from AA verbosity |
| Local NVFP4 | ~$6–10 | power only | See §8. One week of 24/7 is more than this mix. |
Promo matches OpenRouter’s z-ai/glm-5.3-flash as of 26–28 Aug 2026. Budget against list. Cache is the column that matters in agent loops; Flash list cache is $0.03 vs Grok 4.6’s $0.50 (17×).
xAI will not say. Two honest brackets:
You will not hit 100% every week on Heavy without turning Grok Build into a job. Bishop had to work at it on the $30 plan. Heavy’s 16-agent mode also burns the pool faster per wall-clock hour. Treat $5k–$8k of Grok 4.6 API per fully used Heavy month as the range, not $10k.
Task counts, using Bishop’s 31 sessions / $30 week as “a SuperGrok-week of looping agent work”:
$30 SuperGrok
~31
sessions / week if you fill the meter
$300 Heavy (6–10×)
~190–310
sessions / week, if the pool even lets you
AA tasks @ $0.94
~5.3k–8.2k
Index tasks per maxed Heavy month
Same on Flash API
~$290–$480
Bishop-mix list for that Heavy month. AA Index tasks: $460–$770. Sheet in §7.
So: OpenRouter Flash does the same Bishop-shaped week for about $11 list instead of $180 of Grok API or a slice of a $30–$300 subscription. The catch is quality and verbosity, not the invoice.
One ruler, three invoices. The mix is Bishop’s SuperGrok week (9.66M in / 306M cache / 1.23M out) priced at Grok 4.6 API, then repriced at OpenRouter GLM-5.3-Flash. “Same tasks” inflates Flash output 2.1× and loops 1.3× from the AA verbosity gap, so you are comparing finished work, not token count.
A “Bishop-session” is one of those 31 Grok Build sessions (~40k output tokens with that week’s cache ratio). The original “100 tasks = a weekly Grok limit” hypothetical is in the table even though the measured SuperGrok week was 31, not 100.
$300 on Flash list
~$4.8k
Grok 4.6 API-eq. ~830 Bishop-sessions. No weekly cap.
$300 on Flash promo
~$9.6k
Through 9 Sep 2026. About a maxed Heavy month of API value.
Empty a $300 Heavy month
$5–8k
Grok API-eq. Flash list invoice for that mix: $290–$480.
100 Bishop-sessions
$36
Flash list. Same 100 on Grok 4.6 API: $579. SuperGrok week is 31, not 100.
| Workload | Grok 4.6 API | Flash list | Flash promo | Flash, same tasks | How to read it |
|---|---|---|---|---|---|
| 1 Bishop-session | $5.79 | $0.36 | $0.18 | $0.49 | 1/31 of a measured SuperGrok week |
| SuperGrok week (31 sessions) | $180 | $11.24 | $5.62 | $15 | Filled a $30 meter. Included cost: $7.50/week |
| 100 Bishop-sessions | $579 | $36 | $18 | $49 | The “100 tasks = a weekly limit” hypothetical. 3.2 SuperGrok weeks |
| Maxed SuperGrok month ($30) | $773 | $48 | $24 | $65 | 4.3 weeks × $180. Subscription still $30 |
| $300 Heavy month, 6× burn | $4,600 | $288 | $144 | $387 | If Heavy is ~6× SuperGrok and you hit 100% every week |
| $300 Heavy month, 10× price-ratio | $7,700 | $482 | $241 | $647 | If Heavy is 10× SuperGrok. Still under $10k of API |
| Spend $300 on Flash list | $4,794 eq. | $300 | — | $3,568 eq. | ~830 sessions. About the 6× Heavy month, uncapped |
| Spend $300 on Flash promo | $9,586 eq. | — | $300 | $7,136 eq. | Promo only through 9 Sep 2026. Budget list after that |
| $10k of Grok-eq (16% of 4× farm) | $10,000 | $626 | $313 | $841 | Workday looping, or ~4 pegged hours/day at C8 |
| Pegged 4× farm, 24/7 C8 | $61,410 | $3,843 | $1,921 | $5,164 | ~8–13 emptied Heavy months of API value in one calendar month |
| 2× EXL3 replicas, 24/7 | $80k–$100k | $5.0k–$6.3k | $2.5k–$3.1k | $6.7k–$8.4k | Max concurrency. Worse tok/s per stream. Same four cards |
Bishop mix at list: 9.658M × $0.15 + 305.717M × $0.03 + 1.231M × $0.50 = $11.24, vs Grok $179.56 (16.0×). Promo is half. “Same tasks” is $15.10 for that week (11.9×). $10k and $61k rows reprice the Grok token mix, not equal merged PRs. Local power for the pegged farm is $237–$394/month (§8) — that is the $61k row’s real bill, not $3,843.
Yes. $10k is not the ceiling. It is a light-duty month.
Conservative decode aggregate at C8 on 4×: 400 tok/s (the 2× EXL3 box already did 336 aggregate at C16 adaptive; 0xSero’s 4× C1 is 208). Agent loops are not 100% decode — tools, prefill, and sampler waits eat the rest. Use 40% decode duty for a farm that is actually looping without a human in the chair:
400 tok/s × 0.40 × 86,400 s/day × 30.44 days = 421 million output tokens / month
Bishop’s mix costs $145.87 of Grok 4.6 API per million output tokens (because each million of output rides ~7.9M input and ~248M cache). Scale that:
24/7 C8 farm
~$61k
Grok 4.6 API equivalent / month
To hit $10k
16%
of that farm, or ~4 h/day pegged
Bishop-sessions
~10.6k
421M out ÷ 39.7k out/session
vs SuperGrok
~85×
maxed $30 months packed into one calendar month
| Utilization | Grok 4.6 API-eq | Flash list | Flash promo | Bishop-sessions | What it looks like |
|---|---|---|---|---|---|
| Workday 8h, 4 agents, 25% duty | ~$8k | $501 | $250 | ~1.4k | Just-you, looping while you work. Near $10k already. |
| 16% of 24/7 C8 farm | $10k | $626 | $313 | ~1.7k | The question. Not a stretch. OpenRouter invoice: six hundred dollars, not ten thousand. |
| Workday 8h, 8 agents, 40% duty | ~$20k | $1,252 | $626 | ~3.5k | Serious personal farm, nights off. |
| 24/7 C8, 40% duty | ~$61k | $3,843 | $1,921 | ~10.6k | Pegged. OpenRouter is ~$3.8k list. Local bill is power (~$240–$390). |
| 2× EXL3 replicas, 24/7 | ~$80k–$100k | $5.0k–$6.3k | $2.5k–$3.1k | more, slower each | Max concurrency. Worse tok/s/stream. |
“Bishop-session” = 1.23M output / 31 ≈ 39.7k output tokens, with that week’s cache/input ratio. Your loops will differ. If Flash is 2.1× more verbose, you finish fewer tasks per GPU-second, not fewer Grok-equivalent dollars — the Grok-eq column is priced at Grok rates for the same token mix. Flash list is 16.0× cheaper on that mix; “same tasks” is 11.9× (§7). Locally the tokens cost power.
A fully used Heavy month is maybe $5k–$8k of Grok 4.6 API. The 4× box at 24/7 is ~8–12 Heavy months of that value in one calendar month, or ~200 SuperGrok ($30) months. Electricity for 4× 600W + ~300W host ≈ 2.7 kW:
2.7 kW × 24 × 30.44 × $0.12/kWh ≈ $237/month · at $0.20/kWh ≈ $394/month
So the local “subscription” is power plus amortization, not $300. $10k of Grok-equivalent tokens on the box costs a few hundred dollars of electricity, not $300 of xAI. Buying that same $10k of Grok-eq on OpenRouter Flash is $626 list / $313 promo — cheap vs Grok API, still ~10–16× the electricity bill. That is the whole argument for leaving the meter, and for not pretending OpenRouter is as cheap as owning the GPUs.
Same Bishop ruler. Same 40% decode duty. Different silicon: Mac Studio M5 Ultra, 256 GB unified memory, 1.2 TB/s (Apple, 25 Aug 2026). 256GB is the SKU you can order today; 512GB follows in late October. Flash at 4-bit is a 200GB Unsloth GGUF / ~178GB MLX 4-bit. It fits. KV is the squeeze.
| Build | Size | Fits 256 GB? | Agent note |
|---|---|---|---|
| MLX 8-bit | 334 GB | No | 512GB SKU. Not this one. |
| MLX 6-bit | 256 GB | No (weights = the box) | Zero KV. Skip. |
| Unsloth UD-Q4_K_XL | 200 GB | Yes, tight | 93% of BF16 top-1. 1–2 interactive agents. Pin this for quality. |
| MLX mixed 4/8-bit | 182 GB | Yes | Better PPL than uniform 4-bit. More KV than UD-Q4. |
| UD-Q3_K_XL / IQ3_XXS | 148 / 120 GB | Yes, headroom | 4 concurrent on one box. 82–87% accuracy. Use when you need slots, not when you have four Macs. |
| UD-Q2_K_XL | 109 GB | Easy | 78% accuracy. Spark trick, not a 256GB trick. |
Unsloth GGUF sizes 28 Aug 2026; PipeNetwork MLX PPL table (8-bit 334.1 GB / mixed-4/8 181.9 GB / 4-bit 177.6 GB). Unsloth’s 4-bit “162–210 GB total memory” band is why Q4 on 256GB is a 1–2 agent machine, not an 8-agent machine.
Nobody has published a warmed GLM-5.3-Flash decode curve on M5 Ultra yet — machines start arriving 22 Sep 2026. Two M3 Ultra 512GB numbers exist: omlx Q4e 24.1 tok/s (4k prompt), antirez ds4 Q4 35.5 tok/s short / 26.6 tok/s at ~12k. M5 Ultra bandwidth is 1.2 TB/s vs M3 Ultra 819 GB/s = 1.46×. Neural Accelerators (Apple’s 4.3× peak AI compute vs M3 Ultra) help prefill, not decode. Decode stays bandwidth-bound.
Agent C1 pin: 45 tok/s · 26.6 × 1.46 ≈ 39, plus a little Metal/MTP. High-concurrency aggregate per box: 70 tok/s (1.55× C1). NVIDIA SGLang got 1.9× C1→C8; Mac saturates the bus sooner.
| Box | Interactive | Overnight | Per-stream at interactive |
|---|---|---|---|
| 1× M5 Ultra 256GB, Q4 | 2 | 4 | ~22–35 tok/s. KV is the wall, not the scheduler. |
| 1× M5 Ultra 256GB, Q3 | 4 | 8 | More slots, dumber weights. Fine for mechanical loops. |
| 4× Ultra, replica Q4 | 8 | 16 | Same per-stream as one box. Matches the 4× 6000 count. |
| 6× Ultra, replica Q4 | 12 | 24 | More agents than the 4× 6000. Each stream still ~25–35 tok/s. |
| 4× RTX PRO 6000 NVFP4 | 8 | 16 | ~50 tok/s/stream at C8 (400 aggregate). Faster each, same count as 4× Ultra. |
| 4× Ultra, TP (wrong play) | 1–2 | 2–3 | Apple 3× one system ≈ 135 tok/s one stream. Fat Kimi, not Flash agents. |
Interactive pin on one 256GB Q4 replica: 2 agents. That keeps each stream in the 20–35 tok/s band — slower than hosted Flash (~50 tok/s) and slower than the 4× 6000 at C8 (~50 tok/s/stream), still usable. Four agents on one Q4 box drops you under ~18 tok/s/stream. That is overnight, not chat.
Same formula as §8: aggregate decode tok/s × 40% duty × 30.44 days, priced at $145.87 of Grok 4.6 API per million output tokens of the Bishop mix. That is $153.50 of Grok-eq per tok/s of aggregate decode. 1× high-concurrency uses 70 tok/s; replica clusters multiply the boxes.
| Farm (high-duty looping) | Agg. tok/s | Grok 4.6 API-eq | vs $300 Heavy | Flash list | Power / mo |
|---|---|---|---|---|---|
| Empty SuperGrok Heavy month | — | $5k–$8k | 1.0× | $290–$480 | $0 |
| $300 of OpenRouter Flash list | — | $4,794 eq. | ~0.6–1.0× | $300 | $0 |
| 1× M5 Ultra 256GB | 70 | $10.7k | 1.4–2.3× | $672 | $16 |
| 1× Ultra, C1 only (no batching) | 45 | $6.9k | 0.9–1.5× | $432 | $16 |
| 4× 6000, 16% duty ($10k question) | 64 | $10k | 1.3–2.2× | $626 | ~$38 |
| 4× Ultra, replica | 280 | $43k | 5.6–9.3× | $2,690 | $63 |
| 4× RTX PRO 6000 NVFP4 | 400 | $61k | 8.0–13.3× | $3,843 | $237 |
| 6× Ultra, replica | 420 | $64k | 8.4–14.0× | $4,035 | $95 |
Heavy column is emptied-month multiples using the $4.6k–$7.7k API-eq bracket from §6. Power is 180 W/Studio under inference (M5 Ultra community ~180–214 W peak; Apple’s last Ultra max was 270 W) at $0.12/kWh, vs 2.7 kW for the 4× 6000. Promo Flash is half the list column. 1× Ultra C1-only is the honest floor if Metal batching does not show up.
1× Ultra vs Heavy
~2×
One 256GB Studio at high duty ≈ two emptied $300 months, every calendar month
6× Ultra vs 4× 6000
~$64k / $61k
Same Grok-eq band. Mac uses ~40% of the watts and more agent slots.
1× Ultra on OpenRouter
$672
Flash list for that $10.7k mix. Promo $336. Power $16.
4× Ultra on OpenRouter
$2.7k
vs $3.8k to match the 4× 6000 farm on Flash list
Same ruler, six rows, the version made to post. Token-max Heavy and $300 of Flash buy about the same Grok-eq. $300 of flagship GLM-5.3 does not — it is $1.40/$4.40, 1.8× under Grok, not 16×. Pegged silicon is a different category: the box costs about one month of the Grok 4.6 API bill it replaces, then you pay watts.
| Farm | Agents | Grok 4.6 to match | vs $300 Heavy | Flash list | You pay |
|---|---|---|---|---|---|
| Token-max SuperGrok Heavy | — | $5–8k | 1.0× | $290–480 | $300 |
| $300 GLM-5.3 Flash · OpenRouter | — | $4.8k | ~1× | $300 | $300 |
| $300 GLM-5.3 · OpenRouter | — | $550 | 0.07–0.11× | $34 | $300 |
| 1× M5 Ultra 256GB | 2 | $10.7k | 1.4–2.3× | $672 | $16 |
| 2× RTX PRO 6000 | 4–8 | $52k | 6.7–11× | $3.2k | $127 |
| 4× RTX PRO 6000 | 8 | $61k | 8–13× | $3.8k | $237 |
$300 Flash = Bishop mix at list ($11.24/week → $4,794 Grok-eq). $300 GLM-5.3 flagship at $1.40/$0.26/$4.40 = $98.42/week → $547 Grok-eq. 2× 6000 is EXL3 C16 measured 336 tok/s × $153.46/tok/s-month. Power at $0.12/kWh. Current RTX PRO 6000 street ~$14–16k/card: 2× ≈ $30–37k vs $52k of Grok-eq/month; 4× ≈ $60–72k vs $61k; Ultra 256GB ~$10k vs $10.7k. One pegged month ≈ the hardware sticker in Grok 4.6 API dollars.
LibertAIDAI/GLM-5.3-Flash-NVFP4 with the 0xSero SGLang SM120 image, TP4, flashinfer_cutlass, MTP5, FP8 KV, graphs ≤ 8. Parsers: reasoning glm45, tools glm47.max_tokens — max thinking will burn the budget inside <think> and return empty content.low for mechanical edits, high for implement-and-test, max only when a Grok-class problem would have justified Grok. Default is max; that will inflate verbosity toward the AA 2.1× figure.If the goal is maximum concurrent loops rather than nicest single-agent UX: two EXL3 TP2 replicas instead of one NVFP4 TP4 on the 4× 6000, or N Ultra replicas. Same idea either side — more endpoints, slower each stream.
flashinfer_cutlass.