Hardware · Local AI · Apple Silicon

Mac Studio Memory Roadmap

Every serious hint on next-gen unified memory — 512GB gone, 768GB tested, 1.5TB whispered for M7 Ultra — plus bandwidth math and what that means for models you can actually run at home.

July 28, 2026 Rumors + official specs Sources: Bloomberg/Gurman · Apple · NVIDIA · MacRumors · Macworld Audience: local-AI builders

1. Why this matters now

For local LLM work, Apple’s pitch has never been “beats an H100 at FLOPs.” It has been: one coherent memory pool big enough to hold the weights, quiet enough to live under a desk, efficient enough that a full rack of GPUs is overkill for a single researcher’s agent loop.

The March 2025 M3 Ultra Mac Studio made that concrete with a 512GB unified-memory ceiling — Apple’s “most memory ever in a personal computer” line. Then the global DRAM shortage ate the high SKUs. By early 2026 the 512GB option disappeared from the configurator; later storefront snapshots showed M3 Ultra configs pinned near the 96GB floor while delivery estimates stretched months out.

Capacity is the product. Bandwidth is the tax. Supply decides which SKUs exist.

So when Bloomberg-sourced reporting says Apple has tested 768GB on an M5 Ultra Studio and is looking at 1.5TB on an M7 Ultra — possibly never sold to the public — that is not trivia. It is the entire local-AI value proposition for this form factor, under a memory market that is actively hostile to huge consumer configs.

2. What we already know: M3 Ultra 512GB

Anchor numbers from Apple’s own Mac Studio tech specs (current-generation Ultra SKU as shipped in 2025):

Axis M3 Ultra (top) Notes
CPU / GPU Up to 32-core CPU · 80-core GPU · 32-core Neural Engine Two dies fused (UltraFusion lineage)
Memory bandwidth 819 GB/s Apple published figure for Ultra configs
Unified memory (launch) 96 GB base · up to 512 GB 512GB was a ~$4,000 uplift at launch era pricing
512GB SKU status (2026) Removed amid DRAM shortage High-end BTO temporarily / indefinitely unavailable
Base pricing (post June 2026 hike) Ultra base reported ~$5,299 Up from ~$3,999 launch base in secondary coverage

Side-by-side on the same product page: M4 Max Studio configs advertise 410 GB/s (32-core GPU class) or 546 GB/s (40-core GPU class). Ultra is still the bandwidth and capacity king of the consumer Mac line — just not an HBM monster.

For local models: 512GB at ~819 GB/s is a capacity-first machine. You can hold weights that a 24–96GB discrete GPU never will. Prefill and decode still feel the ~0.8 TB/s ceiling hard — more on that below.

3. Hints teased out of the rumor mill

None of the following is Apple-confirmed silicon. It is a chain of secondary reporting (MacRumors, Macworld, 9to5Mac, Tom’s Hardware) paraphrasing Mark Gurman’s Bloomberg Power On notes and related supply-chain pieces. Treat every number as tested / planned / supply-gated, not as a store SKU.

Hint A — M5 Ultra has been tested to 768GB

June 2026 coverage: Apple plans an M5 Ultra Mac Studio with roughly 36 CPU cores and 80 GPU cores, and has tested support for up to 768GB of unified memory — a step above the prior 512GB ceiling. Same articles stress that supply constraints may block shipping that SKU at launch.

Apple has tested support for up to 768GB of unified memory, but supply constraints could prevent it from launching with an option for that much memory. MacRumors, June 25, 2026 (citing Bloomberg)

That is the clean answer to “will there be another 500 / 512 class high end?” The technical ceiling in the rumor set is 768, not a rehash of 512. Whether you can buy 768 is a separate, uglier question.

Hint B — Internal code J246, earlier 2026 slipped to later 2026

Macworld notes an internal code J246 for the M5 Ultra Studio, replacing the March 2025 M3 Ultra machine. Earlier expectations of a first-half 2026 ship slid toward an October 2026 window after memory-chip supply and price shocks. By late July 2026, broader Mac roadmap pieces still put M5 Ultra “this year” but with timing described as memory-supply dependent and currently unclear.

Hint C — Pricing hints are illustrative, not SKUs

After the June 2026 Mac price increases (Ultra base climbing to the mid-$5k range), secondary pieces floated that a fully loaded high-memory M5 Studio could clear $10,000. Forum math goes higher. None of that is a leaked Apple price sheet — it is extrapolation from DRAM cost and prior $4k 512GB uplifts.

Hint D — 1.5TB is an M7 Ultra story, and maybe not a product

Gurman-sourced Macworld reporting: while preparing M5 Ultra, Apple is also working on an M7 Ultra that supports up to 1.5TB of memory — roughly double the planned M5 Ultra high mark. The same coverage warns the public may never be able to buy that configuration; internal AI-server use and RAM-market conditions gate what becomes a BTO option. Arrival talk centers on ~2028–2029, not the next refresh.

Hint E — M5 Ultra for internal AI servers too

Separate from the Studio: Apple is described as building internal AI servers using M5 Ultra. Apple does not sell servers, so this is Apple-for-Apple capacity. It still matters: silicon that must host large on-device / on-prem inference for Apple’s own stack is more likely to keep extreme memory configs in the design envelope even if the public never sees the top bin.

768GB tested (M5 Ultra) ~36 CPU / 80 GPU (rumored) J246 Studio code 1.5TB (M7 Ultra, maybe internal) DRAM shortage still dominates

4. Chip roadmap: M5 → skipped M6 high-end → M7

This is the part people get wrong when they say “wait for M6 Ultra” or “M7 next year.” The rumor consensus as of mid–late July 2026:

Chip Window (reported) Memory notes Studio relevance
M5 Max / M5 Ultra Studio refresh targeted 2026 (slipped) Ultra tested to 768GB; Max lower (laptop/desktop shared lineage) Next Studio — primarily a chip refresh, not a chassis redesign
M6 (base only) Fall 2026 class entry Macs No serious Pro/Max/Ultra memory story — those tiers skipped Not the high-end Studio path
M6 Pro / Max / Ultra Skipped N/A Do not plan a Studio around an M6 Ultra
M7 base ~H1 2027 Mid-line Macs
M7 Pro / M7 Max Late 2027 Major neural / AI uplift (car-team IP lore in secondary coverage) Feeds the Ultra that follows
M7 Ultra ~2028 Studio (some copy says into 2029) Up to 1.5TB support; may stay internal Second Studio hop; possible better heatsink / internals
Why skip M6 Pro/Max/Ultra? Gurman-line reporting: Apple had planned larger neural-processing jumps for the M7 family and decided those AI gains were worth breaking the usual Pro/Max/Ultra cadence. AI priority over tidy generational naming — not a cancellation of high-end Macs forever.

Practical planning rule: if you are buying a desk-side Apple box for models in the next 12–18 months, the decision is M3 Ultra (if you can find memory) vs wait for M5 Ultra. “Wait for M6 Ultra” is a dead branch in the current rumor tree. “Wait for M7 Ultra / 1.5TB” is a 2028-class bet with non-trivial chance the top bin never becomes a consumer SKU.

5. Memory bandwidth by generation

Capacity decides which models fit. Bandwidth decides how fast tokens leave the machine once weights are resident — especially decode, long-context attention traffic, and multi-agent concurrency.

Platform Peak memory bandwidth Capacity (high end) Confidence
M1/M2 Ultra class (historical) ~800 GB/s Up to 128–192 GB era tops Shipped
M3 Ultra Mac Studio 819 GB/s 512 GB (was); lower while shortage Apple official
M4 Max Mac Studio 410 or 546 GB/s Lower than Ultra (config-dependent) Apple official
M5 Max (shipping laptops / pending Studio) Higher than prior Max; exact Ultra bus TBD High laptop configs in the 128 GB class discourse Partial / editorial
M5 Ultra (rumored) Unannounced — expect “above Max, Ultra-class”; wild editorial guesses of multi-hundred GB/s to 1+ TB/s if double-die scales cleanly Tested 768 GB Capacity: high-confidence rumor · BW: speculation
M7 Ultra (rumored) Unannounced; “closer to Blackwell-class AI” is marketing-adjacent rumor language, not a GB/s figure Up to 1.5 TB (maybe not public) Low–medium; far out
NVIDIA RTX PRO 6000 Blackwell 1,792 GB/s GDDR7 96 GB ECC NVIDIA official
NVIDIA DGX Station (GB300 deskside) GPU 7.1 TB/s HBM3e · CPU ~396 GB/s LPDDR5X · NVLink-C2C ~900 GB/s class coherent fabric ~748 GB coherent (252 GB HBM + 496 GB CPU-side in published splits) NVIDIA / partner official
Rule of thumb: Apple wins on single-pool capacity under a quiet desk. Discrete NVIDIA wins on bytes per second past the GPU. DGX Station tries to take both — at server-adjacent cost, power, and software gravity (CUDA + NVIDIA stack).

Editorial pieces have floated M5 Studio configs “exceeding 600 GB/s.” Treat that as a floor-ish narrative for high Max/Ultra class, not a measured Ultra number. Until Apple prints a GB/s figure on a press slide, do not bake 1.2 TB/s into a purchase model.

6. What models you can run (capacity math, not magic)

Rough GGUF / MLX planning weights (order-of-magnitude; KV cache and runtime overhead still bite):

Model scale (dense-ish) ~Q4 weights On 96 GB On 256 GB On 512 GB On 768 GB
70B-class ~35–45 GB Comfortable + context Easy multi-model Overkill alone Overkill alone
120B-class ~65–80 GB Tight / short context Comfortable Easy Easy
200–300B dense ~110–180 GB No Fits mid/high Comfortable Comfortable + longer ctx
405B-class Q4 ~220–260 GB No Borderline / no Yes (hero config) Yes + headroom
~600–700B dense Q4 ~330–420 GB No No Yes if lean Much healthier
Huge MoE (active params small) Varies — total params can still be 400B–1T+ Depends on total resident experts Often yes for open MoE Yes for larger MoE Best chance at “frontier-adjacent open” weights + agents

What 512GB actually unlocked

The M3 Ultra 512GB machine’s superpower was never “fastest tokens in the lab.” It was:

What 768GB adds over 512GB

Capacity wins

  • +50% pool → larger dense Q4 / Q5, or same model with much more KV
  • Room for speculative draft + target pairs resident
  • More concurrent agent models without unloading
  • Heavier multimodal towers (vision + LLM) co-resident

What it does not automatically add

  • 2× tokens/s — that needs bandwidth + matmul engines
  • Training parity with multi-GPU CUDA boxes
  • CUDA ecosystem (MLX / Metal / llama.cpp / ollama paths differ)
  • A guaranteed shippable SKU in a DRAM crisis
Bandwidth-bound reality check: If M5 Ultra lands near ~0.8–1.0 TB/s while weights grow 50%, some workloads will feel slower per token on the bigger model even though they newly fit. Capacity and speed are not the same upgrade.

7. Mac Studio vs RTX PRO 6000 vs DGX Station

Three different answers to “local AI on a desk.” Comparing them as if they were the same product is how people waste money.

Mac Studio M3 Ultra 512GB Mac Studio M5 Ultra 768GB (rumored) RTX PRO 6000 Blackwell DGX Station (GB300)
Memory model Unified LPDDR-class pool Same architecture family 96 GB GDDR7 on GPU ~748 GB coherent CPU+GPU
Peak BW 819 GB/s TBD (expect Ultra-class, not HBM) 1,792 GB/s 7.1 TB/s GPU HBM (+ slower CPU pool)
Best at Large weights, quiet, Mac software Same + bigger ceiling if SKU ships Fast inference / creative pro / mid models Near-cluster deskside, huge models, NVIDIA stack
Weak at Raw tok/s vs discrete, CUDA Same class of limits + price risk Models >~70–90 GB resident Price, power, heat, procurement
Power / form Small, efficient Studio Same chassis (rumored chip refresh) 600 W class card + workstation chassis Deskside supercomputer
Software gravity MLX, Metal, llama.cpp, macOS agents Same + whatever M5 neural bits unlock CUDA, TensorRT, pro viz Full NVIDIA AI developer stack

How to choose without self-gaslighting

Capacity peers: A rumored 768GB M5 Ultra and a ~748GB DGX Station look similar on a spreadsheet. They are not. DGX’s fast pool is still the HBM side; Apple’s pool is uniform and slower. For decode-heavy agent work that is memory-capacity bound, Apple can feel “enough.” For bandwidth-hungry prefill, fine-tune steps, or CUDA-native stacks, DGX / multi-GPU wins even with less total DRAM on a single RTX card.

8. Is 768GB worth upgrading from M3 Ultra 512GB?

Upgrade for capacity and architecture — not for a free 50% tok/s.

Strong upgrade case

  • You actually run models that fill 400–500 GB today
  • You want long context + giant weights co-resident
  • You need multi-agent stacks that thrash on 512
  • M5 neural / GPU improvements land for your stack
  • You never got a 512 SKU because it vanished

Weak upgrade case

  • Your daily driver is 70B–120B Q4
  • You are bandwidth-starved, not capacity-starved
  • You need CUDA more than unified memory
  • You would only buy 768 if it is “cheap” (it will not be)
  • You can wait for M7 Ultra / 1.5TB clarity

Relative to a fully optioned M3 Ultra 512GB that you already own and use: +256GB is meaningful only if you have a workload that hits the wall. Relative to a stripped current storefront (96GB Ultra while 512 is gone): any high-memory M5 Ultra that actually ships is a different machine class again.

Price discipline: if high-memory configs land in the $10k–$15k+ rumor band while used or leftover 512GB Ultras exist, do the dollars-per-GB math including resale. DRAM scarcity has already shown Apple will delete SKUs rather than eat infinite margin compression.

9. Buy / wait checklist

Do not pre-order a fantasy SKU. Tested 768GB ≠ listed 768GB. 1.5TB support ≠ buyable 1.5TB. Memory markets can delete the entire high end of a product page overnight — they already did once for 512GB.

10. Sources & confidence

Official / primary-ish

Rumor / secondary (memory ceiling & roadmap)

Confidence tags

Claim Confidence
M3 Ultra 819 GB/s; launch max 512 GB High — Apple published
512 GB BTO removed during DRAM crunch High — storefront / multi-outlet reporting
M5 Ultra Studio planned; tested to 768 GB Medium–high — multi-outlet Gurman paraphrase
768 GB will be a buyable launch SKU Low–medium — same sources warn supply may block it
M6 Pro/Max/Ultra skipped; M7 Pro/Max 2027; M7 Ultra 2028 Medium — roadmap reporting, not Apple slides
M7 Ultra 1.5 TB public config Low — support claimed; public sale explicitly doubted
Specific M5 Ultra GB/s number Speculation until Apple prints it
RTX PRO 6000 / DGX Station numbers High — vendor specs (watch for SKU variants)