explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

  1. Home
  2. /
  3. Dictionary
  4. /
  5. Memory Bandwidth
Infrastructure & Hardwareaka bandwidth

Memory Bandwidth

Memory bandwidth is the rate at which data moves between memory and compute units, often the true bottleneck for inference since modern models are memory-bound, not compute-bound.

Ask Melo about this← all terms

During autoregressive inference, each token generation reads the entire model's weights from memory but performs relatively little computation per byte loaded, making memory bandwidth — not FLOP/s — the binding constraint. HBM3e on modern GPUs provides around 4–8 TB/s, and the ratio of bandwidth to compute determines the arithmetic intensity crossover. This is why quantization (fewer bytes per weight) speeds up inference almost linearly even though it reduces precision.

Related terms

High-Bandwidth MemoryVRAMQuantizationLatencyThermodynamic ComputingEmbedded AI

Where Memory Bandwidth comes up

  • NVIDIA Vera CPU: Great Silicon, Misleading Whitepaper
  • Gemma 4 26B A4B Runs 2x Faster on Mac — Community MLX Optimization
  • FreeToken Runs a 753B MoE Locally — But the GPU Is Only Half the Story
  • Petals Resurfaces: Why BitTorrent-Style LLM Inference Still Struggles