Memory and the Memory Hierarchy

Context: FIT1047_MOC · how memory is addressed and why it’s the performance bottleneck · registers from Sequential Circuits (Latches, Flip-Flops, Registers) · MUX addressing from Combinational Circuits (Adders, Decoders, MUX, ALU)

Quick Revision

  • 🎯 Objective: memory = addressed boxes whose interpretation belongs to the PROGRAM ➔ address bits reach locations ➔ hierarchy trades speed for size (registers → caches → RAM → disk → network).
  • 📦 Core Components: byte- vs word-addressable ➔ RAM (random access) ➔ cache (recently-used + prefetch) ➔ swapping (RAM’s overflow to disk).
  • ⚡ Key Constraint: RAM is up to slower than registers — caching exists because programs reuse and neighbour-access data; a cache miss can cost .

📝 Core

1. Addressing

  • Interpretation is the program’s job ➔ neither memory nor CPU knows if a word is a number, char, pixel or instruction (stored program).
  • Byte-addressable ➔ one address per byte (most architectures); word-addressable ➔ one address per word (MARIE: 16-bit words).
  • Capacity formula address bits ⟹ locations. MARIE worked example: -bit addresses ⟹ words 2 bytes bytes KiB.
  • RAM = random access ➔ any address, any order, same cost — unlike tape/disk whose heads must move.
  • RAM module internals ➔ chips arranged in rows × columns; an address splits into (row, column, byte-within-chip) fields, each decoded by a MUX — lecture module: bits Mbit, 17-bit addresses split 3+3+11.

2. The Bottleneck and the Cache

  • ProblemLoad 1A0 / Add 1A0 / Store 1A0 touches the same address thrice; waiting on RAM each time stalls the CPU.
  • Cache ➔ memory inside the CPU between registers and RAM: faster than RAM, slower than registers; smaller than RAM, larger than registers; keeps recently used values close.
  • Transparent ➔ same Load 1A0 — hardware serves it from cache when possible; stores may go cache-only (write-back) or cache+RAM (write-through) or RAM-only.
  • Prefetch trick ➔ programs walk consecutive addresses (text scan, image pixels, array loops) ⟹ a miss loads the next few values too; eviction discards oldest.
  • Why it’s tricky ➔ cache misses make identical code up to slower; performance depends on the cache architecture; code itself is cached too (it lives in RAM).

3. Swapping (the other direction)

  • Not enough RAM? ➔ write unused RAM regions to disk, read back on demand — the opposite of caching; hardware + OS implement it (virtual memory, Week 5).

⚖️ Core Decision Matrix — the hierarchy (lecture table)

LevelSpeed vs registersSizeWhere
Registers (in-cycle)a few wordsinside CPU
L1 cache slower~100 KBinside CPU
L2/L3 cache slower~1–8 MBin/near CPU
RAM slowerGBsmain bus
Disk/SSD~ slowerTBsexternal
Network–∞?external

✍️ Practice

⚠️ Common Mistakes

  • 💡 Word- vs byte-addressable changes capacity ➔ same address width, different total bytes; MARIE’s 8 KiB needs BOTH facts ( AND 2-byte words).
  • 💡 Caching ≠ swapping ➔ cache pulls hot data toward the CPU; swapping pushes cold data away to disk — opposite directions, both transparent.
  • 💡 “Random access” is about order-independence ➔ not randomness; the contrast class is sequential media (tape, disk head travel).