Virtual Memory (MMU / TLB / PTW)
July 23, 2026 · View on GitHub
Scope: the Vortex virtual-memory subsystem — the per-core MMU
(hw/rtl/mem/VX_mmu.sv,
VX_mmu_tlb.sv,
VX_mmu_ptw.sv), the SimX model
(sim/simx/mem/mmu.{cpp,h},
mmu_tlb.{cpp,h}), and the runtime VM
software stack (sw/runtime/common/vm.{cpp,h},
sw/common/vm_types.h).
This document is the single reference for the subsystem: the architecture below, plus the usage surface — the perf counters and the randomized-VA testing knobs (§8).
VM is gated by VX_CFG_VM_ENABLE (default off,
VX_config.toml:24).
1. The v3 VM model
VM in v3 lives in two places:
- The compute-core MMU (RTL + SimX) translates VA→PA for kernel LSU/fetch traffic.
- The CP DMA software walker translates VA operands of
CMD_MEM_*commands (seecommand_processor.md§8).
There is no shared device-side MMU and no RTL CP MMU yet. The host runtime API is VA-only; the host never translates at transfer time — the CP does.
kernel LSU/fetch VA runtime (host)
│ ──────────────
▼ VMManager: PA alloc, mint VA,
per-core MMU (after coalescer) build host-shadow page tables,
│ satp[31]? BARE → bypass batched flush of dirty PT pages
▼ │
TLB CAM ── hit ─► PA ─► cache CP_SATP_LO/HI ──► CP cp_translate()
│ miss (Sv32/Sv39 SW walk of device PTs)
▼
PTW walk (PTE loads via same cache port) ─► fill TLB ─► replay as PA
2. ISA, CSRs, and configuration
- SATP: the compute cores use the RISC-V
satpCSR0x180(VX_types.toml:322), programmed once at boot byvx_start.S(csrw satpwith PT-base PPN + mode, per-core, under#ifdef VX_CFG_VM_ENABLE) and surfaced assched_csr_if.csr_satp(VX_csr_data.sv:142-150). The CP DMA carries its own separateCP_SATP_LO/HIat CP regfile offsets0x028/0x02C— same packed value, distinct mechanism (dual-SATP plumbing). - Page-table format (
VX_types.toml:47-54):VX_VM_ADDR_MODE = SV39 if XLEN==64 else SV32,VX_VM_PT_LEVEL = 3/2,VX_VM_PTE_SIZE = 8/4,VX_VM_PAGE_LOG2_SIZE = 12. PT base region atVX_MEM_PAGE_TABLE_BASE_ADDR.VM_ADDR_MODEenum lists[BARE,SV32,SV39,SV48,SV57](:646). - TLB sizing: a single flat
VX_CFG_TLB_SIZE-entry (32) fully- associative CAM, one per dcache MMU + one per icache MMU per core (VX_config.toml:160). No L2/L3. - Perf: 6 VM perf CSRs in the memory-subsystem class
VX_DCR_MPM_CLASS_MEM = 7at 0xB0B–0xB10 (+_H mirrors), alongside off-chip memory / lmem / coalescer ([csr_mpm_mem]in VX_types.toml). - Runtime caps:
VX_CAPS_VM_SUPPORT,VX_MEM_PHYS = 0x8(vortex2.h:74,121).
Sv32 vs Sv39: the SW stack, SimX MMU, and CP walker support both; the RTL PTW is Sv32-only (hardcoded 2-level), so on FPGA only Sv32/RV32 VM is real — consistent with the project's 32-bit-only RTL policy.
3. RTL components
VX_mmu.sv(top) — merges an elastic-buffered TLB path, a bypass path, and the PTW throughVX_mem_arb. Takessatp[31:0];needs_translation()bypasses only onsatp[31](BARE) — there is no address-range bypass (:38-44).VX_mmu_tlb.sv— fully-associative CAM TLB, MRU victim select, 4-state FSM (IDLE/READY/PTW_WAIT/REPLAY). It can match superpages viapage_level/vpn_mask, but fills always setpage_level=0(:282) so megapages are stored as 4 KB entries.VX_mmu_ptw.sv— Sv32-only, hardcoded 2-level walker (L1_REQ/RESP → L0_REQ/RESP → FILL); 4-byte PTE; one walk in flight; does not yet act on V/R/W/X/U flags (page faults un-delivered,:113-120).
Instantiated per core in VX_core.sv:440
(dcache MMU, DCACHE_NUM_REQS ports) and :461 (icache MMU, 1 port),
both under #ifdef VX_CFG_VM_ENABLE. The MMU sits after the
coalescer / LSU adapter (a single per-core MMU, not per-LSU-slice).
4. SimX model
sim/simx/mem/mmu.{cpp,h} is a per-core
Mmu SimObject with 4 channels. Unlike the RTL, its PTW is generalized
over VX_VM_PT_LEVEL via a level-counter FSM
(mmu.cpp:56-167), so it models both
Sv32 and Sv39; it performs real page-fault checks (PTE V, R=0 & W=1)
and aborts on fault, and reconstructs superpage offsets correctly (caching
megapages as single 4 KB TLB entries).
mmu_tlb.{cpp,h} is the matching
fully-associative TLB with reads/hits/misses/evictions counters. Wired in
core.cpp:191-213; set_satp fans out to
both MMUs.
5. Runtime VM software stack
sw/runtime/common/vm.{cpp,h} provides
VMManager: a page-table builder + VA allocator owned by vx::Device.
Page tables are host-shadowed (shadow_pt_, dirty_pt_pages_) and
flushed to the device in one bulk CMD_MEM_WRITE(physical) per dirty PT
page. mem_alloc allocates a PA from global_mem_, then either
phy_to_virt_map (mint a VA, install PTEs) or install_identity_map for
VX_MEM_PHYS buffers. The runtime API is VA-only.
The stack is VM-discovery-driven, not compile-time gated: vm.{h,cpp}
and vm_types.h carry no #ifdef VM_ENABLE, and the runtime discovers VM
at runtime from the CP DEV_CAPS.VM_ENABLED bit
(device.cpp:277), branching on
vm_enabled_. (One compile-time constant remains:
VX_CFG_VM_PINNED_REGION_SIZE for pinned-slab sizing,
device.cpp:34.) Randomized-VA
testing is available via VORTEX_RANDOMIZE_VA/VA_SEED.
sw/common/vm_types.h holds the pure
host-side SATP_t, PTE_t, vAddr_t, and Page_Fault_Exception,
Sv32/Sv39-split, including only the SW-facing VX_types.h.
5.1 CP composition
The CP DMA is MMU-aware: cmd_processor.cpp:cp_translate() performs the
identical Sv32/Sv39 walk as VMManager::page_table_walk, reading PTEs from
device RAM. Every CMD_MEM_* operand is a VA and is translated unless the
MEM_FLAG_PHYSICAL flag is set. This is the only device-side translation
path today.
6. End-to-end flow
- LSU lanes → coalescer/adapter → per-core dcache MMU. If
satpMSB is clear (BARE) the request bypasses translation; otherwise the per-lane VA hits the TLB CAM. - TLB hit → forward to cache as PA. TLB miss → kick the PTW.
- PTW walks the page table by issuing PTE-fetch loads through the same
downstream cache port (RTL:
ptw_mem_ifmerged viaVX_mem_arb; SimX:ReqOut[0]with a PTW tag marker). RTL = fixed 2-level Sv32; SimX =VX_VM_PT_LEVEL-deep loop. - On fill, the TLB caches
{vpn→ppn, flags}and the request replays as a PA to the cache. (SimX delivers page faults; RTL does not yet.) - SATP is programmed once at boot by
vx_start.S; the CP'sCP_SATPis programmed by the host atcp_init.
7. Deliberate simplifications (current model)
These are architectural choices in the as-built v3 VM, recorded so they are not mistaken for bugs:
- No ASID.
SATP_tparses anasidfield but it is unused; the RTL PTW drops thesatpASID bits. Single global VA space, no per-dispatch SATP reprogramming. - No A/D-bit writeback. The runtime pre-sets
A=D=1; the TLBs are read-only. - Page-fault delivery is partial. SimX aborts on a fault; the RTL PTW
does not check
V/R/W/X/Uyet, and no fault is routed to the LSU as an exception. - No PMP / page protection enforcement.
- SV48/SV57 enum values exist but are unimplemented.
8. Usage: perf counters and randomized-VA testing
8.1 Perf counters
Six MMU counters live in the memory-subsystem MPM class
(VX_DCR_MPM_CLASS_MEM, §2); the hardware sums the icache- and dcache-MMU
counters into one bank.
sw/runtime/common/perf.cpp:579-584
reads them and prints a per-core vm: line in the memory report (the MEM
class, selected with --perf=7 to blackbox.sh):
| CSR | Meaning |
|---|---|
VX_CSR_MPM_TLB_READS | Total TLB lookups (icache + dcache MMU) |
VX_CSR_MPM_TLB_HITS | TLB hits |
VX_CSR_MPM_TLB_MISSES | TLB misses (each triggers a PTW) |
VX_CSR_MPM_TLB_EVICTS | TLB evictions on fill |
VX_CSR_MPM_PTW_WALKS | Completed PTW walks |
VX_CSR_MPM_PTW_LATENCY | Total PTW latency in cycles (avg = LATENCY / WALKS) |
PERF: vm: tlb_reads=96, hit=96%, evicts=0, ptw_walks=4, ptw_avg_lat=84.75
8.2 Randomized-VA testing knobs
VMManager's constructor
(vm.cpp:49-53) reads two environment
variables to stress the translation path:
VORTEX_RANDOMIZE_VA=0(default) — identity mapping:vx_mem_allocreturns a VA equal to the underlying PA, so the pipeline is exercised without address remapping (a clean baseline).VORTEX_RANDOMIZE_VA=1— eachvx_mem_allocmints a random page-aligned base VA, reserves the contiguous range, and installs per-page PTEs; the caller receives the random VA while the PA stays inglobal_mem_. Multi-page buffers stay VA-contiguous.VORTEX_VA_SEED=N— seed for thestd::mt19937_64RNG (default0x12345678). A fixed seed reproduces the same VA stream across runs.
9. Proposed but not yet implemented
- GPU-aligned multi-level TLB hierarchy — a 3-level L1/L2/L3 TLB
hierarchy (multi-banked post-coalescer L1 DTLB, per-core L1 ITLB,
per-cluster L2, chip-wide L3), a centralized multi-walker PTW with a
PTE walk-cache, non-blocking TLB MSHRs, TLB inclusion-policy knobs
(NINE/INCL/EXCL), broadcast
SFENCE.VMAinvalidation, and aVX_dmablock split out of the CP with its own TLB into the shared hierarchy. None of it is implemented yet; the current MMU is a single flat 32-entry FA TLB per cache port. - RTL PTW Sv39 + superpage fills — the RTL walker is Sv32-only and stores megapages as 4 KB entries. Generalizing it (as SimX already is) is required for RV64 VM on FPGA.
- RTL page-fault delivery — check PTE
V/R/W/X/Uand route a fault to the LSU as an exception (VX_mmu_ptw.sv:113stub). - RTL CP shared device-side MMU — add the SATP regfile decode + a
hardware walker so the
CP DMA honors VM in RTL, matching the SimX/CP-software path (see
command_processor.md§10 item 2). configure --vmfirst-class flag — VM is still forced per build viaCONFIGS=-DVX_CFG_VM_ENABLE.- RTL VM in CI — the
vm()regression runs SimX-only; the rtlsim/xrt lines are commented out pending RTL PTW completion.
Superseded directions (recorded to avoid revival): a per-LSU-slice
MMU placement (replaced by a single per-core MMU after
the coalescer); wiring the orphaned sim/common/mem.cpp MemoryUnit
(replaced by the dedicated sim/simx/mem/mmu.cpp SimObject); and the
original compile-time VM_ENABLE + per-transfer host-side translation
model (replaced by runtime DEV_CAPS.VM_ENABLED discovery + MMU-aware CP
DMA — the host no longer translates per transfer).