# WASM-only reader panic in uzlib inflate Living record of the investigation into the browser-only crash when opening a book. Native QEMU (same submodule) does **not** crash, so this is a wasm32 TCG backend bug. Update as things are ruled in/out. ## Symptom Opening a book in the browser panics; native opens the same book fine. ```text Invalid read at addr 0x3FCB7999, size 1, region '(null)' size 0x0 owner '(none)' ram=0, reason: rejected Guru Meditation Error: Core 0 panic'ed (Load access fault). Exception was unhandled. MEPC : 0x4201e210 RA : 0x4201e18c SP : 0x3fcb3810 ``` ## Root cause (established) `wasm_tci_only_tb = true` in `tcg_out_qemu_st` is **upstream qemu-wasm's own design**, not a workaround this project invented. A full diff of our `tcg/wasm32/tcg-target.c.inc` against pristine upstream qemu-wasm shows the two files are identical except for that one line. Upstream unconditionally forces every store-containing TB to the TCI interpreter, so upstream's *compiled* (JIT) store path is emitted but never executed — it is effectively unfinished, untested dead code, and it corrupts guest state the moment it runs. Mechanism: the wasm32 backend compiles a TB as a chain of wasm blocks guarded by `if (BLOCK_PTR <= block_idx)`. When a softmmu slow-path helper yields (asyncify unwind for I/O / interrupt), the wasm function returns and re-enters, skipping already-executed blocks via `BLOCK_PTR`. This restart is safe to replay for loads but not for a compiled store, so upstream punts all store TBs to TCI (which has precise instruction-level checkpointing). The real fix is therefore not a port bug hunt but completing upstream's compiled store path to be rewind-safe (e.g. defer the memory commit past the TB's yield points / buffer stores until block commit), then dropping the blanket TCI. Large TCG-backend task; ~15 min per browser iteration; no native repro. ## Established facts - **Fault site (via same-version local ELF addr2line, not byte-identical):** `uzlib_uncompress` (tinflate.c:613) ← `tinf_inflate_block_data` (tinflate.c:476) ← `InflateReader::read` ← `FontDecompressor::decompressGroup` ← `prewarmCache`. "Indexing" = the reader inflating compressed glyph groups from flash into a `malloc`'d temp buffer. - **Input is static flash data**, independent of SD. The compressed font bank is in the firmware image; the SD file only selects *which* glyph groups get prewarmed (the characters present in the text). - **Native renders the reader** ("chapter one") with no fault — proven with `_scratch/native-reader-repro.py` + QMP screendumps. Confirms the firmware, the font data, and the decompression logic are all correct. - **Fault address is non-deterministic:** `0x3FCB7999` and `0x3FCB7997` across runs. A static machine-map boundary would fault at a fixed address. Drift ⇒ a computed pointer is slightly wrong, i.e. runtime/codegen, not the memory map. - **Rejected region is bogus:** `region '(null)' size 0x0 owner '(none)' ram=0`. The address lands in mapped DRAM (`0x3fc80000`+`0x60000`) but the wasm softmmu TLB routed it to the IO slow path with a zero-size section. Hallmark of a bad address/TLB result from a JIT-compiled TB. - **Menus work.** Home/file-browser render crisp text through the same font system and the same TLB fast path, so the TLB codegen is not broadly broken — the bug is pattern/data specific in a hot, JIT-promoted TB. - **User observation:** firmware-written crash logs open fine; prepopulated files do not. Consistent with data-dependence — a crash log's character set prewarms different glyph groups than lorem-ipsum, avoiding the group whose compressed byte pattern trips the miscompiled path. ## Architecture notes (why JIT is implicated) - `tcg/wasm32.c`: TBs run on the TCI interpreter until executed `INSTANTIATE_NUM` times, then get compiled to native WASM (`instantiate_wasm`) and run as added table functions. The inflate hot loop gets promoted. - `WASM_TCI_ONLY_ENTRY` forces a TB to stay on TCI forever; `wasm_tci_only_tb` (per-TB, reset each translation) requests that. - Dynamic promotion must stay on for usable performance (all-TCI starves the guest watchdog), so a permanent fix must correct codegen, not disable JIT. ## Hypotheses | # | Hypothesis | Status | |---|-----------|--------| | H1 | A store-containing TB is miscompiled (the cleanup removed `wasm_tci_only_tb = true` from `tcg_out_qemu_st`); the compiled TB corrupts a pointer a later load dereferences | **confirmed** (E1) | | H2 | Compiled **load** codegen bug (`tcg_wasm_out_qemu_ld` / `_direct`) | untested | | H3 | TLB fast-path codegen edge case (`tcg_wasm_out_tlb_load`) hit only by certain access patterns | untested | | H4 | asyncify/fiber unwind corrupts vCPU state mid-inflate | unlikely (inflate makes no async import calls) | | H5 | mimalloc / linear-memory / threads issue | unlikely (native heap layout differs, no crash) | ## Experiments ### E1 — force store-containing TBs to TCI (tests H1) ✅ CONFIRMED Re-added `wasm_tci_only_tb = true;` to `tcg_out_qemu_st` in `tcg/wasm32/tcg-target.c.inc`. - Result: **panic gone.** `make browser-test` reaches the reader and renders a correct page (black lorem-ipsum on white, footer `80% example 1/3 33%`). No reject, no `Guru Meditation`, 1 boot, 0 watchdog panics. New refresh frames (`dtm2_white=0.891`) appear past the old `0.960` stall. - **Conclusion: the compiled `qemu_st` codegen is the bug.** Forcing store TBs to TCI is a working *workaround*, not the fix — every store-containing TB then runs interpreted (broad perf cost; prior attempts reported watchdog starvation under sustained load, though this finite-burst run was clean). - The reader first paints a correct BW page (black text on white), then a follow-up **grayscale** refresh pass inverts it into muddled white-on-black text. That second issue is real and still open (see ISSUES.md "Reader renders inverted grayscale"); it is independent of this panic fix. The `browser-test` wait-for-stable happened to screenshot the initial white page. ### Caveat on E1's specificity Forcing store-*containing* TBs to TCI drops the whole TB (its loads, stores, and arithmetic) to the interpreter. So the bug is *in a TB that also has a store*, not provably the store op. `tcg_wasm_out_qemu_st_direct` and `tcg_wasm_out_qemu_ld_direct` are byte-for-byte symmetric and look correct in isolation, which points away from the direct emit and toward the store path's surrounding block/rewind/helper scaffolding in `tcg_wasm_out_qemu_st`, or an interaction with a co-located op. ### E2 — crash is icount-independent; loads-TCI also masks it (inconclusive on op) - Removed the store flag and ran at the new `noicount` default: **same crash** (`Invalid read at 0x3FCB80A5`, drifts, `MEPC 0x4201e210`). So icount timing is not the trigger. - Moved the flag to `tcg_out_qemu_ld` (loads TCI, stores compiled): **passed** (33s). Not decisive — load-containing TBs are a superset of store-containing TBs, so forcing loads to TCI drops nearly every TB (including all store TBs) to the interpreter. It masks the bug the same way store-TCI does. - **Takeaway:** whole-TB TCI cannot isolate the offending op. Precise localization needs per-op TCI infrastructure (backend is whole-TB) or reading the generated wasm for one miscompiled TB. Large task; reader already works at ~31s with the store flag, so this is headroom, not a blocker. ### E3 — PC-windowed bisection: not a single TB, it's aggregate Added a compile-window to `tcg_out_qemu_st` (`s->gen_tb->pc` gate): store-TBs inside `[LO,HI)` compile, all others stay TCI. Each run rebuilt + browser-tested. | window | region | result | |--------|--------|--------| | 0x4201d000–0x4201f800 | uzlib inflate + FontDecompressor | **pass** | | 0x40000000–0x40400000 | ROM (memcpy/memset) + IRAM (tlsf/malloc) | **pass** | | 0x42000000–0x420d0000 | flash text lower half (prewarm callers) | **pass** | | 0x420d0000–0x421a0000 | flash text upper half | **pass** | | (none / all store-TBs compiled) | everything | **crash** | Every region compiles safely *in isolation*; only the full set corrupts. So the bug is **not a localizable miscompiled op/TB** — it is an aggregate effect that appears once enough compiled store-TBs run together. The faulting load is `d->dict_ring[d->lzOff]` (tinflate.c:476) off a stack-resident inflate struct, i.e. some store clobbers a *neighbouring* stack/heap pointer, and which store depends on timing. This matches a **rewind/replay or TB-instance-management race** in the compiled restartable-block store path, not a codegen typo — and it is exactly why upstream forces every store-TB to TCI. Fixing it is architectural (rewind/instance-safe compiled stores), not a one-op correction. ### E4-E7 — deterministic minimal repro, and what the bug is NOT With a per-op gate in `tcg_out_qemu_st` (compile only store-TBs matching a criterion, TCI the rest) plus targeted instrumentation, each run rebuilt + browser-tested against the exact 1.4.1 firmware: | experiment | change | result | |-----------|--------|--------| | E4 lifecycle | never `removeFunction` retired TB instances (pin all) | **crash** | | E5 helper-only | force every store through the softmmu helper (no direct store codegen) | **crash** | | E6 width: byte only | compile TBs whose store is 8-bit | **pass** | | E6 width: 16 only | 16-bit | **pass** | | E6 width: 32 only | 32-bit | **pass** | | E6 width: byte+16 | | **pass** | | E6 width: 16+32 | | **pass** | | E6 width: **byte+32** | | **crash** | | E7 bytewise-32 | emit 32-bit store as four 8-bit stores, compile all | **crash** | | precheck | skip store helper on rewind when DONE_FLAG already set | **crash** | | nochain | `-d nochain` (disable TB chaining), compile all stores | **crash** | | determinism | rerun the crashing all-compiled build 3x | **crash x3, same gen 16, same MEPC** | **Established:** the crash is *deterministic* (same point every run) but depends on the exact compiled-TB set. The minimal reproducer is a byte-store TB **and** a 32-bit-store TB compiled together; it is *counter-monotonic* — additionally compiling 16-bit-store TBs hides it. **Ruled out:** single miscompiled TB (E1-E3), a specific memory region (inflate/ROM/IRAM/flash halves), TB chaining (nochain), instance/function lifecycle & `removeFunction` reuse (E4 pin), the store *data* codegen itself (E5 helper-only and E7 bytewise both crash), and store double-execution on rewind (precheck to skip the replayed helper still crashes). **Correction:** the rule is monotonic, not counter-monotonic — `{8,16,32}` (original) also crashes, so 16-bit is irrelevant. The rule is simply: **crash iff an 8-bit-store TB and a 32-bit-store TB are both compiled.** ### E8 — the 8-bit requirement is aggregate, not one TB Compiled all 32-bit-store TBs, then compiled 8-bit-store TBs only inside one region at a time: | 8-bit region compiled (with all 32-bit compiled) | result | |---|---| | flash lower 0x42000000-0x420d0000 | **pass** | | ROM+IRAM 0x40000000-0x40400000 (incl. memset) | **pass** | | flash upper 0x420d0000-0x421a0000 | **pass** | | all 8-bit (no region gate) | **crash** | Those three regions cover essentially all executable memory, yet each passes while their union crashes. So the 8-bit side is **aggregate** too — there is no single interacting 8-bit TB. Combined with E1-E3 (region-aggregate) this proves the corruption is an **emergent property of running many compiled store-TBs**, which whole-TB TCI gating (the backend's only control) cannot dissect. The remaining paths are generated-wasm inspection of the BLOCK_PTR/register restart sequence, or building per-op TCI. A deterministic reproducer exists (remove `wasm_tci_only_tb`, open example.txt, faults at MEPC 0x4201e210, gen 16). ### E9 — static analysis of generated wasm (dump + wasm2wat) Added a node-only dump hook (`wasm_dump_module`, tcg.c finalize) that writes the first compiled store-TB and load-TB modules; `_scratch/dump-tbs.mjs` boots the firmware under node far enough to emit them. `wasm2wat --enable-threads` gives readable IR. **Finding 1 (real bug, fixed-then-shelved):** the ld/st slow paths detect an asyncify yield with the thread-global `DONE_FLAG` (ctx+20), while the general call path uses `UNWINDING` (ctx+28), which is reset before each call and set *only* by the coroutine switch. `DONE_FLAG` is set by *any* nested ld/st helper, so a nested guest access inside a store/load helper that then yields makes the outer op falsely read "completed" and swallow the unwind. Converting ld/st to the `UNWINDING` mechanism is the correct fix. **But** rebuilding with compiled stores + the UNWINDING fix still crashes identically (gen 16, same MEPC), so the store helpers on this path never actually yield — this bug is latent here. Change shelved (not shipped) to avoid touching the working compiled-load path without a green test; recorded for a future upstream fix. **Finding 2 (reframes the root cause):** combine E5 with the load result. E5 routes *every* store through the C softmmu helper (correct memory write) and still crashes; loads share the identical restart/block structure and work. So the corruption is **not** in the store data codegen, the restart machinery, or the yield handling. What remains is that a store *writes* guest memory using an address/value taken from guest registers (wasm globals). A compiled TB that computes a wrong register value corrupts memory only through a store — a wrong address fed to a *load* just reads harmless garbage. This matches every prior result: emergent (depends which ops precede a store), deterministic, and non-localizable by store region or width. The remaining hunt is a register-producing op whose compiled codegen diverges from TCI, exposed destructively via stores — which needs a compiled-vs-TCI per-TB register differential, not more whole-TB gating. Tooling retained: `wasm_dump_module` + `wasm_tb_had_store/load` tags + `_scratch/dump-tbs.mjs`. ### Next options 1. **Ship E1 as a guarded workaround** — reader works now; measure real navigation/page-turn perf before judging it unacceptable. 2. **Localize precisely** — bisect by making the TCI trigger conditional (store width; presence of the miss/helper path; TB size) to shrink the suspect set, then read the generated wasm for that TB. ## Diagnostics currently in the tree (uncommitted) - `system/memory.c`: reject log prints region size/owner/ram to identify the faulting region. - `hw/display/xteink_x3_eink.c`: `[eink] refresh` popcount trace (per-plane white ratio) — also used for the separate grayscale-rendering issue. ## Reproduction - **WASM:** `make browser-test` (the test does 4 confirm presses: Browse Files → Test → highlight → Open). Panic appears in `_scratch/browser-raw-sd.log`. - **Native (control, no panic):** `nix develop -c python3 _scratch/native-reader-repro.py`, then inspect `_scratch/native-step4-reader.ppm`. ## Ruled out - Not SD / filesystem: fixed separately (raw FAT32, SSI-SD token, CMD13); the file reads correctly and "Indexing" begins. - Not firmware / font data: native decompresses and renders identical input. - Not the machine memory map: fault address drifts; native map is the same. ## E10: compiled-vs-TCI per-TB differential (opt-in `?tcgdiff`) — store codegen exonerated Built a per-TB differential checker to settle whether single-TB compiled store codegen diverges from TCI. It runs authoritative TCI first, snapshots its result, rewinds guest state, runs the compiled TB as a disposable shadow, compares, then restores the TCI result. Opt-in via `?tcgdiff` (`wasm_diff_enable()` from `web/app.js` `preRun`); normal/browser execution is unchanged. Making the comparison *valid* — not the codegen — was the whole task. Five harness defects each masqueraded as a "register divergence" until fixed: 1. **Execution order** — running the compiled shadow first contaminated TCI's machine state. Fixed: TCI first, shadow second, restore TCI. 2. **Partial-TB comparison** — helper/slow-path TBs don't reach `exit_tb`. Fixed with a `done_flag == 2` sentinel: only fully-inline store TBs that complete (`done_flag == 2 && tb_ptr == 0`) are compared. 3. **Negative-offset state** — the TB-entry interrupt check reads `env-8` (`CPUNegativeOffsetState`, before `CPUArchState`). Snapshot/restore now covers the full `CPUNegativeOffsetState`. 4. **Deep TLB state** — `CPUTLB.f[].table` / `d[].fulltlb` are separately allocated; TCI warms the soft-TLB via its ld/st helper path, so the shadow must start from an identical TLB. Now deep-copied per MMU mode. 5. **Async exit-flag race** — the IO thread asynchronously sets `icount_decr.high = -1` via `cpu_exit()`. TCI can take the TB-entry interrupt exit while the restored shadow sees a cleared flag and runs fully. These are unavoidable concurrency races, so comparisons where the force-exit flag was set around the TCI run are skipped. Only side-effect-safe candidates are compared: store-containing TBs with no general helper call (`wasm_tb_had_helper`), restricted to the app-flash PC window. Live dispatch stays TCI for all store TBs; the compiled function runs only as a shadow. **Result: 160,000+ store-TB comparisons during a full reader-open, zero register or DRAM divergence.** Single-TB compiled store codegen is bit-for-bit identical to TCI. This exonerates per-TB store codegen and confirms the corruption is emergent across TBs (consistent with E1–E8). A per-TB differential is structurally blind to a multi-TB interaction bug; catching it needs a whole-execution differential (two full machines stepped in lockstep from reset), not more per-TB work. The `?tcgdiff` harness is expensive (full DRAM + TLB snapshot and double execution per candidate; heap grows to ~1 GiB, frame gen drops) but is debug-only and does not affect the production store-TCI path. ## E11: whole-execution differential (opt-in `?tcgdiff2`) — live compiled vs TCI shadow Mode 2 inverts E10: the **compiled** path runs LIVE and authoritative (the real buggy path, so corruption *accumulates* across TBs exactly as in the crash), while a per-TB **TCI shadow** — restored from the pre-snapshot — is the known-good reference. Store-TCI is disabled under this mode so compiled stores actually execute. Per eligible store TB: snapshot pre (env, neg, deep TLB, DRAM) -> run compiled live, capture full compiled-post -> rewind -> run TCI shadow -> compare -> restore compiled-post -> continue. Opt-in via `?tcgdiff2`. ### Result - **153,000+ live-compiled-vs-TCI comparisons with accumulation, zero divergence** (bootloader region). Combined with E10 (160k reader-path), inline compiled stores are correct even under live cross-TB accumulation. - Revealed a large blind spot: ~90% of store TBs complete via the store **slow-path** (`done_flag == 1`, TLB-miss -> ld/st helper), not the pure inline path (`done_flag == 2`). Mode 2 was extended to compare slow-path completions too. ### Why it cannot reach the reader corruption (impractical in-browser) The 2x execution + full DRAM (0x60000) + deep-TLB snapshot per eligible TB is ~1000x slower than normal. Consequences: - The guest **interrupt watchdog** fires first: gating comparison to app-flash produced an `Interrupt wdt timeout` panic (MEPC 0x40380264, IRAM idle) during cold boot — a harness-overhead artifact, not the corruption Guru Meditation (MEPC 0x4201e210). - Ungated, it reached only ~153k comparisons (bootloader, gen 2) before the 1100s timeout. Full reader-open is millions of TBs away. Defanging the watchdog would not help; raw throughput is the wall. ### Narrowed suspects (what the differential structurally cannot cheaply test) 1. **Store slow-path** (TLB miss -> `helper_stN_mmu` -> Asyncify unwind/resume), the ~90% of stores mode 2 excluded by the `done_flag == 2` gate. Aligns with E9's `DONE_FLAG` (thread-global) vs `UNWINDING` (per-call) finding. 2. **Helper-containing store TBs**, excluded by the side-effect-safety filter. 3. **8/32-bit store-width interaction** (E1-E8's "needs both widths"). ### Decision In-browser whole-execution differential is retired as a repro path: it proves live-accumulation correctness of inline stores but cannot reach the crash. Next instrument is a deterministic, framework-free **slow-path store differential unit test** (no boot, no watchdog, no perf wall) targeting suspect (1)/(3). Option A (sequential deterministic two-pass) remains a documented future option if a full-execution comparison is later required. Both `?tcgdiff` (E10) and `?tcgdiff2` (E11) are retained as opt-in tools; production store-TCI path is unchanged. ## E12/E13: normal-speed compiled stores and file-open-armed E11 ### `?nostoretci` probe Added opt-in `?nostoretci` to disable the store-TCI workaround without enabling any differential overhead. Result: the exact original corruption still reproduces at normal speed during reader open: - `Load access fault` - `MEPC 0x4201e210` - `MTVAL 0x3fcb8004` / invalid read near the inflate/text buffer region So store-TCI is still required in production. ### `?tcgdiff2=open` Arming E11 from boot was too slow or perturbed app init. Added a browser-test hook so `?tcgdiff2=open` arms `wasm_diff_enable2()` immediately before the test presses Confirm to open `example.txt`. Boot and app init therefore run on the normal production path, and compiled stores become live only over the repro interval. With the all-app-flash window (`0x42000000..0x42200000`), the first observed mode-2 divergence after arming is register state, not DRAM: ```text TCGDIFF pc=0x4201d6ea gpr[2] compiled=0x3fcb42f0 tci=0x3fcb42e0 pre-ra=0x42029374 pre-sp=0x3fcb42f0 ``` This is `utf8NextCodepoint()`: TCI applies the stack-frame prologue (`sp -= 16`) while the compiled-live path has stale `sp` in `CPUArchState`. The later crash is still the known invalid read at `MEPC 0x4201e210`. ### Current interpretation E10 remains green because it forces no-chain and runs TCI live, then compiled as a disposable shadow. E11 is the real compiled-live path and exposes state that is stale across compiled execution/chaining boundaries. The remaining suspect is no longer individual store codegen; it is **compiled TB state materialization**: WASM register globals and `CPUArchState` can diverge across real compiled-live entry/exit/chaining, and store-TCI masks the bad store-containing TBs by routing them through TCI. A tested broad fix candidate — making store TBs return to C instead of using the same-module `BLOCK_PTR` chain — caused an app-init interrupt-watchdog timeout, so it is not viable as-is. A narrower fix should target state materialization (flush/reload live guest register state at compiled TB boundaries or before helper/exception-visible points) rather than disabling chaining wholesale. ## E14: narrowed production fix — only 32-bit store TBs stay TCI The blanket upstream qemu-wasm fallback interpreted every store-containing TB. The width interaction tests show that is broader than necessary: - all stores compiled (`?nostoretci`) -> exact crash remains (`MEPC 0x4201e210`, invalid read near `0x3FCB80xx`) - only 8-bit qemu store TBs interpreted -> reader-open passes - only 32-bit qemu store TBs interpreted -> reader-open passes Production now keeps only **32-bit qemu store TBs** on TCI; 8/16/64-bit store TBs compile normally. This breaks the bad 8-bit/32-bit compiled-store interaction while preserving more dynamic WASM promotion than the blanket store-TCI fallback. `?nostoretci` remains as the normal-speed all-compiled negative-control repro and still crashes at the original panic. Validation: ```text normal URL: browser raw-sd test passed, open example.txt took 36.8s ?nostoretci: Load access fault, MEPC 0x4201e210, invalid read 0x3FCB80A5 ``` This fixes the browser panic without disabling dynamic promotion globally. The remaining upstream-quality improvement would be eliminating the 32-bit store-TCI subset entirely by finding the exact cross-width/shared-memory state bug. ### E14 benchmark correction — 32-bit-only fallback rejected as default A 3-run benchmark changed the decision: | fallback | result | open times | | --- | --- | --- | | 32-bit-store-only TCI | 2 pass, 1 app-init watchdog | 34.0s, 33.8s | | blanket store-TCI | 3 pass | 35.3s, 36.6s, 34.3s | The narrowed fallback is a little faster when it works, but it is not reliable enough: one run failed before controls enabled with an interrupt watchdog. The default is restored to the upstream blanket store-TCI fallback. Keep the 32-bit-only result as a useful diagnostic, not a production fix. ## E15/E16: the corruption is a nondeterministic race, not a miscompiled TB Added runtime-settable, opt-in store gating to isolate the culprit without rebuilds: - `?sthash=1|2` — compile one hash-parity half of *all* store TBs (both widths present in the compiled half), interpret the other. - `?stcw=&stcmask=&stcval=` — keep every width compiled except the selected width `stcw`, which compiles only when `(hash(pc) & m) == v`. Intended for single-width binary search; logs `STBISECT compile pc=... w=...`. ### Attempted bisection and what it actually showed - `sthash=1` (compile odd-hash half) -> crash; `sthash=2` (even half) -> pass. Initial read: a specific pair of TBs. - Width bisection contradicted that: with all other widths compiled, **both** 8-bit halves crash and **both** 32-bit halves crash. Identity of the compiled store TBs does not matter; only that enough of them run. ### Decisive test: determinism Re-running a single fixed config exposes nondeterminism: ```text sthash=2 run 1 -> CRASH sthash=2 run 2 -> PASS sthash=2 run 3 -> CRASH ``` So the earlier single-run "pass" results (including the width fallbacks and the first `sthash=2`) were probabilistic, not fixes. **The compiled-store corruption is a nondeterministic race**, not a deterministic miscompilation of any specific TB. E10/E11 (per-TB compiled == TCI, 300k+ comparisons) are fully consistent: each TB is individually correct; the fault is emergent timing/ordering. ### Reliability of the production default Blanket store-TCI over 5 runs: 4 clean reader-opens, 1 failure — and that 1 was a **cold-boot interrupt-watchdog timeout** (`MEPC 0x4219a37e`), a separate host-scheduling boot flake, not the `0x4201e210` corruption. So blanket store-TCI shows 0 corruption crashes here; it suppresses the race in practice but the underlying race remains. ### Revised suspect Prime suspect is now shared-memory ordering / concurrency between the compiled guest-store path and another agent touching guest RAM (IO/main thread device models, dirty-bit/notdirty handling, or the Asyncify store slow-path `DONE_FLAG`/`UNWINDING` re-entrancy flagged in E9), rather than store codegen. The crash is a corrupted *pointer* (invalid read at `0x3FCB80xx`, region null), consistent with a torn/stale store of a pointer-sized value under a race. ### Instrumentation status `?sthash` and `?stcw/stcmask/stcval` are opt-in debug tools; the default build path is unchanged (blanket store-TCI). The narrowed 32-bit-only fallback remains rejected as a default because it only lowers, not removes, the race probability. ## E17: not the store block-replay/redispatch path Added opt-in `?yieldtrace`: the dispatch loop counts compiled-TB "yield/resume" re-entries, defined as a compiled TB returning non-zero with the *same* `tb_ptr` and `do_init == 0` (the `DONE_FLAG`-driven redispatch/resume path; chains set `do_init = 1`, completion sets `tb_ptr = 0`). Result over three crashing `?nostoretci` runs: **all crashed at `0x4201e210`, zero yield re-entries**. So the store-specific block-replay-via-redispatch mechanism (upstream's stated reason for blanket store-TCI) does not fire during this workload. That path is not the trigger here. Caveat: Asyncify-based unwinds (coroutine yields) restore the whole C stack and are transparent to this counter, so a coroutine-mediated replay is not excluded. But the store `DONE_FLAG` redispatch path is definitively ruled out. ### Consolidated conclusion The `0x4201e210` corruption is: - nondeterministic (same config crashes/passes across runs), - independent of which specific store TBs are compiled, - gated only by having enough compiled stores of both widths (a timing knob), - not the store `DONE_FLAG` redispatch/replay path. The consistent remaining explanation is a **data race on guest RAM** between the CPU worker (compiled `i64.storeN`) and another agent that runs concurrently — main-thread `QEMUTimer`/device-model callbacks (SPI/SD/GDMA/cache) reading or writing guest DRAM. TCI's C store path masks it by changing timing/serialization, not by being semantically different. This matches upstream qemu-wasm shipping blanket store-TCI rather than a targeted fix. A real fix is a threading-correctness project: ensure guest-RAM accesses from device/timer callbacks are properly serialized against vCPU stores (BQL/memory barriers), or run device memory access on the vCPU thread. Opt-in repro/debug controls now available: `?nostoretci`, `?sthash`, `?stcw/stcmask/stcval`, `?yieldtrace`. Production default remains blanket store-TCI. ## E18: not dynamic function-table churn either Confirmed the threading model: the build uses `-pthread -sPROXY_TO_PTHREAD=1 -sASYNCIFY=1`, and `qemu_thread_create` runs the vCPU in its own pthread, concurrent with the main/IO-loop thread over shared memory. So a genuine guest-RAM data race is possible. Tested the most tractable non-thread hypothesis first: each compiled TB adds a WASM table function and batch-`removeFunction`s once >500 removals are pending; more compiled stores (both widths) means more churn, which is timing-dependent. Added opt-in `?noremove` (`wasm_set_disable_tb_removal`) to leak functions instead of recycling table indices. Result: `?nostoretci&noremove` crashed 4/4 (Load access faults). Recycling of WASM table indices is **not** the corruption source. ### Ruled out so far - specific store-TB identity (E15) - store `DONE_FLAG` block-replay/redispatch (E17) - WASM function-table index reuse (E18) ### Standing conclusion The corruption is a nondeterministic concurrency issue between the vCPU pthread (compiled `i64.storeN`) and the concurrent main/IO thread (timer/device callbacks, AIO/block completions) over shared guest memory or QEMU softmmu metadata — the class of bug upstream qemu-wasm avoids with blanket store-TCI. Confirming the exact racing access needs thread-aware instrumentation (e.g., logging IO-thread guest-RAM writes and TLB mutations during reader open) rather than more store-gating experiments. Production stays on blanket store-TCI. Opt-in debug knobs now available: `?nostoretci`, `?sthash`, `?stcw/stcmask/stcval`, `?yieldtrace`, `?noremove`. ## E19: IO-thread guest-RAM access is not present in the failing window Added opt-in `?ramrace` instrumentation for off-vCPU guest-DRAM access: - `address_space_read` / `address_space_write` - cached read/write slow paths - `address_space_map(..., is_write)` / DMA-style map reads - direct mapped write unmap dirtying - off-vCPU TLB flush entry points Also added `?ramrace=open`, armed by the browser harness immediately before opening `example.txt`, to avoid the known cold-boot watchdog flake from always-on probe overhead. Results: - Always-on `?nostoretci&ramrace` crash: two off-vCPU full TLB flushes during early boot only; zero off-vCPU DRAM reads/writes/maps/unmaps. - Open-window `?nostoretci&ramrace=open` crash at the original signature (`Invalid read 0x3FCB80A6`, `MEPC 0x4201e210`): **zero** off-vCPU DRAM reads, writes, maps, unmaps, and TLB flushes. So the specific hypothesis from E18 — concurrent main/IO-thread device or timer callbacks touching guest DRAM during inflate — is ruled out for the corruption window. The remaining race is not an IO-thread guest-RAM accessor. The fault address is inside the configured ESP32-C3 DRAM aperture (`0x3fc80000..0x3fce0000`), so the access fault is still best read as corrupted vCPU execution state/translation, not an unmapped legitimate pointer. ### Revised suspect space after E19 - vCPU-thread-only nondeterminism in the wasm backend around async exits, helper unwinds, or context/global state restoration that is invisible to the `DONE_FLAG` redispatch counter. - A data race on QEMU/TCG metadata or CPUState (not guest RAM), such as async exit/interrupt state observed by the vCPU while compiled stores are active. - SharedArrayBuffer memory-ordering differences between compiled WASM stores and the C/TCI softmmu path, without another QEMU guest-RAM writer in the same window. Production default remains blanket store-TCI. The next useful probe is not more IO-thread DRAM logging; it is targeted vCPU-state/TLB-entry instrumentation at the failing translated block or around async exit handling. ## E20-E22: state materialization probes E11 had already identified the first live divergence after arming at open: `pc=0x4201d6ea` in `utf8NextCodepoint()`, where TCI had applied the prologue `sp -= 16` to `CPUArchState` but compiled-live still had stale `env->gpr[2]`. The disassembly confirms `0x4201d6d6..0x4201d6e8` is exactly the function prologue and first branch: ```text 4201d6d6: addi sp,sp,-16 4201d6d8: sw s2,0(sp) ... 4201d6e4: lbu s1,0(s2) 4201d6e8: beqz s1,4201d7a8 4201d6ea: mv s0,a0 ``` ### E20: no-chain boundary probe Added opt-in `?stateflush=open`, armed by the browser harness immediately before opening `example.txt`. It reuses the existing `wasm_diff_needs_nochain_at(pc)` app-flash window to force app TBs to return to C without running the expensive compiled-vs-TCI differential shadow. Result with `?nostoretci&stateflush=open`: original corruption still reproduced (`Invalid read 0x3FCB80A6`, `MEPC 0x4201e210`). Returning to C between app TBs is not enough; the wasm backend does not materialize `CPUArchState` merely by leaving a compiled TB. ### E21/E22: blind helper spill attempts are invalid Tested a direct backend fix candidate: spill TCG global-mem temps before C helper calls so slow-path load/store helpers and exceptions see coherent architectural state. - spilling all global-mem temps before+after helpers caused an earlier invalid read around `0x07585d82`; - spilling before helpers only caused the same failure; - narrowing to RISC-V integer GPR globals (`xN/...`) still caused immediate invalid writes around `0x075f....`. Conclusion: helper-boundary materialization is still the right suspect, but it must respect TCG liveness/state. Blind backend spills at helper emission can write stale globals back into `CPUArchState` and make corruption worse. The fix belongs in the wasm backend's handling of TCG sync/liveness around calls (or in targeted generated `sync_i32/i64` ops), not in an unconditional helper wrapper spill. ### Revised next step Trace the TCG liveness/sync ops for the divergent prologue TB and the later slow-path helper TB: verify whether `la_global_sync()` marks `x2/sp` `TS_MEM` across the relevant call/branch, and whether the wasm backend emits the corresponding store. The observed bug is now specifically: **compiled globals and `CPUArchState` are allowed to diverge, and the wasm backend exposes stale env to helper/exception-visible code when store TBs stay compiled.** ## E23-E31: concrete TLB-array corruption Added `FAULTTRACE`/`TLBTRACE` under `?ramrace=open` to inspect the exact failing load path. Key observations from crashing `?nostoretci&ramrace=open` runs: 1. The faulting guest address (`0x3fcb8004`/`0x3fcb80a6`) is valid ESP32-C3 DRAM. 2. TLB fill resolves the page correctly as RAM: ```text FAULTTRACE tlb_fill vaddr=0x3fcb8000 paddr=0x3fcb8000 asidx=0 prot=0x3 mr=esp32c3.iram ram=1 romd=0 xlat=0x3c000 FAULTTRACE tlb_store index=56 iotlb=0x9c000 addr_page=0x3fcb8000 stored_xlat=0xffffffffc03e4000 phys=0x3fcb8000 FAULTTRACE mmu_lookup1 addr=0x3fcb8000 index=56 maybe_resized=1 tlb_addr=0x3fcb8000 flags=0x0 full_xlat=0xffffffffc03e4000 addend=0xc19d4000 ``` 3. Before the next byte load on the same page, the compact TLB and full TLB entry are corrupted: ```text FAULTTRACE mmu_lookup1 addr=0x3fcb8004 index=56 maybe_resized=0 tlb_addr=0x3fcb8200 flags=0x200 full_xlat=0x0 addend=0xc0348000 FAULTTRACE mmio_read addr=0x3fcb8004 xlat_section=0x0 mr=(null) ram=0 ``` `0x200` is `TLB_MMIO`, so the vCPU incorrectly dispatches a DRAM read through the MMIO/unassigned path. This is why the access is rejected even though the guest address is in DRAM. The corruption happens inside the active compiled TB before it returns, so the post-TB scanner only sees the last good state. The faulting firmware sequence is in `uzlib_uncompress`: ```text 4201e18c: lw a3,56(s0) 4201e18e: lw a5,24(s0) 4201e190: lw a4,52(s0) 4201e194: addi a2,a5,1 4201e198: sw a2,24(s0) 4201e19a: add a3,a3,a4 4201e19c: lbu a4,0(a3) # faults after TLB entry was corrupted 4201e1a0: sb a4,0(a5) ``` The compiled store at `0x4201e198` is now the prime suspect: it likely writes through a bad host address into the TLB arrays, changing the same page's TLB entry from RAM to MMIO and zeroing `fulltlb[index].xlat_section`. That matches why blanket store-TCI works: the C/TCI store path does not corrupt the TLB arrays. ### Current root-cause statement The bug is no longer just "nondeterministic race". It is specifically: > compiled wasm store fast-path sometimes computes/uses a bad host address and > writes into the vCPU softmmu TLB arrays (`CPUTLBEntry`/`CPUTLBEntryFull`), > causing the immediately following load from valid DRAM to be treated as MMIO > with `xlat_section=0`. Still unknown: why the store fast path obtains the bad host address. Candidate remaining causes: stale/corrupted TLB addend for the store's page, wasm32 address truncation/sign-extension around addends, or an in-TB ordering bug in the compiled store TLB lookup. The next focused probe is to trace the store fast-path host address for the `0x4201e198` TB before the `i64.store32` executes. ## E32-E34: fault-path TB isolation Tried a narrow compiled-store guard (`?tlbguard=open`) that forces a store fast path to the C helper if the computed host address lands inside the vCPU TLB arrays. It did **not** suppress the crash, even though the guard range included `fulltlb[56]`: ```text TLBGUARD range lo=0x13e8000 hi=0x353c000 ``` So the corrupting write is not caught as a normal compiled `qemu_st` fast-path whose final host pointer falls directly inside that TLB-array range. The corruption still appears between the TLB fill and the following byte load. Added `?stpc=LO-HI`, an opt-in PC-window override that interprets only store-containing TBs whose start PC lands in the window, while `?nostoretci` keeps every other store TB compiled. Results: - plain `?nostoretci`: still crashes at the original signature. - `?nostoretci&stpc=0x4201e180-0x4201e220`: **3/3 pass** (`open example.txt` about 46-52s). - wider/narrowed windows around `uzlib_uncompress` also pass as long as they include the fault-path TB start. This pins the corruption to compiled execution of the `uzlib_uncompress` TB that contains the copy-loop sequence around: ```text 4201e18c: lw a3,56(s0) 4201e18e: lw a5,24(s0) 4201e190: lw a4,52(s0) 4201e194: addi a2,a5,1 4201e198: sw a2,24(s0) 4201e19a: add a3,a3,a4 4201e19c: lbu a4,0(a3) 4201e1a0: sb a4,0(a5) ``` The PC-window override is firmware-specific and **not** a production fix, but it is now the smallest positive isolation: one compiled store-containing TB in the inflate copy loop is sufficient to corrupt the vCPU TLB entry; interpreting that TB avoids the crash while preserving compiled stores elsewhere. Next useful step: dump or instrument the generated WASM for this exact TB and inspect the `sw a2,24(s0)` / following `lbu` fast-path sequence. The current node dump helper only captures early boot TBs, so late-reader TB dumping needs a browser/open-window trigger or a PC-filtered `wasm_dump_module` path.