39 KiB
WASM-only reader panic in uzlib inflate
Living record of the investigation into the browser-only crash when opening a book. Native QEMU (same submodule) does not crash, so this is a wasm32 TCG backend bug. Update as things are ruled in/out.
Symptom
Opening a book in the browser panics; native opens the same book fine.
Invalid read at addr 0x3FCB7999, size 1, region '(null)' size 0x0 owner '(none)' ram=0, reason: rejected
Guru Meditation Error: Core 0 panic'ed (Load access fault). Exception was unhandled.
MEPC : 0x4201e210 RA : 0x4201e18c SP : 0x3fcb3810
Root cause (established)
wasm_tci_only_tb = true in tcg_out_qemu_st is upstream qemu-wasm's own
design, not a workaround this project invented. A full diff of our
tcg/wasm32/tcg-target.c.inc against pristine upstream qemu-wasm shows the two
files are identical except for that one line. Upstream unconditionally forces
every store-containing TB to the TCI interpreter, so upstream's compiled
(JIT) store path is emitted but never executed — it is effectively unfinished,
untested dead code, and it corrupts guest state the moment it runs.
Mechanism: the wasm32 backend compiles a TB as a chain of wasm blocks guarded
by if (BLOCK_PTR <= block_idx). When a softmmu slow-path helper yields
(asyncify unwind for I/O / interrupt), the wasm function returns and re-enters,
skipping already-executed blocks via BLOCK_PTR. This restart is safe to
replay for loads but not for a compiled store, so upstream punts all store TBs
to TCI (which has precise instruction-level checkpointing).
The real fix is therefore not a port bug hunt but completing upstream's compiled store path to be rewind-safe (e.g. defer the memory commit past the TB's yield points / buffer stores until block commit), then dropping the blanket TCI. Large TCG-backend task; ~15 min per browser iteration; no native repro.
Established facts
- Fault site (via same-version local ELF addr2line, not byte-identical):
uzlib_uncompress(tinflate.c:613) ←tinf_inflate_block_data(tinflate.c:476) ←InflateReader::read←FontDecompressor::decompressGroup←prewarmCache. "Indexing" = the reader inflating compressed glyph groups from flash into amalloc'd temp buffer. - Input is static flash data, independent of SD. The compressed font bank is in the firmware image; the SD file only selects which glyph groups get prewarmed (the characters present in the text).
- Native renders the reader ("chapter one") with no fault — proven with
_scratch/native-reader-repro.py+ QMP screendumps. Confirms the firmware, the font data, and the decompression logic are all correct. - Fault address is non-deterministic:
0x3FCB7999and0x3FCB7997across runs. A static machine-map boundary would fault at a fixed address. Drift ⇒ a computed pointer is slightly wrong, i.e. runtime/codegen, not the memory map. - Rejected region is bogus:
region '(null)' size 0x0 owner '(none)' ram=0. The address lands in mapped DRAM (0x3fc80000+0x60000) but the wasm softmmu TLB routed it to the IO slow path with a zero-size section. Hallmark of a bad address/TLB result from a JIT-compiled TB. - Menus work. Home/file-browser render crisp text through the same font system and the same TLB fast path, so the TLB codegen is not broadly broken — the bug is pattern/data specific in a hot, JIT-promoted TB.
- User observation: firmware-written crash logs open fine; prepopulated files do not. Consistent with data-dependence — a crash log's character set prewarms different glyph groups than lorem-ipsum, avoiding the group whose compressed byte pattern trips the miscompiled path.
Architecture notes (why JIT is implicated)
tcg/wasm32.c: TBs run on the TCI interpreter until executedINSTANTIATE_NUMtimes, then get compiled to native WASM (instantiate_wasm) and run as added table functions. The inflate hot loop gets promoted.WASM_TCI_ONLY_ENTRYforces a TB to stay on TCI forever;wasm_tci_only_tb(per-TB, reset each translation) requests that.- Dynamic promotion must stay on for usable performance (all-TCI starves the guest watchdog), so a permanent fix must correct codegen, not disable JIT.
Hypotheses
| # | Hypothesis | Status |
|---|---|---|
| H1 | A store-containing TB is miscompiled (the cleanup removed wasm_tci_only_tb = true from tcg_out_qemu_st); the compiled TB corrupts a pointer a later load dereferences |
confirmed (E1) |
| H2 | Compiled load codegen bug (tcg_wasm_out_qemu_ld / _direct) |
untested |
| H3 | TLB fast-path codegen edge case (tcg_wasm_out_tlb_load) hit only by certain access patterns |
untested |
| H4 | asyncify/fiber unwind corrupts vCPU state mid-inflate | unlikely (inflate makes no async import calls) |
| H5 | mimalloc / linear-memory / threads issue | unlikely (native heap layout differs, no crash) |
Experiments
E1 — force store-containing TBs to TCI (tests H1) ✅ CONFIRMED
Re-added wasm_tci_only_tb = true; to tcg_out_qemu_st in
tcg/wasm32/tcg-target.c.inc.
- Result: panic gone.
make browser-testreaches the reader and renders a correct page (black lorem-ipsum on white, footer80% example 1/3 33%). No reject, noGuru Meditation, 1 boot, 0 watchdog panics. New refresh frames (dtm2_white=0.891) appear past the old0.960stall. - Conclusion: the compiled
qemu_stcodegen is the bug. Forcing store TBs to TCI is a working workaround, not the fix — every store-containing TB then runs interpreted (broad perf cost; prior attempts reported watchdog starvation under sustained load, though this finite-burst run was clean). - The reader first paints a correct BW page (black text on white), then a
follow-up grayscale refresh pass inverts it into muddled white-on-black
text. That second issue is real and still open (see ISSUES.md "Reader renders
inverted grayscale"); it is independent of this panic fix. The
browser-testwait-for-stable happened to screenshot the initial white page.
Caveat on E1's specificity
Forcing store-containing TBs to TCI drops the whole TB (its loads, stores,
and arithmetic) to the interpreter. So the bug is in a TB that also has a
store, not provably the store op. tcg_wasm_out_qemu_st_direct and
tcg_wasm_out_qemu_ld_direct are byte-for-byte symmetric and look correct in
isolation, which points away from the direct emit and toward the store path's
surrounding block/rewind/helper scaffolding in tcg_wasm_out_qemu_st, or an
interaction with a co-located op.
E2 — crash is icount-independent; loads-TCI also masks it (inconclusive on op)
- Removed the store flag and ran at the new
noicountdefault: same crash (Invalid read at 0x3FCB80A5, drifts,MEPC 0x4201e210). So icount timing is not the trigger. - Moved the flag to
tcg_out_qemu_ld(loads TCI, stores compiled): passed (33s). Not decisive — load-containing TBs are a superset of store-containing TBs, so forcing loads to TCI drops nearly every TB (including all store TBs) to the interpreter. It masks the bug the same way store-TCI does. - Takeaway: whole-TB TCI cannot isolate the offending op. Precise localization needs per-op TCI infrastructure (backend is whole-TB) or reading the generated wasm for one miscompiled TB. Large task; reader already works at ~31s with the store flag, so this is headroom, not a blocker.
E3 — PC-windowed bisection: not a single TB, it's aggregate
Added a compile-window to tcg_out_qemu_st (s->gen_tb->pc gate): store-TBs
inside [LO,HI) compile, all others stay TCI. Each run rebuilt + browser-tested.
| window | region | result |
|---|---|---|
| 0x4201d000–0x4201f800 | uzlib inflate + FontDecompressor | pass |
| 0x40000000–0x40400000 | ROM (memcpy/memset) + IRAM (tlsf/malloc) | pass |
| 0x42000000–0x420d0000 | flash text lower half (prewarm callers) | pass |
| 0x420d0000–0x421a0000 | flash text upper half | pass |
| (none / all store-TBs compiled) | everything | crash |
Every region compiles safely in isolation; only the full set corrupts. So the
bug is not a localizable miscompiled op/TB — it is an aggregate effect that
appears once enough compiled store-TBs run together. The faulting load is
d->dict_ring[d->lzOff] (tinflate.c:476) off a stack-resident inflate struct,
i.e. some store clobbers a neighbouring stack/heap pointer, and which store
depends on timing. This matches a rewind/replay or TB-instance-management
race in the compiled restartable-block store path, not a codegen typo — and it
is exactly why upstream forces every store-TB to TCI. Fixing it is architectural
(rewind/instance-safe compiled stores), not a one-op correction.
E4-E7 — deterministic minimal repro, and what the bug is NOT
With a per-op gate in tcg_out_qemu_st (compile only store-TBs matching a
criterion, TCI the rest) plus targeted instrumentation, each run rebuilt +
browser-tested against the exact 1.4.1 firmware:
| experiment | change | result |
|---|---|---|
| E4 lifecycle | never removeFunction retired TB instances (pin all) |
crash |
| E5 helper-only | force every store through the softmmu helper (no direct store codegen) | crash |
| E6 width: byte only | compile TBs whose store is 8-bit | pass |
| E6 width: 16 only | 16-bit | pass |
| E6 width: 32 only | 32-bit | pass |
| E6 width: byte+16 | pass | |
| E6 width: 16+32 | pass | |
| E6 width: byte+32 | crash | |
| E7 bytewise-32 | emit 32-bit store as four 8-bit stores, compile all | crash |
| precheck | skip store helper on rewind when DONE_FLAG already set | crash |
| nochain | -d nochain (disable TB chaining), compile all stores |
crash |
| determinism | rerun the crashing all-compiled build 3x | crash x3, same gen 16, same MEPC |
Established: the crash is deterministic (same point every run) but depends on the exact compiled-TB set. The minimal reproducer is a byte-store TB and a 32-bit-store TB compiled together; it is counter-monotonic — additionally compiling 16-bit-store TBs hides it.
Ruled out: single miscompiled TB (E1-E3), a specific memory region
(inflate/ROM/IRAM/flash halves), TB chaining (nochain), instance/function
lifecycle & removeFunction reuse (E4 pin), the store data codegen itself
(E5 helper-only and E7 bytewise both crash), and store double-execution on
rewind (precheck to skip the replayed helper still crashes).
Correction: the rule is monotonic, not counter-monotonic — {8,16,32}
(original) also crashes, so 16-bit is irrelevant. The rule is simply: crash iff
an 8-bit-store TB and a 32-bit-store TB are both compiled.
E8 — the 8-bit requirement is aggregate, not one TB
Compiled all 32-bit-store TBs, then compiled 8-bit-store TBs only inside one region at a time:
| 8-bit region compiled (with all 32-bit compiled) | result |
|---|---|
| flash lower 0x42000000-0x420d0000 | pass |
| ROM+IRAM 0x40000000-0x40400000 (incl. memset) | pass |
| flash upper 0x420d0000-0x421a0000 | pass |
| all 8-bit (no region gate) | crash |
Those three regions cover essentially all executable memory, yet each passes
while their union crashes. So the 8-bit side is aggregate too — there is no
single interacting 8-bit TB. Combined with E1-E3 (region-aggregate) this proves
the corruption is an emergent property of running many compiled store-TBs,
which whole-TB TCI gating (the backend's only control) cannot dissect. The
remaining paths are generated-wasm inspection of the BLOCK_PTR/register restart
sequence, or building per-op TCI. A deterministic reproducer exists (remove
wasm_tci_only_tb, open example.txt, faults at MEPC 0x4201e210, gen 16).
E9 — static analysis of generated wasm (dump + wasm2wat)
Added a node-only dump hook (wasm_dump_module, tcg.c finalize) that writes the
first compiled store-TB and load-TB modules; _scratch/dump-tbs.mjs boots the
firmware under node far enough to emit them. wasm2wat --enable-threads gives
readable IR.
Finding 1 (real bug, fixed-then-shelved): the ld/st slow paths detect an
asyncify yield with the thread-global DONE_FLAG (ctx+20), while the general
call path uses UNWINDING (ctx+28), which is reset before each call and set
only by the coroutine switch. DONE_FLAG is set by any nested ld/st helper,
so a nested guest access inside a store/load helper that then yields makes the
outer op falsely read "completed" and swallow the unwind. Converting ld/st to
the UNWINDING mechanism is the correct fix. But rebuilding with compiled
stores + the UNWINDING fix still crashes identically (gen 16, same MEPC), so the
store helpers on this path never actually yield — this bug is latent here. Change
shelved (not shipped) to avoid touching the working compiled-load path without a
green test; recorded for a future upstream fix.
Finding 2 (reframes the root cause): combine E5 with the load result. E5 routes every store through the C softmmu helper (correct memory write) and still crashes; loads share the identical restart/block structure and work. So the corruption is not in the store data codegen, the restart machinery, or the yield handling. What remains is that a store writes guest memory using an address/value taken from guest registers (wasm globals). A compiled TB that computes a wrong register value corrupts memory only through a store — a wrong address fed to a load just reads harmless garbage. This matches every prior result: emergent (depends which ops precede a store), deterministic, and non-localizable by store region or width. The remaining hunt is a register-producing op whose compiled codegen diverges from TCI, exposed destructively via stores — which needs a compiled-vs-TCI per-TB register differential, not more whole-TB gating.
Tooling retained: wasm_dump_module + wasm_tb_had_store/load tags +
_scratch/dump-tbs.mjs.
Next options
- Ship E1 as a guarded workaround — reader works now; measure real navigation/page-turn perf before judging it unacceptable.
- Localize precisely — bisect by making the TCI trigger conditional (store width; presence of the miss/helper path; TB size) to shrink the suspect set, then read the generated wasm for that TB.
Diagnostics currently in the tree (uncommitted)
system/memory.c: reject log prints region size/owner/ram to identify the faulting region.hw/display/xteink_x3_eink.c:[eink] refreshpopcount trace (per-plane white ratio) — also used for the separate grayscale-rendering issue.
Reproduction
- WASM:
make browser-test(the test does 4 confirm presses: Browse Files → Test → highlight → Open). Panic appears in_scratch/browser-raw-sd.log. - Native (control, no panic):
nix develop -c python3 _scratch/native-reader-repro.py, then inspect_scratch/native-step4-reader.ppm.
Ruled out
- Not SD / filesystem: fixed separately (raw FAT32, SSI-SD token, CMD13); the file reads correctly and "Indexing" begins.
- Not firmware / font data: native decompresses and renders identical input.
- Not the machine memory map: fault address drifts; native map is the same.
E10: compiled-vs-TCI per-TB differential (opt-in ?tcgdiff) — store codegen exonerated
Built a per-TB differential checker to settle whether single-TB compiled
store codegen diverges from TCI. It runs authoritative TCI first, snapshots
its result, rewinds guest state, runs the compiled TB as a disposable shadow,
compares, then restores the TCI result. Opt-in via ?tcgdiff
(wasm_diff_enable() from web/app.js preRun); normal/browser execution is
unchanged.
Making the comparison valid — not the codegen — was the whole task. Five harness defects each masqueraded as a "register divergence" until fixed:
- Execution order — running the compiled shadow first contaminated TCI's machine state. Fixed: TCI first, shadow second, restore TCI.
- Partial-TB comparison — helper/slow-path TBs don't reach
exit_tb. Fixed with adone_flag == 2sentinel: only fully-inline store TBs that complete (done_flag == 2 && tb_ptr == 0) are compared. - Negative-offset state — the TB-entry interrupt check reads
env-8(CPUNegativeOffsetState, beforeCPUArchState). Snapshot/restore now covers the fullCPUNegativeOffsetState. - Deep TLB state —
CPUTLB.f[].table/d[].fulltlbare separately allocated; TCI warms the soft-TLB via its ld/st helper path, so the shadow must start from an identical TLB. Now deep-copied per MMU mode. - Async exit-flag race — the IO thread asynchronously sets
icount_decr.high = -1viacpu_exit(). TCI can take the TB-entry interrupt exit while the restored shadow sees a cleared flag and runs fully. These are unavoidable concurrency races, so comparisons where the force-exit flag was set around the TCI run are skipped.
Only side-effect-safe candidates are compared: store-containing TBs with no
general helper call (wasm_tb_had_helper), restricted to the app-flash PC
window. Live dispatch stays TCI for all store TBs; the compiled function runs
only as a shadow.
Result: 160,000+ store-TB comparisons during a full reader-open, zero register or DRAM divergence. Single-TB compiled store codegen is bit-for-bit identical to TCI. This exonerates per-TB store codegen and confirms the corruption is emergent across TBs (consistent with E1–E8). A per-TB differential is structurally blind to a multi-TB interaction bug; catching it needs a whole-execution differential (two full machines stepped in lockstep from reset), not more per-TB work.
The ?tcgdiff harness is expensive (full DRAM + TLB snapshot and double
execution per candidate; heap grows to ~1 GiB, frame gen drops) but is
debug-only and does not affect the production store-TCI path.
E11: whole-execution differential (opt-in ?tcgdiff2) — live compiled vs TCI shadow
Mode 2 inverts E10: the compiled path runs LIVE and authoritative (the real
buggy path, so corruption accumulates across TBs exactly as in the crash),
while a per-TB TCI shadow — restored from the pre-snapshot — is the
known-good reference. Store-TCI is disabled under this mode so compiled stores
actually execute. Per eligible store TB: snapshot pre (env, neg, deep TLB,
DRAM) -> run compiled live, capture full compiled-post -> rewind -> run TCI
shadow -> compare -> restore compiled-post -> continue. Opt-in via ?tcgdiff2.
Result
- 153,000+ live-compiled-vs-TCI comparisons with accumulation, zero divergence (bootloader region). Combined with E10 (160k reader-path), inline compiled stores are correct even under live cross-TB accumulation.
- Revealed a large blind spot: ~90% of store TBs complete via the store
slow-path (
done_flag == 1, TLB-miss -> ld/st helper), not the pure inline path (done_flag == 2). Mode 2 was extended to compare slow-path completions too.
Why it cannot reach the reader corruption (impractical in-browser)
The 2x execution + full DRAM (0x60000) + deep-TLB snapshot per eligible TB is ~1000x slower than normal. Consequences:
- The guest interrupt watchdog fires first: gating comparison to app-flash
produced an
Interrupt wdt timeoutpanic (MEPC 0x40380264, IRAM idle) during cold boot — a harness-overhead artifact, not the corruption Guru Meditation (MEPC 0x4201e210). - Ungated, it reached only ~153k comparisons (bootloader, gen 2) before the 1100s timeout. Full reader-open is millions of TBs away. Defanging the watchdog would not help; raw throughput is the wall.
Narrowed suspects (what the differential structurally cannot cheaply test)
- Store slow-path (TLB miss ->
helper_stN_mmu-> Asyncify unwind/resume), the ~90% of stores mode 2 excluded by thedone_flag == 2gate. Aligns with E9'sDONE_FLAG(thread-global) vsUNWINDING(per-call) finding. - Helper-containing store TBs, excluded by the side-effect-safety filter.
- 8/32-bit store-width interaction (E1-E8's "needs both widths").
Decision
In-browser whole-execution differential is retired as a repro path: it proves live-accumulation correctness of inline stores but cannot reach the crash. Next instrument is a deterministic, framework-free slow-path store differential unit test (no boot, no watchdog, no perf wall) targeting suspect (1)/(3). Option A (sequential deterministic two-pass) remains a documented future option if a full-execution comparison is later required.
Both ?tcgdiff (E10) and ?tcgdiff2 (E11) are retained as opt-in tools;
production store-TCI path is unchanged.
E12/E13: normal-speed compiled stores and file-open-armed E11
?nostoretci probe
Added opt-in ?nostoretci to disable the store-TCI workaround without enabling
any differential overhead. Result: the exact original corruption still
reproduces at normal speed during reader open:
Load access faultMEPC 0x4201e210MTVAL 0x3fcb8004/ invalid read near the inflate/text buffer region
So store-TCI is still required in production.
?tcgdiff2=open
Arming E11 from boot was too slow or perturbed app init. Added a browser-test
hook so ?tcgdiff2=open arms wasm_diff_enable2() immediately before the test
presses Confirm to open example.txt. Boot and app init therefore run on the
normal production path, and compiled stores become live only over the repro
interval.
With the all-app-flash window (0x42000000..0x42200000), the first observed
mode-2 divergence after arming is register state, not DRAM:
TCGDIFF pc=0x4201d6ea gpr[2] compiled=0x3fcb42f0 tci=0x3fcb42e0
pre-ra=0x42029374 pre-sp=0x3fcb42f0
This is utf8NextCodepoint(): TCI applies the stack-frame prologue (sp -= 16)
while the compiled-live path has stale sp in CPUArchState. The later crash
is still the known invalid read at MEPC 0x4201e210.
Current interpretation
E10 remains green because it forces no-chain and runs TCI live, then compiled as
a disposable shadow. E11 is the real compiled-live path and exposes state that is
stale across compiled execution/chaining boundaries. The remaining suspect is no
longer individual store codegen; it is compiled TB state materialization:
WASM register globals and CPUArchState can diverge across real compiled-live
entry/exit/chaining, and store-TCI masks the bad store-containing TBs by routing
them through TCI.
A tested broad fix candidate — making store TBs return to C instead of using the
same-module BLOCK_PTR chain — caused an app-init interrupt-watchdog timeout,
so it is not viable as-is. A narrower fix should target state materialization
(flush/reload live guest register state at compiled TB boundaries or before
helper/exception-visible points) rather than disabling chaining wholesale.
E14: narrowed production fix — only 32-bit store TBs stay TCI
The blanket upstream qemu-wasm fallback interpreted every store-containing TB. The width interaction tests show that is broader than necessary:
- all stores compiled (
?nostoretci) -> exact crash remains (MEPC 0x4201e210, invalid read near0x3FCB80xx) - only 8-bit qemu store TBs interpreted -> reader-open passes
- only 32-bit qemu store TBs interpreted -> reader-open passes
Production now keeps only 32-bit qemu store TBs on TCI; 8/16/64-bit store
TBs compile normally. This breaks the bad 8-bit/32-bit compiled-store
interaction while preserving more dynamic WASM promotion than the blanket
store-TCI fallback. ?nostoretci remains as the normal-speed all-compiled
negative-control repro and still crashes at the original panic.
Validation:
normal URL: browser raw-sd test passed, open example.txt took 36.8s
?nostoretci: Load access fault, MEPC 0x4201e210, invalid read 0x3FCB80A5
This fixes the browser panic without disabling dynamic promotion globally. The remaining upstream-quality improvement would be eliminating the 32-bit store-TCI subset entirely by finding the exact cross-width/shared-memory state bug.
E14 benchmark correction — 32-bit-only fallback rejected as default
A 3-run benchmark changed the decision:
| fallback | result | open times |
|---|---|---|
| 32-bit-store-only TCI | 2 pass, 1 app-init watchdog | 34.0s, 33.8s |
| blanket store-TCI | 3 pass | 35.3s, 36.6s, 34.3s |
The narrowed fallback is a little faster when it works, but it is not reliable enough: one run failed before controls enabled with an interrupt watchdog. The default is restored to the upstream blanket store-TCI fallback. Keep the 32-bit-only result as a useful diagnostic, not a production fix.
E15/E16: the corruption is a nondeterministic race, not a miscompiled TB
Added runtime-settable, opt-in store gating to isolate the culprit without rebuilds:
?sthash=1|2— compile one hash-parity half of all store TBs (both widths present in the compiled half), interpret the other.?stcw=<bytes>&stcmask=<m>&stcval=<v>— keep every width compiled except the selected widthstcw, which compiles only when(hash(pc) & m) == v. Intended for single-width binary search; logsSTBISECT compile pc=... w=....
Attempted bisection and what it actually showed
sthash=1(compile odd-hash half) -> crash;sthash=2(even half) -> pass. Initial read: a specific pair of TBs.- Width bisection contradicted that: with all other widths compiled, both 8-bit halves crash and both 32-bit halves crash. Identity of the compiled store TBs does not matter; only that enough of them run.
Decisive test: determinism
Re-running a single fixed config exposes nondeterminism:
sthash=2 run 1 -> CRASH
sthash=2 run 2 -> PASS
sthash=2 run 3 -> CRASH
So the earlier single-run "pass" results (including the width fallbacks and the
first sthash=2) were probabilistic, not fixes. The compiled-store corruption
is a nondeterministic race, not a deterministic miscompilation of any specific
TB. E10/E11 (per-TB compiled == TCI, 300k+ comparisons) are fully consistent:
each TB is individually correct; the fault is emergent timing/ordering.
Reliability of the production default
Blanket store-TCI over 5 runs: 4 clean reader-opens, 1 failure — and that 1 was
a cold-boot interrupt-watchdog timeout (MEPC 0x4219a37e), a separate
host-scheduling boot flake, not the 0x4201e210 corruption. So blanket
store-TCI shows 0 corruption crashes here; it suppresses the race in practice
but the underlying race remains.
Revised suspect
Prime suspect is now shared-memory ordering / concurrency between the compiled
guest-store path and another agent touching guest RAM (IO/main thread device
models, dirty-bit/notdirty handling, or the Asyncify store slow-path
DONE_FLAG/UNWINDING re-entrancy flagged in E9), rather than store codegen.
The crash is a corrupted pointer (invalid read at 0x3FCB80xx, region null),
consistent with a torn/stale store of a pointer-sized value under a race.
Instrumentation status
?sthash and ?stcw/stcmask/stcval are opt-in debug tools; the default build
path is unchanged (blanket store-TCI). The narrowed 32-bit-only fallback remains
rejected as a default because it only lowers, not removes, the race probability.
E17: not the store block-replay/redispatch path
Added opt-in ?yieldtrace: the dispatch loop counts compiled-TB "yield/resume"
re-entries, defined as a compiled TB returning non-zero with the same tb_ptr
and do_init == 0 (the DONE_FLAG-driven redispatch/resume path; chains set
do_init = 1, completion sets tb_ptr = 0).
Result over three crashing ?nostoretci runs: all crashed at 0x4201e210,
zero yield re-entries. So the store-specific block-replay-via-redispatch
mechanism (upstream's stated reason for blanket store-TCI) does not fire during
this workload. That path is not the trigger here.
Caveat: Asyncify-based unwinds (coroutine yields) restore the whole C stack and
are transparent to this counter, so a coroutine-mediated replay is not excluded.
But the store DONE_FLAG redispatch path is definitively ruled out.
Consolidated conclusion
The 0x4201e210 corruption is:
- nondeterministic (same config crashes/passes across runs),
- independent of which specific store TBs are compiled,
- gated only by having enough compiled stores of both widths (a timing knob),
- not the store
DONE_FLAGredispatch/replay path.
The consistent remaining explanation is a data race on guest RAM between the
CPU worker (compiled i64.storeN) and another agent that runs concurrently —
main-thread QEMUTimer/device-model callbacks (SPI/SD/GDMA/cache) reading or
writing guest DRAM. TCI's C store path masks it by changing timing/serialization,
not by being semantically different. This matches upstream qemu-wasm shipping
blanket store-TCI rather than a targeted fix.
A real fix is a threading-correctness project: ensure guest-RAM accesses from
device/timer callbacks are properly serialized against vCPU stores (BQL/memory
barriers), or run device memory access on the vCPU thread. Opt-in repro/debug
controls now available: ?nostoretci, ?sthash, ?stcw/stcmask/stcval,
?yieldtrace. Production default remains blanket store-TCI.
E18: not dynamic function-table churn either
Confirmed the threading model: the build uses -pthread -sPROXY_TO_PTHREAD=1 -sASYNCIFY=1, and qemu_thread_create runs the vCPU in its own pthread,
concurrent with the main/IO-loop thread over shared memory. So a genuine
guest-RAM data race is possible.
Tested the most tractable non-thread hypothesis first: each compiled TB adds a
WASM table function and batch-removeFunctions once >500 removals are pending;
more compiled stores (both widths) means more churn, which is timing-dependent.
Added opt-in ?noremove (wasm_set_disable_tb_removal) to leak functions
instead of recycling table indices.
Result: ?nostoretci&noremove crashed 4/4 (Load access faults). Recycling of
WASM table indices is not the corruption source.
Ruled out so far
- specific store-TB identity (E15)
- store
DONE_FLAGblock-replay/redispatch (E17) - WASM function-table index reuse (E18)
Standing conclusion
The corruption is a nondeterministic concurrency issue between the vCPU pthread
(compiled i64.storeN) and the concurrent main/IO thread (timer/device
callbacks, AIO/block completions) over shared guest memory or QEMU softmmu
metadata — the class of bug upstream qemu-wasm avoids with blanket store-TCI.
Confirming the exact racing access needs thread-aware instrumentation (e.g.,
logging IO-thread guest-RAM writes and TLB mutations during reader open) rather
than more store-gating experiments. Production stays on blanket store-TCI.
Opt-in debug knobs now available: ?nostoretci, ?sthash,
?stcw/stcmask/stcval, ?yieldtrace, ?noremove.
E19: IO-thread guest-RAM access is not present in the failing window
Added opt-in ?ramrace instrumentation for off-vCPU guest-DRAM access:
address_space_read/address_space_write- cached read/write slow paths
address_space_map(..., is_write)/ DMA-style map reads- direct mapped write unmap dirtying
- off-vCPU TLB flush entry points
Also added ?ramrace=open, armed by the browser harness immediately before
opening example.txt, to avoid the known cold-boot watchdog flake from always-on
probe overhead.
Results:
- Always-on
?nostoretci&ramracecrash: two off-vCPU full TLB flushes during early boot only; zero off-vCPU DRAM reads/writes/maps/unmaps. - Open-window
?nostoretci&ramrace=opencrash at the original signature (Invalid read 0x3FCB80A6,MEPC 0x4201e210): zero off-vCPU DRAM reads, writes, maps, unmaps, and TLB flushes.
So the specific hypothesis from E18 — concurrent main/IO-thread device or timer callbacks touching guest DRAM during inflate — is ruled out for the corruption window. The remaining race is not an IO-thread guest-RAM accessor.
The fault address is inside the configured ESP32-C3 DRAM aperture
(0x3fc80000..0x3fce0000), so the access fault is still best read as corrupted
vCPU execution state/translation, not an unmapped legitimate pointer.
Revised suspect space after E19
- vCPU-thread-only nondeterminism in the wasm backend around async exits,
helper unwinds, or context/global state restoration that is invisible to the
DONE_FLAGredispatch counter. - A data race on QEMU/TCG metadata or CPUState (not guest RAM), such as async exit/interrupt state observed by the vCPU while compiled stores are active.
- SharedArrayBuffer memory-ordering differences between compiled WASM stores and the C/TCI softmmu path, without another QEMU guest-RAM writer in the same window.
Production default remains blanket store-TCI. The next useful probe is not more IO-thread DRAM logging; it is targeted vCPU-state/TLB-entry instrumentation at the failing translated block or around async exit handling.
E20-E22: state materialization probes
E11 had already identified the first live divergence after arming at open:
pc=0x4201d6ea in utf8NextCodepoint(), where TCI had applied the prologue
sp -= 16 to CPUArchState but compiled-live still had stale env->gpr[2].
The disassembly confirms 0x4201d6d6..0x4201d6e8 is exactly the function
prologue and first branch:
4201d6d6: addi sp,sp,-16
4201d6d8: sw s2,0(sp)
...
4201d6e4: lbu s1,0(s2)
4201d6e8: beqz s1,4201d7a8
4201d6ea: mv s0,a0
E20: no-chain boundary probe
Added opt-in ?stateflush=open, armed by the browser harness immediately before
opening example.txt. It reuses the existing wasm_diff_needs_nochain_at(pc)
app-flash window to force app TBs to return to C without running the expensive
compiled-vs-TCI differential shadow.
Result with ?nostoretci&stateflush=open: original corruption still reproduced
(Invalid read 0x3FCB80A6, MEPC 0x4201e210). Returning to C between app TBs
is not enough; the wasm backend does not materialize CPUArchState merely by
leaving a compiled TB.
E21/E22: blind helper spill attempts are invalid
Tested a direct backend fix candidate: spill TCG global-mem temps before C helper calls so slow-path load/store helpers and exceptions see coherent architectural state.
- spilling all global-mem temps before+after helpers caused an earlier invalid
read around
0x07585d82; - spilling before helpers only caused the same failure;
- narrowing to RISC-V integer GPR globals (
xN/...) still caused immediate invalid writes around0x075f.....
Conclusion: helper-boundary materialization is still the right suspect, but it
must respect TCG liveness/state. Blind backend spills at helper emission can
write stale globals back into CPUArchState and make corruption worse. The fix
belongs in the wasm backend's handling of TCG sync/liveness around calls (or in
targeted generated sync_i32/i64 ops), not in an unconditional helper wrapper
spill.
Revised next step
Trace the TCG liveness/sync ops for the divergent prologue TB and the later
slow-path helper TB: verify whether la_global_sync() marks x2/sp TS_MEM
across the relevant call/branch, and whether the wasm backend emits the
corresponding store. The observed bug is now specifically: compiled globals and
CPUArchState are allowed to diverge, and the wasm backend exposes stale env to
helper/exception-visible code when store TBs stay compiled.
E23-E31: concrete TLB-array corruption
Added FAULTTRACE/TLBTRACE under ?ramrace=open to inspect the exact failing
load path.
Key observations from crashing ?nostoretci&ramrace=open runs:
- The faulting guest address (
0x3fcb8004/0x3fcb80a6) is valid ESP32-C3 DRAM. - TLB fill resolves the page correctly as RAM:
FAULTTRACE tlb_fill vaddr=0x3fcb8000 paddr=0x3fcb8000 asidx=0 prot=0x3
mr=esp32c3.iram ram=1 romd=0 xlat=0x3c000
FAULTTRACE tlb_store index=56 iotlb=0x9c000 addr_page=0x3fcb8000
stored_xlat=0xffffffffc03e4000 phys=0x3fcb8000
FAULTTRACE mmu_lookup1 addr=0x3fcb8000 index=56 maybe_resized=1
tlb_addr=0x3fcb8000 flags=0x0 full_xlat=0xffffffffc03e4000 addend=0xc19d4000
- Before the next byte load on the same page, the compact TLB and full TLB entry are corrupted:
FAULTTRACE mmu_lookup1 addr=0x3fcb8004 index=56 maybe_resized=0
tlb_addr=0x3fcb8200 flags=0x200 full_xlat=0x0 addend=0xc0348000
FAULTTRACE mmio_read addr=0x3fcb8004 xlat_section=0x0 mr=(null) ram=0
0x200 is TLB_MMIO, so the vCPU incorrectly dispatches a DRAM read through
the MMIO/unassigned path. This is why the access is rejected even though the
guest address is in DRAM.
The corruption happens inside the active compiled TB before it returns, so the
post-TB scanner only sees the last good state. The faulting firmware sequence is
in uzlib_uncompress:
4201e18c: lw a3,56(s0)
4201e18e: lw a5,24(s0)
4201e190: lw a4,52(s0)
4201e194: addi a2,a5,1
4201e198: sw a2,24(s0)
4201e19a: add a3,a3,a4
4201e19c: lbu a4,0(a3) # faults after TLB entry was corrupted
4201e1a0: sb a4,0(a5)
The compiled store at 0x4201e198 is now the prime suspect: it likely writes
through a bad host address into the TLB arrays, changing the same page's TLB
entry from RAM to MMIO and zeroing fulltlb[index].xlat_section. That matches
why blanket store-TCI works: the C/TCI store path does not corrupt the TLB arrays.
Current root-cause statement
The bug is no longer just "nondeterministic race". It is specifically:
compiled wasm store fast-path sometimes computes/uses a bad host address and writes into the vCPU softmmu TLB arrays (
CPUTLBEntry/CPUTLBEntryFull), causing the immediately following load from valid DRAM to be treated as MMIO withxlat_section=0.
Still unknown: why the store fast path obtains the bad host address. Candidate
remaining causes: stale/corrupted TLB addend for the store's page, wasm32
address truncation/sign-extension around addends, or an in-TB ordering bug in the
compiled store TLB lookup. The next focused probe is to trace the store fast-path
host address for the 0x4201e198 TB before the i64.store32 executes.
E32-E34: fault-path TB isolation
Tried a narrow compiled-store guard (?tlbguard=open) that forces a store fast
path to the C helper if the computed host address lands inside the vCPU TLB
arrays. It did not suppress the crash, even though the guard range included
fulltlb[56]:
TLBGUARD range lo=0x13e8000 hi=0x353c000
So the corrupting write is not caught as a normal compiled qemu_st fast-path
whose final host pointer falls directly inside that TLB-array range. The
corruption still appears between the TLB fill and the following byte load.
Added ?stpc=LO-HI, an opt-in PC-window override that interprets only
store-containing TBs whose start PC lands in the window, while ?nostoretci
keeps every other store TB compiled.
Results:
- plain
?nostoretci: still crashes at the original signature. ?nostoretci&stpc=0x4201e180-0x4201e220: 3/3 pass (open example.txtabout 46-52s).- wider/narrowed windows around
uzlib_uncompressalso pass as long as they include the fault-path TB start.
This pins the corruption to compiled execution of the uzlib_uncompress TB that
contains the copy-loop sequence around:
4201e18c: lw a3,56(s0)
4201e18e: lw a5,24(s0)
4201e190: lw a4,52(s0)
4201e194: addi a2,a5,1
4201e198: sw a2,24(s0)
4201e19a: add a3,a3,a4
4201e19c: lbu a4,0(a3)
4201e1a0: sb a4,0(a5)
The PC-window override is firmware-specific and not a production fix, but it is now the smallest positive isolation: one compiled store-containing TB in the inflate copy loop is sufficient to corrupt the vCPU TLB entry; interpreting that TB avoids the crash while preserving compiled stores elsewhere.
Next useful step: dump or instrument the generated WASM for this exact TB and
inspect the sw a2,24(s0) / following lbu fast-path sequence. The current
node dump helper only captures early boot TBs, so late-reader TB dumping needs a
browser/open-window trigger or a PC-filtered wasm_dump_module path.