797 lines
39 KiB
Markdown
797 lines
39 KiB
Markdown
# WASM-only reader panic in uzlib inflate
|
||
|
||
Living record of the investigation into the browser-only crash when opening a
|
||
book. Native QEMU (same submodule) does **not** crash, so this is a wasm32 TCG
|
||
backend bug. Update as things are ruled in/out.
|
||
|
||
## Symptom
|
||
|
||
Opening a book in the browser panics; native opens the same book fine.
|
||
|
||
```text
|
||
Invalid read at addr 0x3FCB7999, size 1, region '(null)' size 0x0 owner '(none)' ram=0, reason: rejected
|
||
Guru Meditation Error: Core 0 panic'ed (Load access fault). Exception was unhandled.
|
||
MEPC : 0x4201e210 RA : 0x4201e18c SP : 0x3fcb3810
|
||
```
|
||
|
||
## Root cause (established)
|
||
|
||
`wasm_tci_only_tb = true` in `tcg_out_qemu_st` is **upstream qemu-wasm's own
|
||
design**, not a workaround this project invented. A full diff of our
|
||
`tcg/wasm32/tcg-target.c.inc` against pristine upstream qemu-wasm shows the two
|
||
files are identical except for that one line. Upstream unconditionally forces
|
||
every store-containing TB to the TCI interpreter, so upstream's *compiled*
|
||
(JIT) store path is emitted but never executed — it is effectively unfinished,
|
||
untested dead code, and it corrupts guest state the moment it runs.
|
||
|
||
Mechanism: the wasm32 backend compiles a TB as a chain of wasm blocks guarded
|
||
by `if (BLOCK_PTR <= block_idx)`. When a softmmu slow-path helper yields
|
||
(asyncify unwind for I/O / interrupt), the wasm function returns and re-enters,
|
||
skipping already-executed blocks via `BLOCK_PTR`. This restart is safe to
|
||
replay for loads but not for a compiled store, so upstream punts all store TBs
|
||
to TCI (which has precise instruction-level checkpointing).
|
||
|
||
The real fix is therefore not a port bug hunt but completing upstream's
|
||
compiled store path to be rewind-safe (e.g. defer the memory commit past the
|
||
TB's yield points / buffer stores until block commit), then dropping the
|
||
blanket TCI. Large TCG-backend task; ~15 min per browser iteration; no native
|
||
repro.
|
||
|
||
## Established facts
|
||
|
||
- **Fault site (via same-version local ELF addr2line, not byte-identical):**
|
||
`uzlib_uncompress` (tinflate.c:613) ← `tinf_inflate_block_data` (tinflate.c:476)
|
||
← `InflateReader::read` ← `FontDecompressor::decompressGroup` ← `prewarmCache`.
|
||
"Indexing" = the reader inflating compressed glyph groups from flash into a
|
||
`malloc`'d temp buffer.
|
||
- **Input is static flash data**, independent of SD. The compressed font bank is
|
||
in the firmware image; the SD file only selects *which* glyph groups get
|
||
prewarmed (the characters present in the text).
|
||
- **Native renders the reader** ("chapter one") with no fault — proven with
|
||
`_scratch/native-reader-repro.py` + QMP screendumps. Confirms the firmware,
|
||
the font data, and the decompression logic are all correct.
|
||
- **Fault address is non-deterministic:** `0x3FCB7999` and `0x3FCB7997` across
|
||
runs. A static machine-map boundary would fault at a fixed address. Drift ⇒ a
|
||
computed pointer is slightly wrong, i.e. runtime/codegen, not the memory map.
|
||
- **Rejected region is bogus:** `region '(null)' size 0x0 owner '(none)' ram=0`.
|
||
The address lands in mapped DRAM (`0x3fc80000`+`0x60000`) but the wasm softmmu
|
||
TLB routed it to the IO slow path with a zero-size section. Hallmark of a bad
|
||
address/TLB result from a JIT-compiled TB.
|
||
- **Menus work.** Home/file-browser render crisp text through the same font
|
||
system and the same TLB fast path, so the TLB codegen is not broadly broken —
|
||
the bug is pattern/data specific in a hot, JIT-promoted TB.
|
||
- **User observation:** firmware-written crash logs open fine; prepopulated
|
||
files do not. Consistent with data-dependence — a crash log's character set
|
||
prewarms different glyph groups than lorem-ipsum, avoiding the group whose
|
||
compressed byte pattern trips the miscompiled path.
|
||
|
||
## Architecture notes (why JIT is implicated)
|
||
|
||
- `tcg/wasm32.c`: TBs run on the TCI interpreter until executed
|
||
`INSTANTIATE_NUM` times, then get compiled to native WASM (`instantiate_wasm`)
|
||
and run as added table functions. The inflate hot loop gets promoted.
|
||
- `WASM_TCI_ONLY_ENTRY` forces a TB to stay on TCI forever;
|
||
`wasm_tci_only_tb` (per-TB, reset each translation) requests that.
|
||
- Dynamic promotion must stay on for usable performance (all-TCI starves the
|
||
guest watchdog), so a permanent fix must correct codegen, not disable JIT.
|
||
|
||
## Hypotheses
|
||
|
||
| # | Hypothesis | Status |
|
||
|---|-----------|--------|
|
||
| H1 | A store-containing TB is miscompiled (the cleanup removed `wasm_tci_only_tb = true` from `tcg_out_qemu_st`); the compiled TB corrupts a pointer a later load dereferences | **confirmed** (E1) |
|
||
| H2 | Compiled **load** codegen bug (`tcg_wasm_out_qemu_ld` / `_direct`) | untested |
|
||
| H3 | TLB fast-path codegen edge case (`tcg_wasm_out_tlb_load`) hit only by certain access patterns | untested |
|
||
| H4 | asyncify/fiber unwind corrupts vCPU state mid-inflate | unlikely (inflate makes no async import calls) |
|
||
| H5 | mimalloc / linear-memory / threads issue | unlikely (native heap layout differs, no crash) |
|
||
|
||
## Experiments
|
||
|
||
### E1 — force store-containing TBs to TCI (tests H1) ✅ CONFIRMED
|
||
Re-added `wasm_tci_only_tb = true;` to `tcg_out_qemu_st` in
|
||
`tcg/wasm32/tcg-target.c.inc`.
|
||
|
||
- Result: **panic gone.** `make browser-test` reaches the reader and renders a
|
||
correct page (black lorem-ipsum on white, footer `80% example 1/3 33%`). No
|
||
reject, no `Guru Meditation`, 1 boot, 0 watchdog panics. New refresh frames
|
||
(`dtm2_white=0.891`) appear past the old `0.960` stall.
|
||
- **Conclusion: the compiled `qemu_st` codegen is the bug.** Forcing store TBs
|
||
to TCI is a working *workaround*, not the fix — every store-containing TB
|
||
then runs interpreted (broad perf cost; prior attempts reported watchdog
|
||
starvation under sustained load, though this finite-burst run was clean).
|
||
- The reader first paints a correct BW page (black text on white), then a
|
||
follow-up **grayscale** refresh pass inverts it into muddled white-on-black
|
||
text. That second issue is real and still open (see ISSUES.md "Reader renders
|
||
inverted grayscale"); it is independent of this panic fix. The `browser-test`
|
||
wait-for-stable happened to screenshot the initial white page.
|
||
|
||
### Caveat on E1's specificity
|
||
Forcing store-*containing* TBs to TCI drops the whole TB (its loads, stores,
|
||
and arithmetic) to the interpreter. So the bug is *in a TB that also has a
|
||
store*, not provably the store op. `tcg_wasm_out_qemu_st_direct` and
|
||
`tcg_wasm_out_qemu_ld_direct` are byte-for-byte symmetric and look correct in
|
||
isolation, which points away from the direct emit and toward the store path's
|
||
surrounding block/rewind/helper scaffolding in `tcg_wasm_out_qemu_st`, or an
|
||
interaction with a co-located op.
|
||
|
||
### E2 — crash is icount-independent; loads-TCI also masks it (inconclusive on op)
|
||
- Removed the store flag and ran at the new `noicount` default: **same crash**
|
||
(`Invalid read at 0x3FCB80A5`, drifts, `MEPC 0x4201e210`). So icount timing
|
||
is not the trigger.
|
||
- Moved the flag to `tcg_out_qemu_ld` (loads TCI, stores compiled): **passed**
|
||
(33s). Not decisive — load-containing TBs are a superset of store-containing
|
||
TBs, so forcing loads to TCI drops nearly every TB (including all store TBs)
|
||
to the interpreter. It masks the bug the same way store-TCI does.
|
||
- **Takeaway:** whole-TB TCI cannot isolate the offending op. Precise
|
||
localization needs per-op TCI infrastructure (backend is whole-TB) or reading
|
||
the generated wasm for one miscompiled TB. Large task; reader already works at
|
||
~31s with the store flag, so this is headroom, not a blocker.
|
||
|
||
### E3 — PC-windowed bisection: not a single TB, it's aggregate
|
||
Added a compile-window to `tcg_out_qemu_st` (`s->gen_tb->pc` gate): store-TBs
|
||
inside `[LO,HI)` compile, all others stay TCI. Each run rebuilt + browser-tested.
|
||
|
||
| window | region | result |
|
||
|--------|--------|--------|
|
||
| 0x4201d000–0x4201f800 | uzlib inflate + FontDecompressor | **pass** |
|
||
| 0x40000000–0x40400000 | ROM (memcpy/memset) + IRAM (tlsf/malloc) | **pass** |
|
||
| 0x42000000–0x420d0000 | flash text lower half (prewarm callers) | **pass** |
|
||
| 0x420d0000–0x421a0000 | flash text upper half | **pass** |
|
||
| (none / all store-TBs compiled) | everything | **crash** |
|
||
|
||
Every region compiles safely *in isolation*; only the full set corrupts. So the
|
||
bug is **not a localizable miscompiled op/TB** — it is an aggregate effect that
|
||
appears once enough compiled store-TBs run together. The faulting load is
|
||
`d->dict_ring[d->lzOff]` (tinflate.c:476) off a stack-resident inflate struct,
|
||
i.e. some store clobbers a *neighbouring* stack/heap pointer, and which store
|
||
depends on timing. This matches a **rewind/replay or TB-instance-management
|
||
race** in the compiled restartable-block store path, not a codegen typo — and it
|
||
is exactly why upstream forces every store-TB to TCI. Fixing it is architectural
|
||
(rewind/instance-safe compiled stores), not a one-op correction.
|
||
|
||
### E4-E7 — deterministic minimal repro, and what the bug is NOT
|
||
With a per-op gate in `tcg_out_qemu_st` (compile only store-TBs matching a
|
||
criterion, TCI the rest) plus targeted instrumentation, each run rebuilt +
|
||
browser-tested against the exact 1.4.1 firmware:
|
||
|
||
| experiment | change | result |
|
||
|-----------|--------|--------|
|
||
| E4 lifecycle | never `removeFunction` retired TB instances (pin all) | **crash** |
|
||
| E5 helper-only | force every store through the softmmu helper (no direct store codegen) | **crash** |
|
||
| E6 width: byte only | compile TBs whose store is 8-bit | **pass** |
|
||
| E6 width: 16 only | 16-bit | **pass** |
|
||
| E6 width: 32 only | 32-bit | **pass** |
|
||
| E6 width: byte+16 | | **pass** |
|
||
| E6 width: 16+32 | | **pass** |
|
||
| E6 width: **byte+32** | | **crash** |
|
||
| E7 bytewise-32 | emit 32-bit store as four 8-bit stores, compile all | **crash** |
|
||
| precheck | skip store helper on rewind when DONE_FLAG already set | **crash** |
|
||
| nochain | `-d nochain` (disable TB chaining), compile all stores | **crash** |
|
||
| determinism | rerun the crashing all-compiled build 3x | **crash x3, same gen 16, same MEPC** |
|
||
|
||
**Established:** the crash is *deterministic* (same point every run) but depends
|
||
on the exact compiled-TB set. The minimal reproducer is a byte-store TB **and**
|
||
a 32-bit-store TB compiled together; it is *counter-monotonic* — additionally
|
||
compiling 16-bit-store TBs hides it.
|
||
|
||
**Ruled out:** single miscompiled TB (E1-E3), a specific memory region
|
||
(inflate/ROM/IRAM/flash halves), TB chaining (nochain), instance/function
|
||
lifecycle & `removeFunction` reuse (E4 pin), the store *data* codegen itself
|
||
(E5 helper-only and E7 bytewise both crash), and store double-execution on
|
||
rewind (precheck to skip the replayed helper still crashes).
|
||
|
||
**Correction:** the rule is monotonic, not counter-monotonic — `{8,16,32}`
|
||
(original) also crashes, so 16-bit is irrelevant. The rule is simply: **crash iff
|
||
an 8-bit-store TB and a 32-bit-store TB are both compiled.**
|
||
|
||
### E8 — the 8-bit requirement is aggregate, not one TB
|
||
Compiled all 32-bit-store TBs, then compiled 8-bit-store TBs only inside one
|
||
region at a time:
|
||
|
||
| 8-bit region compiled (with all 32-bit compiled) | result |
|
||
|---|---|
|
||
| flash lower 0x42000000-0x420d0000 | **pass** |
|
||
| ROM+IRAM 0x40000000-0x40400000 (incl. memset) | **pass** |
|
||
| flash upper 0x420d0000-0x421a0000 | **pass** |
|
||
| all 8-bit (no region gate) | **crash** |
|
||
|
||
Those three regions cover essentially all executable memory, yet each passes
|
||
while their union crashes. So the 8-bit side is **aggregate** too — there is no
|
||
single interacting 8-bit TB. Combined with E1-E3 (region-aggregate) this proves
|
||
the corruption is an **emergent property of running many compiled store-TBs**,
|
||
which whole-TB TCI gating (the backend's only control) cannot dissect. The
|
||
remaining paths are generated-wasm inspection of the BLOCK_PTR/register restart
|
||
sequence, or building per-op TCI. A deterministic reproducer exists (remove
|
||
`wasm_tci_only_tb`, open example.txt, faults at MEPC 0x4201e210, gen 16).
|
||
|
||
### E9 — static analysis of generated wasm (dump + wasm2wat)
|
||
Added a node-only dump hook (`wasm_dump_module`, tcg.c finalize) that writes the
|
||
first compiled store-TB and load-TB modules; `_scratch/dump-tbs.mjs` boots the
|
||
firmware under node far enough to emit them. `wasm2wat --enable-threads` gives
|
||
readable IR.
|
||
|
||
**Finding 1 (real bug, fixed-then-shelved):** the ld/st slow paths detect an
|
||
asyncify yield with the thread-global `DONE_FLAG` (ctx+20), while the general
|
||
call path uses `UNWINDING` (ctx+28), which is reset before each call and set
|
||
*only* by the coroutine switch. `DONE_FLAG` is set by *any* nested ld/st helper,
|
||
so a nested guest access inside a store/load helper that then yields makes the
|
||
outer op falsely read "completed" and swallow the unwind. Converting ld/st to
|
||
the `UNWINDING` mechanism is the correct fix. **But** rebuilding with compiled
|
||
stores + the UNWINDING fix still crashes identically (gen 16, same MEPC), so the
|
||
store helpers on this path never actually yield — this bug is latent here. Change
|
||
shelved (not shipped) to avoid touching the working compiled-load path without a
|
||
green test; recorded for a future upstream fix.
|
||
|
||
**Finding 2 (reframes the root cause):** combine E5 with the load result. E5
|
||
routes *every* store through the C softmmu helper (correct memory write) and
|
||
still crashes; loads share the identical restart/block structure and work. So
|
||
the corruption is **not** in the store data codegen, the restart machinery, or
|
||
the yield handling. What remains is that a store *writes* guest memory using an
|
||
address/value taken from guest registers (wasm globals). A compiled TB that
|
||
computes a wrong register value corrupts memory only through a store — a wrong
|
||
address fed to a *load* just reads harmless garbage. This matches every prior
|
||
result: emergent (depends which ops precede a store), deterministic, and
|
||
non-localizable by store region or width. The remaining hunt is a
|
||
register-producing op whose compiled codegen diverges from TCI, exposed
|
||
destructively via stores — which needs a compiled-vs-TCI per-TB register
|
||
differential, not more whole-TB gating.
|
||
|
||
Tooling retained: `wasm_dump_module` + `wasm_tb_had_store/load` tags +
|
||
`_scratch/dump-tbs.mjs`.
|
||
|
||
### Next options
|
||
1. **Ship E1 as a guarded workaround** — reader works now; measure real
|
||
navigation/page-turn perf before judging it unacceptable.
|
||
2. **Localize precisely** — bisect by making the TCI trigger conditional
|
||
(store width; presence of the miss/helper path; TB size) to shrink the
|
||
suspect set, then read the generated wasm for that TB.
|
||
|
||
## Diagnostics currently in the tree (uncommitted)
|
||
|
||
- `system/memory.c`: reject log prints region size/owner/ram to identify the
|
||
faulting region.
|
||
- `hw/display/xteink_x3_eink.c`: `[eink] refresh` popcount trace (per-plane
|
||
white ratio) — also used for the separate grayscale-rendering issue.
|
||
|
||
## Reproduction
|
||
|
||
- **WASM:** `make browser-test` (the test does 4 confirm presses: Browse Files →
|
||
Test → highlight → Open). Panic appears in `_scratch/browser-raw-sd.log`.
|
||
- **Native (control, no panic):** `nix develop -c python3 _scratch/native-reader-repro.py`,
|
||
then inspect `_scratch/native-step4-reader.ppm`.
|
||
|
||
## Ruled out
|
||
|
||
- Not SD / filesystem: fixed separately (raw FAT32, SSI-SD token, CMD13); the
|
||
file reads correctly and "Indexing" begins.
|
||
- Not firmware / font data: native decompresses and renders identical input.
|
||
- Not the machine memory map: fault address drifts; native map is the same.
|
||
|
||
## E10: compiled-vs-TCI per-TB differential (opt-in `?tcgdiff`) — store codegen exonerated
|
||
|
||
Built a per-TB differential checker to settle whether single-TB compiled
|
||
store codegen diverges from TCI. It runs authoritative TCI first, snapshots
|
||
its result, rewinds guest state, runs the compiled TB as a disposable shadow,
|
||
compares, then restores the TCI result. Opt-in via `?tcgdiff`
|
||
(`wasm_diff_enable()` from `web/app.js` `preRun`); normal/browser execution is
|
||
unchanged.
|
||
|
||
Making the comparison *valid* — not the codegen — was the whole task. Five
|
||
harness defects each masqueraded as a "register divergence" until fixed:
|
||
|
||
1. **Execution order** — running the compiled shadow first contaminated TCI's
|
||
machine state. Fixed: TCI first, shadow second, restore TCI.
|
||
2. **Partial-TB comparison** — helper/slow-path TBs don't reach `exit_tb`.
|
||
Fixed with a `done_flag == 2` sentinel: only fully-inline store TBs that
|
||
complete (`done_flag == 2 && tb_ptr == 0`) are compared.
|
||
3. **Negative-offset state** — the TB-entry interrupt check reads `env-8`
|
||
(`CPUNegativeOffsetState`, before `CPUArchState`). Snapshot/restore now
|
||
covers the full `CPUNegativeOffsetState`.
|
||
4. **Deep TLB state** — `CPUTLB.f[].table` / `d[].fulltlb` are separately
|
||
allocated; TCI warms the soft-TLB via its ld/st helper path, so the shadow
|
||
must start from an identical TLB. Now deep-copied per MMU mode.
|
||
5. **Async exit-flag race** — the IO thread asynchronously sets
|
||
`icount_decr.high = -1` via `cpu_exit()`. TCI can take the TB-entry
|
||
interrupt exit while the restored shadow sees a cleared flag and runs fully.
|
||
These are unavoidable concurrency races, so comparisons where the force-exit
|
||
flag was set around the TCI run are skipped.
|
||
|
||
Only side-effect-safe candidates are compared: store-containing TBs with no
|
||
general helper call (`wasm_tb_had_helper`), restricted to the app-flash PC
|
||
window. Live dispatch stays TCI for all store TBs; the compiled function runs
|
||
only as a shadow.
|
||
|
||
**Result: 160,000+ store-TB comparisons during a full reader-open, zero
|
||
register or DRAM divergence.** Single-TB compiled store codegen is bit-for-bit
|
||
identical to TCI. This exonerates per-TB store codegen and confirms the
|
||
corruption is emergent across TBs (consistent with E1–E8). A per-TB
|
||
differential is structurally blind to a multi-TB interaction bug; catching it
|
||
needs a whole-execution differential (two full machines stepped in lockstep
|
||
from reset), not more per-TB work.
|
||
|
||
The `?tcgdiff` harness is expensive (full DRAM + TLB snapshot and double
|
||
execution per candidate; heap grows to ~1 GiB, frame gen drops) but is
|
||
debug-only and does not affect the production store-TCI path.
|
||
|
||
## E11: whole-execution differential (opt-in `?tcgdiff2`) — live compiled vs TCI shadow
|
||
|
||
Mode 2 inverts E10: the **compiled** path runs LIVE and authoritative (the real
|
||
buggy path, so corruption *accumulates* across TBs exactly as in the crash),
|
||
while a per-TB **TCI shadow** — restored from the pre-snapshot — is the
|
||
known-good reference. Store-TCI is disabled under this mode so compiled stores
|
||
actually execute. Per eligible store TB: snapshot pre (env, neg, deep TLB,
|
||
DRAM) -> run compiled live, capture full compiled-post -> rewind -> run TCI
|
||
shadow -> compare -> restore compiled-post -> continue. Opt-in via `?tcgdiff2`.
|
||
|
||
### Result
|
||
- **153,000+ live-compiled-vs-TCI comparisons with accumulation, zero
|
||
divergence** (bootloader region). Combined with E10 (160k reader-path), inline
|
||
compiled stores are correct even under live cross-TB accumulation.
|
||
- Revealed a large blind spot: ~90% of store TBs complete via the store
|
||
**slow-path** (`done_flag == 1`, TLB-miss -> ld/st helper), not the pure
|
||
inline path (`done_flag == 2`). Mode 2 was extended to compare slow-path
|
||
completions too.
|
||
|
||
### Why it cannot reach the reader corruption (impractical in-browser)
|
||
The 2x execution + full DRAM (0x60000) + deep-TLB snapshot per eligible TB is
|
||
~1000x slower than normal. Consequences:
|
||
- The guest **interrupt watchdog** fires first: gating comparison to app-flash
|
||
produced an `Interrupt wdt timeout` panic (MEPC 0x40380264, IRAM idle) during
|
||
cold boot — a harness-overhead artifact, not the corruption Guru Meditation
|
||
(MEPC 0x4201e210).
|
||
- Ungated, it reached only ~153k comparisons (bootloader, gen 2) before the
|
||
1100s timeout. Full reader-open is millions of TBs away. Defanging the
|
||
watchdog would not help; raw throughput is the wall.
|
||
|
||
### Narrowed suspects (what the differential structurally cannot cheaply test)
|
||
1. **Store slow-path** (TLB miss -> `helper_stN_mmu` -> Asyncify unwind/resume),
|
||
the ~90% of stores mode 2 excluded by the `done_flag == 2` gate. Aligns with
|
||
E9's `DONE_FLAG` (thread-global) vs `UNWINDING` (per-call) finding.
|
||
2. **Helper-containing store TBs**, excluded by the side-effect-safety filter.
|
||
3. **8/32-bit store-width interaction** (E1-E8's "needs both widths").
|
||
|
||
### Decision
|
||
In-browser whole-execution differential is retired as a repro path: it proves
|
||
live-accumulation correctness of inline stores but cannot reach the crash. Next
|
||
instrument is a deterministic, framework-free **slow-path store differential
|
||
unit test** (no boot, no watchdog, no perf wall) targeting suspect (1)/(3).
|
||
Option A (sequential deterministic two-pass) remains a documented future option
|
||
if a full-execution comparison is later required.
|
||
|
||
Both `?tcgdiff` (E10) and `?tcgdiff2` (E11) are retained as opt-in tools;
|
||
production store-TCI path is unchanged.
|
||
|
||
## E12/E13: normal-speed compiled stores and file-open-armed E11
|
||
|
||
### `?nostoretci` probe
|
||
Added opt-in `?nostoretci` to disable the store-TCI workaround without enabling
|
||
any differential overhead. Result: the exact original corruption still
|
||
reproduces at normal speed during reader open:
|
||
|
||
- `Load access fault`
|
||
- `MEPC 0x4201e210`
|
||
- `MTVAL 0x3fcb8004` / invalid read near the inflate/text buffer region
|
||
|
||
So store-TCI is still required in production.
|
||
|
||
### `?tcgdiff2=open`
|
||
Arming E11 from boot was too slow or perturbed app init. Added a browser-test
|
||
hook so `?tcgdiff2=open` arms `wasm_diff_enable2()` immediately before the test
|
||
presses Confirm to open `example.txt`. Boot and app init therefore run on the
|
||
normal production path, and compiled stores become live only over the repro
|
||
interval.
|
||
|
||
With the all-app-flash window (`0x42000000..0x42200000`), the first observed
|
||
mode-2 divergence after arming is register state, not DRAM:
|
||
|
||
```text
|
||
TCGDIFF pc=0x4201d6ea gpr[2] compiled=0x3fcb42f0 tci=0x3fcb42e0
|
||
pre-ra=0x42029374 pre-sp=0x3fcb42f0
|
||
```
|
||
|
||
This is `utf8NextCodepoint()`: TCI applies the stack-frame prologue (`sp -= 16`)
|
||
while the compiled-live path has stale `sp` in `CPUArchState`. The later crash
|
||
is still the known invalid read at `MEPC 0x4201e210`.
|
||
|
||
### Current interpretation
|
||
E10 remains green because it forces no-chain and runs TCI live, then compiled as
|
||
a disposable shadow. E11 is the real compiled-live path and exposes state that is
|
||
stale across compiled execution/chaining boundaries. The remaining suspect is no
|
||
longer individual store codegen; it is **compiled TB state materialization**:
|
||
WASM register globals and `CPUArchState` can diverge across real compiled-live
|
||
entry/exit/chaining, and store-TCI masks the bad store-containing TBs by routing
|
||
them through TCI.
|
||
|
||
A tested broad fix candidate — making store TBs return to C instead of using the
|
||
same-module `BLOCK_PTR` chain — caused an app-init interrupt-watchdog timeout,
|
||
so it is not viable as-is. A narrower fix should target state materialization
|
||
(flush/reload live guest register state at compiled TB boundaries or before
|
||
helper/exception-visible points) rather than disabling chaining wholesale.
|
||
|
||
## E14: narrowed production fix — only 32-bit store TBs stay TCI
|
||
|
||
The blanket upstream qemu-wasm fallback interpreted every store-containing TB.
|
||
The width interaction tests show that is broader than necessary:
|
||
|
||
- all stores compiled (`?nostoretci`) -> exact crash remains (`MEPC 0x4201e210`,
|
||
invalid read near `0x3FCB80xx`)
|
||
- only 8-bit qemu store TBs interpreted -> reader-open passes
|
||
- only 32-bit qemu store TBs interpreted -> reader-open passes
|
||
|
||
Production now keeps only **32-bit qemu store TBs** on TCI; 8/16/64-bit store
|
||
TBs compile normally. This breaks the bad 8-bit/32-bit compiled-store
|
||
interaction while preserving more dynamic WASM promotion than the blanket
|
||
store-TCI fallback. `?nostoretci` remains as the normal-speed all-compiled
|
||
negative-control repro and still crashes at the original panic.
|
||
|
||
Validation:
|
||
|
||
```text
|
||
normal URL: browser raw-sd test passed, open example.txt took 36.8s
|
||
?nostoretci: Load access fault, MEPC 0x4201e210, invalid read 0x3FCB80A5
|
||
```
|
||
|
||
This fixes the browser panic without disabling dynamic promotion globally. The
|
||
remaining upstream-quality improvement would be eliminating the 32-bit store-TCI
|
||
subset entirely by finding the exact cross-width/shared-memory state bug.
|
||
|
||
### E14 benchmark correction — 32-bit-only fallback rejected as default
|
||
|
||
A 3-run benchmark changed the decision:
|
||
|
||
| fallback | result | open times |
|
||
| --- | --- | --- |
|
||
| 32-bit-store-only TCI | 2 pass, 1 app-init watchdog | 34.0s, 33.8s |
|
||
| blanket store-TCI | 3 pass | 35.3s, 36.6s, 34.3s |
|
||
|
||
The narrowed fallback is a little faster when it works, but it is not reliable
|
||
enough: one run failed before controls enabled with an interrupt watchdog. The
|
||
default is restored to the upstream blanket store-TCI fallback. Keep the
|
||
32-bit-only result as a useful diagnostic, not a production fix.
|
||
|
||
## E15/E16: the corruption is a nondeterministic race, not a miscompiled TB
|
||
|
||
Added runtime-settable, opt-in store gating to isolate the culprit without
|
||
rebuilds:
|
||
|
||
- `?sthash=1|2` — compile one hash-parity half of *all* store TBs (both widths
|
||
present in the compiled half), interpret the other.
|
||
- `?stcw=<bytes>&stcmask=<m>&stcval=<v>` — keep every width compiled except the
|
||
selected width `stcw`, which compiles only when `(hash(pc) & m) == v`. Intended
|
||
for single-width binary search; logs `STBISECT compile pc=... w=...`.
|
||
|
||
### Attempted bisection and what it actually showed
|
||
- `sthash=1` (compile odd-hash half) -> crash; `sthash=2` (even half) -> pass.
|
||
Initial read: a specific pair of TBs.
|
||
- Width bisection contradicted that: with all other widths compiled, **both**
|
||
8-bit halves crash and **both** 32-bit halves crash. Identity of the compiled
|
||
store TBs does not matter; only that enough of them run.
|
||
|
||
### Decisive test: determinism
|
||
Re-running a single fixed config exposes nondeterminism:
|
||
|
||
```text
|
||
sthash=2 run 1 -> CRASH
|
||
sthash=2 run 2 -> PASS
|
||
sthash=2 run 3 -> CRASH
|
||
```
|
||
|
||
So the earlier single-run "pass" results (including the width fallbacks and the
|
||
first `sthash=2`) were probabilistic, not fixes. **The compiled-store corruption
|
||
is a nondeterministic race**, not a deterministic miscompilation of any specific
|
||
TB. E10/E11 (per-TB compiled == TCI, 300k+ comparisons) are fully consistent:
|
||
each TB is individually correct; the fault is emergent timing/ordering.
|
||
|
||
### Reliability of the production default
|
||
Blanket store-TCI over 5 runs: 4 clean reader-opens, 1 failure — and that 1 was
|
||
a **cold-boot interrupt-watchdog timeout** (`MEPC 0x4219a37e`), a separate
|
||
host-scheduling boot flake, not the `0x4201e210` corruption. So blanket
|
||
store-TCI shows 0 corruption crashes here; it suppresses the race in practice
|
||
but the underlying race remains.
|
||
|
||
### Revised suspect
|
||
Prime suspect is now shared-memory ordering / concurrency between the compiled
|
||
guest-store path and another agent touching guest RAM (IO/main thread device
|
||
models, dirty-bit/notdirty handling, or the Asyncify store slow-path
|
||
`DONE_FLAG`/`UNWINDING` re-entrancy flagged in E9), rather than store codegen.
|
||
The crash is a corrupted *pointer* (invalid read at `0x3FCB80xx`, region null),
|
||
consistent with a torn/stale store of a pointer-sized value under a race.
|
||
|
||
### Instrumentation status
|
||
`?sthash` and `?stcw/stcmask/stcval` are opt-in debug tools; the default build
|
||
path is unchanged (blanket store-TCI). The narrowed 32-bit-only fallback remains
|
||
rejected as a default because it only lowers, not removes, the race probability.
|
||
|
||
## E17: not the store block-replay/redispatch path
|
||
|
||
Added opt-in `?yieldtrace`: the dispatch loop counts compiled-TB "yield/resume"
|
||
re-entries, defined as a compiled TB returning non-zero with the *same* `tb_ptr`
|
||
and `do_init == 0` (the `DONE_FLAG`-driven redispatch/resume path; chains set
|
||
`do_init = 1`, completion sets `tb_ptr = 0`).
|
||
|
||
Result over three crashing `?nostoretci` runs: **all crashed at `0x4201e210`,
|
||
zero yield re-entries**. So the store-specific block-replay-via-redispatch
|
||
mechanism (upstream's stated reason for blanket store-TCI) does not fire during
|
||
this workload. That path is not the trigger here.
|
||
|
||
Caveat: Asyncify-based unwinds (coroutine yields) restore the whole C stack and
|
||
are transparent to this counter, so a coroutine-mediated replay is not excluded.
|
||
But the store `DONE_FLAG` redispatch path is definitively ruled out.
|
||
|
||
### Consolidated conclusion
|
||
The `0x4201e210` corruption is:
|
||
- nondeterministic (same config crashes/passes across runs),
|
||
- independent of which specific store TBs are compiled,
|
||
- gated only by having enough compiled stores of both widths (a timing knob),
|
||
- not the store `DONE_FLAG` redispatch/replay path.
|
||
|
||
The consistent remaining explanation is a **data race on guest RAM** between the
|
||
CPU worker (compiled `i64.storeN`) and another agent that runs concurrently —
|
||
main-thread `QEMUTimer`/device-model callbacks (SPI/SD/GDMA/cache) reading or
|
||
writing guest DRAM. TCI's C store path masks it by changing timing/serialization,
|
||
not by being semantically different. This matches upstream qemu-wasm shipping
|
||
blanket store-TCI rather than a targeted fix.
|
||
|
||
A real fix is a threading-correctness project: ensure guest-RAM accesses from
|
||
device/timer callbacks are properly serialized against vCPU stores (BQL/memory
|
||
barriers), or run device memory access on the vCPU thread. Opt-in repro/debug
|
||
controls now available: `?nostoretci`, `?sthash`, `?stcw/stcmask/stcval`,
|
||
`?yieldtrace`. Production default remains blanket store-TCI.
|
||
|
||
## E18: not dynamic function-table churn either
|
||
|
||
Confirmed the threading model: the build uses `-pthread -sPROXY_TO_PTHREAD=1
|
||
-sASYNCIFY=1`, and `qemu_thread_create` runs the vCPU in its own pthread,
|
||
concurrent with the main/IO-loop thread over shared memory. So a genuine
|
||
guest-RAM data race is possible.
|
||
|
||
Tested the most tractable non-thread hypothesis first: each compiled TB adds a
|
||
WASM table function and batch-`removeFunction`s once >500 removals are pending;
|
||
more compiled stores (both widths) means more churn, which is timing-dependent.
|
||
Added opt-in `?noremove` (`wasm_set_disable_tb_removal`) to leak functions
|
||
instead of recycling table indices.
|
||
|
||
Result: `?nostoretci&noremove` crashed 4/4 (Load access faults). Recycling of
|
||
WASM table indices is **not** the corruption source.
|
||
|
||
### Ruled out so far
|
||
- specific store-TB identity (E15)
|
||
- store `DONE_FLAG` block-replay/redispatch (E17)
|
||
- WASM function-table index reuse (E18)
|
||
|
||
### Standing conclusion
|
||
The corruption is a nondeterministic concurrency issue between the vCPU pthread
|
||
(compiled `i64.storeN`) and the concurrent main/IO thread (timer/device
|
||
callbacks, AIO/block completions) over shared guest memory or QEMU softmmu
|
||
metadata — the class of bug upstream qemu-wasm avoids with blanket store-TCI.
|
||
Confirming the exact racing access needs thread-aware instrumentation (e.g.,
|
||
logging IO-thread guest-RAM writes and TLB mutations during reader open) rather
|
||
than more store-gating experiments. Production stays on blanket store-TCI.
|
||
|
||
Opt-in debug knobs now available: `?nostoretci`, `?sthash`,
|
||
`?stcw/stcmask/stcval`, `?yieldtrace`, `?noremove`.
|
||
|
||
|
||
## E19: IO-thread guest-RAM access is not present in the failing window
|
||
|
||
Added opt-in `?ramrace` instrumentation for off-vCPU guest-DRAM access:
|
||
|
||
- `address_space_read` / `address_space_write`
|
||
- cached read/write slow paths
|
||
- `address_space_map(..., is_write)` / DMA-style map reads
|
||
- direct mapped write unmap dirtying
|
||
- off-vCPU TLB flush entry points
|
||
|
||
Also added `?ramrace=open`, armed by the browser harness immediately before
|
||
opening `example.txt`, to avoid the known cold-boot watchdog flake from always-on
|
||
probe overhead.
|
||
|
||
Results:
|
||
|
||
- Always-on `?nostoretci&ramrace` crash: two off-vCPU full TLB flushes during
|
||
early boot only; zero off-vCPU DRAM reads/writes/maps/unmaps.
|
||
- Open-window `?nostoretci&ramrace=open` crash at the original signature
|
||
(`Invalid read 0x3FCB80A6`, `MEPC 0x4201e210`): **zero** off-vCPU DRAM
|
||
reads, writes, maps, unmaps, and TLB flushes.
|
||
|
||
So the specific hypothesis from E18 — concurrent main/IO-thread device or timer
|
||
callbacks touching guest DRAM during inflate — is ruled out for the corruption
|
||
window. The remaining race is not an IO-thread guest-RAM accessor.
|
||
|
||
The fault address is inside the configured ESP32-C3 DRAM aperture
|
||
(`0x3fc80000..0x3fce0000`), so the access fault is still best read as corrupted
|
||
vCPU execution state/translation, not an unmapped legitimate pointer.
|
||
|
||
### Revised suspect space after E19
|
||
|
||
- vCPU-thread-only nondeterminism in the wasm backend around async exits,
|
||
helper unwinds, or context/global state restoration that is invisible to the
|
||
`DONE_FLAG` redispatch counter.
|
||
- A data race on QEMU/TCG metadata or CPUState (not guest RAM), such as async
|
||
exit/interrupt state observed by the vCPU while compiled stores are active.
|
||
- SharedArrayBuffer memory-ordering differences between compiled WASM stores and
|
||
the C/TCI softmmu path, without another QEMU guest-RAM writer in the same
|
||
window.
|
||
|
||
Production default remains blanket store-TCI. The next useful probe is not more
|
||
IO-thread DRAM logging; it is targeted vCPU-state/TLB-entry instrumentation at
|
||
the failing translated block or around async exit handling.
|
||
|
||
|
||
## E20-E22: state materialization probes
|
||
|
||
E11 had already identified the first live divergence after arming at open:
|
||
`pc=0x4201d6ea` in `utf8NextCodepoint()`, where TCI had applied the prologue
|
||
`sp -= 16` to `CPUArchState` but compiled-live still had stale `env->gpr[2]`.
|
||
The disassembly confirms `0x4201d6d6..0x4201d6e8` is exactly the function
|
||
prologue and first branch:
|
||
|
||
```text
|
||
4201d6d6: addi sp,sp,-16
|
||
4201d6d8: sw s2,0(sp)
|
||
...
|
||
4201d6e4: lbu s1,0(s2)
|
||
4201d6e8: beqz s1,4201d7a8
|
||
4201d6ea: mv s0,a0
|
||
```
|
||
|
||
### E20: no-chain boundary probe
|
||
|
||
Added opt-in `?stateflush=open`, armed by the browser harness immediately before
|
||
opening `example.txt`. It reuses the existing `wasm_diff_needs_nochain_at(pc)`
|
||
app-flash window to force app TBs to return to C without running the expensive
|
||
compiled-vs-TCI differential shadow.
|
||
|
||
Result with `?nostoretci&stateflush=open`: original corruption still reproduced
|
||
(`Invalid read 0x3FCB80A6`, `MEPC 0x4201e210`). Returning to C between app TBs
|
||
is not enough; the wasm backend does not materialize `CPUArchState` merely by
|
||
leaving a compiled TB.
|
||
|
||
### E21/E22: blind helper spill attempts are invalid
|
||
|
||
Tested a direct backend fix candidate: spill TCG global-mem temps before C helper
|
||
calls so slow-path load/store helpers and exceptions see coherent architectural
|
||
state.
|
||
|
||
- spilling all global-mem temps before+after helpers caused an earlier invalid
|
||
read around `0x07585d82`;
|
||
- spilling before helpers only caused the same failure;
|
||
- narrowing to RISC-V integer GPR globals (`xN/...`) still caused immediate
|
||
invalid writes around `0x075f....`.
|
||
|
||
Conclusion: helper-boundary materialization is still the right suspect, but it
|
||
must respect TCG liveness/state. Blind backend spills at helper emission can
|
||
write stale globals back into `CPUArchState` and make corruption worse. The fix
|
||
belongs in the wasm backend's handling of TCG sync/liveness around calls (or in
|
||
targeted generated `sync_i32/i64` ops), not in an unconditional helper wrapper
|
||
spill.
|
||
|
||
### Revised next step
|
||
|
||
Trace the TCG liveness/sync ops for the divergent prologue TB and the later
|
||
slow-path helper TB: verify whether `la_global_sync()` marks `x2/sp` `TS_MEM`
|
||
across the relevant call/branch, and whether the wasm backend emits the
|
||
corresponding store. The observed bug is now specifically: **compiled globals and
|
||
`CPUArchState` are allowed to diverge, and the wasm backend exposes stale env to
|
||
helper/exception-visible code when store TBs stay compiled.**
|
||
|
||
|
||
## E23-E31: concrete TLB-array corruption
|
||
|
||
Added `FAULTTRACE`/`TLBTRACE` under `?ramrace=open` to inspect the exact failing
|
||
load path.
|
||
|
||
Key observations from crashing `?nostoretci&ramrace=open` runs:
|
||
|
||
1. The faulting guest address (`0x3fcb8004`/`0x3fcb80a6`) is valid ESP32-C3 DRAM.
|
||
2. TLB fill resolves the page correctly as RAM:
|
||
|
||
```text
|
||
FAULTTRACE tlb_fill vaddr=0x3fcb8000 paddr=0x3fcb8000 asidx=0 prot=0x3
|
||
mr=esp32c3.iram ram=1 romd=0 xlat=0x3c000
|
||
FAULTTRACE tlb_store index=56 iotlb=0x9c000 addr_page=0x3fcb8000
|
||
stored_xlat=0xffffffffc03e4000 phys=0x3fcb8000
|
||
FAULTTRACE mmu_lookup1 addr=0x3fcb8000 index=56 maybe_resized=1
|
||
tlb_addr=0x3fcb8000 flags=0x0 full_xlat=0xffffffffc03e4000 addend=0xc19d4000
|
||
```
|
||
|
||
3. Before the next byte load on the same page, the compact TLB and full TLB entry
|
||
are corrupted:
|
||
|
||
```text
|
||
FAULTTRACE mmu_lookup1 addr=0x3fcb8004 index=56 maybe_resized=0
|
||
tlb_addr=0x3fcb8200 flags=0x200 full_xlat=0x0 addend=0xc0348000
|
||
FAULTTRACE mmio_read addr=0x3fcb8004 xlat_section=0x0 mr=(null) ram=0
|
||
```
|
||
|
||
`0x200` is `TLB_MMIO`, so the vCPU incorrectly dispatches a DRAM read through
|
||
the MMIO/unassigned path. This is why the access is rejected even though the
|
||
guest address is in DRAM.
|
||
|
||
The corruption happens inside the active compiled TB before it returns, so the
|
||
post-TB scanner only sees the last good state. The faulting firmware sequence is
|
||
in `uzlib_uncompress`:
|
||
|
||
```text
|
||
4201e18c: lw a3,56(s0)
|
||
4201e18e: lw a5,24(s0)
|
||
4201e190: lw a4,52(s0)
|
||
4201e194: addi a2,a5,1
|
||
4201e198: sw a2,24(s0)
|
||
4201e19a: add a3,a3,a4
|
||
4201e19c: lbu a4,0(a3) # faults after TLB entry was corrupted
|
||
4201e1a0: sb a4,0(a5)
|
||
```
|
||
|
||
The compiled store at `0x4201e198` is now the prime suspect: it likely writes
|
||
through a bad host address into the TLB arrays, changing the same page's TLB
|
||
entry from RAM to MMIO and zeroing `fulltlb[index].xlat_section`. That matches
|
||
why blanket store-TCI works: the C/TCI store path does not corrupt the TLB arrays.
|
||
|
||
### Current root-cause statement
|
||
|
||
The bug is no longer just "nondeterministic race". It is specifically:
|
||
|
||
> compiled wasm store fast-path sometimes computes/uses a bad host address and
|
||
> writes into the vCPU softmmu TLB arrays (`CPUTLBEntry`/`CPUTLBEntryFull`),
|
||
> causing the immediately following load from valid DRAM to be treated as MMIO
|
||
> with `xlat_section=0`.
|
||
|
||
Still unknown: why the store fast path obtains the bad host address. Candidate
|
||
remaining causes: stale/corrupted TLB addend for the store's page, wasm32
|
||
address truncation/sign-extension around addends, or an in-TB ordering bug in the
|
||
compiled store TLB lookup. The next focused probe is to trace the store fast-path
|
||
host address for the `0x4201e198` TB before the `i64.store32` executes.
|
||
|
||
|
||
## E32-E34: fault-path TB isolation
|
||
|
||
Tried a narrow compiled-store guard (`?tlbguard=open`) that forces a store fast
|
||
path to the C helper if the computed host address lands inside the vCPU TLB
|
||
arrays. It did **not** suppress the crash, even though the guard range included
|
||
`fulltlb[56]`:
|
||
|
||
```text
|
||
TLBGUARD range lo=0x13e8000 hi=0x353c000
|
||
```
|
||
|
||
So the corrupting write is not caught as a normal compiled `qemu_st` fast-path
|
||
whose final host pointer falls directly inside that TLB-array range. The
|
||
corruption still appears between the TLB fill and the following byte load.
|
||
|
||
Added `?stpc=LO-HI`, an opt-in PC-window override that interprets only
|
||
store-containing TBs whose start PC lands in the window, while `?nostoretci`
|
||
keeps every other store TB compiled.
|
||
|
||
Results:
|
||
|
||
- plain `?nostoretci`: still crashes at the original signature.
|
||
- `?nostoretci&stpc=0x4201e180-0x4201e220`: **3/3 pass**
|
||
(`open example.txt` about 46-52s).
|
||
- wider/narrowed windows around `uzlib_uncompress` also pass as long as they
|
||
include the fault-path TB start.
|
||
|
||
This pins the corruption to compiled execution of the `uzlib_uncompress` TB that
|
||
contains the copy-loop sequence around:
|
||
|
||
```text
|
||
4201e18c: lw a3,56(s0)
|
||
4201e18e: lw a5,24(s0)
|
||
4201e190: lw a4,52(s0)
|
||
4201e194: addi a2,a5,1
|
||
4201e198: sw a2,24(s0)
|
||
4201e19a: add a3,a3,a4
|
||
4201e19c: lbu a4,0(a3)
|
||
4201e1a0: sb a4,0(a5)
|
||
```
|
||
|
||
The PC-window override is firmware-specific and **not** a production fix, but it
|
||
is now the smallest positive isolation: one compiled store-containing TB in the
|
||
inflate copy loop is sufficient to corrupt the vCPU TLB entry; interpreting that
|
||
TB avoids the crash while preserving compiled stores elsewhere.
|
||
|
||
Next useful step: dump or instrument the generated WASM for this exact TB and
|
||
inspect the `sw a2,24(s0)` / following `lbu` fast-path sequence. The current
|
||
node dump helper only captures early boot TBs, so late-reader TB dumping needs a
|
||
browser/open-window trigger or a PC-filtered `wasm_dump_module` path.
|