diff --git a/corpus/DIVERGENCES.md b/corpus/DIVERGENCES.md index 1ffe9e6..615aefc 100644 --- a/corpus/DIVERGENCES.md +++ b/corpus/DIVERGENCES.md @@ -6,7 +6,11 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus. ## Open -### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale +(none) + +## Closed (recent) + +### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale (fixed in fork) - **Found**: narOS N0 rootfs validation (first real-world install run under the divert): with `/bin/sh → nash`, `ca-certificates`' postinst @@ -25,8 +29,32 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus. keeping raw bytes; on output the char sequence is re-encoded as UTF-8. POSIX shells treat variable values as byte strings. - **Severity**: silent data corruption (not a parse failure, so the §4.1 - bash-fallback cannot catch it). Needs a corpus case (`utf8-bytes-passthru`) - and a byte-preservation sweep of read/expansion/heredoc paths. + bash-fallback cannot catch it). +- **Root cause (confirmed)**: `brush-builtins/src/read.rs` `InputReader::read_event` + read one byte at a time and did `self.buffer[0] as char` — a Latin-1 decode — + before pushing into the result `String` (upstream had a `TODO(utf-8)` on the + field). Valid UTF-8 input was thus re-encoded byte-by-byte. +- **Fix**: incremental UTF-8 decoding in `read_event` (`// nash(D2)` markers): + lead byte classifies the sequence length (0xC2–0xF4), continuation bytes are + read (blocking, matching `read`'s wait-for-delimiter behavior; a + non-continuation byte is pushed back), and valid sequences round-trip + byte-identically. bash's `-n`-counts-bytes semantics are preserved for free + because the char-limit check measures `line.len()` (UTF-8 byte length). + Candidate for upstreaming (fixes upstream's own TODO). +- **Residual (accepted)**: *invalid* UTF-8 input bytes still degrade to the + per-byte Latin-1 mapping (a Rust `String` cannot hold raw invalid bytes; + byte-fidelity there would need `Vec` plumbing shell-wide). bash keeps raw + bytes. Not corpus-gated; revisit only if it bites in practice. +- **Verified**: repro now byte-identical to bash; split-across-writes + continuation reassembly correct; `read -n 3` byte counting identical to bash + (C locale); the original real-world failure (`update-ca-certificates` under + nash) exits 0 with all 141 certs processed; corpus 97/97 = 100% including + three new cases (`utf8-read-passthru`, `utf8-while-read-file`, + `utf8-cmdsub-roundtrip`); brush-builtins read tests 20/20; brush compat + suite failure set **identical to the pristine-HEAD baseline rebuilt in the + same container** (29 shared environmental failures, ±1 flaky SIGPIPE-timing + case that passes 10/10 standalone). +- **Regression tests**: the three `utf8-*` corpus cases above. ## Closed diff --git a/corpus/corpus.jsonl b/corpus/corpus.jsonl index 0bcdeee..52d5f3c 100644 --- a/corpus/corpus.jsonl +++ b/corpus/corpus.jsonl @@ -92,3 +92,6 @@ {"id": "curl-sh-shape", "cmd": "printf 'echo piped-script\\n' | bash"} {"id": "seq-paste", "cmd": "seq 1 4 | paste -sd+ | bc 2>/dev/null || seq 1 4 | paste -sd,"} {"id": "nested-quotes-cmdsub", "cmd": "echo \"outer $(echo \"inner $(echo deepest)\")\""} +{"id": "utf8-read-passthru", "cmd": "printf 'F\\305\\221tan\\303\\272s\\303\\255tv\\303\\241ny\\n' | { read x; printf '%s\\n' \"$x\"; }"} +{"id": "utf8-while-read-file", "cmd": "printf 'caf\\303\\251.crt\\nna\\303\\257ve.pem\\n' > names.txt; while read n; do echo \"got:$n\"; done < names.txt"} +{"id": "utf8-cmdsub-roundtrip", "cmd": "x=$(printf '\\303\\251l\\303\\251gant'); echo \"$x\""}