Merge nucleic/olive-ember-seal-q7vk into dev

This commit is contained in:
2026-07-18 15:33:08 -07:00
parent 5ff6fd8631
commit 490fd75319
2 changed files with 34 additions and 3 deletions
+31 -3
View File
@@ -6,7 +6,11 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus.
## Open
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale
(none)
## Closed (recent)
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale (fixed in fork)
- **Found**: narOS N0 rootfs validation (first real-world install run under the
divert): with `/bin/sh → nash`, `ca-certificates`' postinst
@@ -25,8 +29,32 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus.
keeping raw bytes; on output the char sequence is re-encoded as UTF-8.
POSIX shells treat variable values as byte strings.
- **Severity**: silent data corruption (not a parse failure, so the §4.1
bash-fallback cannot catch it). Needs a corpus case (`utf8-bytes-passthru`)
and a byte-preservation sweep of read/expansion/heredoc paths.
bash-fallback cannot catch it).
- **Root cause (confirmed)**: `brush-builtins/src/read.rs` `InputReader::read_event`
read one byte at a time and did `self.buffer[0] as char` — a Latin-1 decode —
before pushing into the result `String` (upstream had a `TODO(utf-8)` on the
field). Valid UTF-8 input was thus re-encoded byte-by-byte.
- **Fix**: incremental UTF-8 decoding in `read_event` (`// nash(D2)` markers):
lead byte classifies the sequence length (0xC2–0xF4), continuation bytes are
read (blocking, matching `read`'s wait-for-delimiter behavior; a
non-continuation byte is pushed back), and valid sequences round-trip
byte-identically. bash's `-n`-counts-bytes semantics are preserved for free
because the char-limit check measures `line.len()` (UTF-8 byte length).
Candidate for upstreaming (fixes upstream's own TODO).
- **Residual (accepted)**: *invalid* UTF-8 input bytes still degrade to the
per-byte Latin-1 mapping (a Rust `String` cannot hold raw invalid bytes;
byte-fidelity there would need `Vec<u8>` plumbing shell-wide). bash keeps raw
bytes. Not corpus-gated; revisit only if it bites in practice.
- **Verified**: repro now byte-identical to bash; split-across-writes
continuation reassembly correct; `read -n 3` byte counting identical to bash
(C locale); the original real-world failure (`update-ca-certificates` under
nash) exits 0 with all 141 certs processed; corpus 97/97 = 100% including
three new cases (`utf8-read-passthru`, `utf8-while-read-file`,
`utf8-cmdsub-roundtrip`); brush-builtins read tests 20/20; brush compat
suite failure set **identical to the pristine-HEAD baseline rebuilt in the
same container** (29 shared environmental failures, ±1 flaky SIGPIPE-timing
case that passes 10/10 standalone).
- **Regression tests**: the three `utf8-*` corpus cases above.
## Closed
+3
View File
@@ -92,3 +92,6 @@
{"id": "curl-sh-shape", "cmd": "printf 'echo piped-script\\n' | bash"}
{"id": "seq-paste", "cmd": "seq 1 4 | paste -sd+ | bc 2>/dev/null || seq 1 4 | paste -sd,"}
{"id": "nested-quotes-cmdsub", "cmd": "echo \"outer $(echo \"inner $(echo deepest)\")\""}
{"id": "utf8-read-passthru", "cmd": "printf 'F\\305\\221tan\\303\\272s\\303\\255tv\\303\\241ny\\n' | { read x; printf '%s\\n' \"$x\"; }"}
{"id": "utf8-while-read-file", "cmd": "printf 'caf\\303\\251.crt\\nna\\303\\257ve.pem\\n' > names.txt; while read n; do echo \"got:$n\"; done < names.txt"}
{"id": "utf8-cmdsub-roundtrip", "cmd": "x=$(printf '\\303\\251l\\303\\251gant'); echo \"$x\""}