Merge nucleic/olive-ember-seal-q7vk into dev
This commit is contained in:
+31
-3
@@ -6,7 +6,11 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus.
|
||||
|
||||
## Open
|
||||
|
||||
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale
|
||||
(none)
|
||||
|
||||
## Closed (recent)
|
||||
|
||||
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale (fixed in fork)
|
||||
|
||||
- **Found**: narOS N0 rootfs validation (first real-world install run under the
|
||||
divert): with `/bin/sh → nash`, `ca-certificates`' postinst
|
||||
@@ -25,8 +29,32 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus.
|
||||
keeping raw bytes; on output the char sequence is re-encoded as UTF-8.
|
||||
POSIX shells treat variable values as byte strings.
|
||||
- **Severity**: silent data corruption (not a parse failure, so the §4.1
|
||||
bash-fallback cannot catch it). Needs a corpus case (`utf8-bytes-passthru`)
|
||||
and a byte-preservation sweep of read/expansion/heredoc paths.
|
||||
bash-fallback cannot catch it).
|
||||
- **Root cause (confirmed)**: `brush-builtins/src/read.rs` `InputReader::read_event`
|
||||
read one byte at a time and did `self.buffer[0] as char` — a Latin-1 decode —
|
||||
before pushing into the result `String` (upstream had a `TODO(utf-8)` on the
|
||||
field). Valid UTF-8 input was thus re-encoded byte-by-byte.
|
||||
- **Fix**: incremental UTF-8 decoding in `read_event` (`// nash(D2)` markers):
|
||||
lead byte classifies the sequence length (0xC2–0xF4), continuation bytes are
|
||||
read (blocking, matching `read`'s wait-for-delimiter behavior; a
|
||||
non-continuation byte is pushed back), and valid sequences round-trip
|
||||
byte-identically. bash's `-n`-counts-bytes semantics are preserved for free
|
||||
because the char-limit check measures `line.len()` (UTF-8 byte length).
|
||||
Candidate for upstreaming (fixes upstream's own TODO).
|
||||
- **Residual (accepted)**: *invalid* UTF-8 input bytes still degrade to the
|
||||
per-byte Latin-1 mapping (a Rust `String` cannot hold raw invalid bytes;
|
||||
byte-fidelity there would need `Vec<u8>` plumbing shell-wide). bash keeps raw
|
||||
bytes. Not corpus-gated; revisit only if it bites in practice.
|
||||
- **Verified**: repro now byte-identical to bash; split-across-writes
|
||||
continuation reassembly correct; `read -n 3` byte counting identical to bash
|
||||
(C locale); the original real-world failure (`update-ca-certificates` under
|
||||
nash) exits 0 with all 141 certs processed; corpus 97/97 = 100% including
|
||||
three new cases (`utf8-read-passthru`, `utf8-while-read-file`,
|
||||
`utf8-cmdsub-roundtrip`); brush-builtins read tests 20/20; brush compat
|
||||
suite failure set **identical to the pristine-HEAD baseline rebuilt in the
|
||||
same container** (29 shared environmental failures, ±1 flaky SIGPIPE-timing
|
||||
case that passes 10/10 standalone).
|
||||
- **Regression tests**: the three `utf8-*` corpus cases above.
|
||||
|
||||
## Closed
|
||||
|
||||
|
||||
@@ -92,3 +92,6 @@
|
||||
{"id": "curl-sh-shape", "cmd": "printf 'echo piped-script\\n' | bash"}
|
||||
{"id": "seq-paste", "cmd": "seq 1 4 | paste -sd+ | bc 2>/dev/null || seq 1 4 | paste -sd,"}
|
||||
{"id": "nested-quotes-cmdsub", "cmd": "echo \"outer $(echo \"inner $(echo deepest)\")\""}
|
||||
{"id": "utf8-read-passthru", "cmd": "printf 'F\\305\\221tan\\303\\272s\\303\\255tv\\303\\241ny\\n' | { read x; printf '%s\\n' \"$x\"; }"}
|
||||
{"id": "utf8-while-read-file", "cmd": "printf 'caf\\303\\251.crt\\nna\\303\\257ve.pem\\n' > names.txt; while read n; do echo \"got:$n\"; done < names.txt"}
|
||||
{"id": "utf8-cmdsub-roundtrip", "cmd": "x=$(printf '\\303\\251l\\303\\251gant'); echo \"$x\""}
|
||||
|
||||
Reference in New Issue
Block a user