Merge nucleic/olive-ember-seal-q7vk into dev

This commit is contained in:
2026-07-18 15:33:08 -07:00
parent 5ff6fd8631
commit 490fd75319
2 changed files with 34 additions and 3 deletions
+31 -3
View File
@@ -6,7 +6,11 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus.
## Open
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale
(none)
## Closed (recent)
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale (fixed in fork)
- **Found**: narOS N0 rootfs validation (first real-world install run under the
divert): with `/bin/sh → nash`, `ca-certificates`' postinst
@@ -25,8 +29,32 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus.
keeping raw bytes; on output the char sequence is re-encoded as UTF-8.
POSIX shells treat variable values as byte strings.
- **Severity**: silent data corruption (not a parse failure, so the §4.1
bash-fallback cannot catch it). Needs a corpus case (`utf8-bytes-passthru`)
and a byte-preservation sweep of read/expansion/heredoc paths.
bash-fallback cannot catch it).
- **Root cause (confirmed)**: `brush-builtins/src/read.rs` `InputReader::read_event`
read one byte at a time and did `self.buffer[0] as char` — a Latin-1 decode —
before pushing into the result `String` (upstream had a `TODO(utf-8)` on the
field). Valid UTF-8 input was thus re-encoded byte-by-byte.
- **Fix**: incremental UTF-8 decoding in `read_event` (`// nash(D2)` markers):
lead byte classifies the sequence length (0xC2–0xF4), continuation bytes are
read (blocking, matching `read`'s wait-for-delimiter behavior; a
non-continuation byte is pushed back), and valid sequences round-trip
byte-identically. bash's `-n`-counts-bytes semantics are preserved for free
because the char-limit check measures `line.len()` (UTF-8 byte length).
Candidate for upstreaming (fixes upstream's own TODO).
- **Residual (accepted)**: *invalid* UTF-8 input bytes still degrade to the
per-byte Latin-1 mapping (a Rust `String` cannot hold raw invalid bytes;
byte-fidelity there would need `Vec<u8>` plumbing shell-wide). bash keeps raw
bytes. Not corpus-gated; revisit only if it bites in practice.
- **Verified**: repro now byte-identical to bash; split-across-writes
continuation reassembly correct; `read -n 3` byte counting identical to bash
(C locale); the original real-world failure (`update-ca-certificates` under
nash) exits 0 with all 141 certs processed; corpus 97/97 = 100% including
three new cases (`utf8-read-passthru`, `utf8-while-read-file`,
`utf8-cmdsub-roundtrip`); brush-builtins read tests 20/20; brush compat
suite failure set **identical to the pristine-HEAD baseline rebuilt in the
same container** (29 shared environmental failures, ±1 flaky SIGPIPE-timing
case that passes 10/10 standalone).
- **Regression tests**: the three `utf8-*` corpus cases above.
## Closed