Files
nash/corpus/DIVERGENCES.md
T

58 lines
3.4 KiB
Markdown

# nash ↔ bash divergence log
Live record of behavioral divergences found by the replay harness
(`replay.py`, corpus in `corpus.jsonl`). Per docs/NASH.md §10.3/§11, each entry
stays open until fixed in the fork (or upstream) and re-verified by the corpus.
## Open
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale
- **Found**: narOS N0 rootfs validation (first real-world install run under the
divert): with `/bin/sh → nash`, `ca-certificates`' postinst
(`update-ca-certificates`, a `while read`-over-conf loop) mangled the UTF-8
filename `NetLock_Arany_=Class_Gold=_Főtanúsítvány.crt` and failed the
whole image build. Worked around structurally in `os/mkimage/build-rootfs.sh`
(the divert is now always the image's last configure step), but any
*runtime* `apt install` inside narOS still runs postinsts under nash, so this
blocks narOS/M3 forcing gates until fixed.
- **Repro** (nash `0.4.0` musl arm64, empty locale):
`echo 'Főtanúsítvány' | nash -c 'while read x; do echo "$x"; done'` —
bash emits the input bytes unchanged (`F \305\221 t a n \303\272 …`); nash
emits each byte Latin-1→UTF-8 double-encoded (`F \303\205 \302\221 …`).
- **Suspected root cause**: brush decodes input bytes to `String` with a lossy/
Latin-1 assumption on the `read` path (or at word-splitting) instead of
keeping raw bytes; on output the char sequence is re-encoded as UTF-8.
POSIX shells treat variable values as byte strings.
- **Severity**: silent data corruption (not a parse failure, so the §4.1
bash-fallback cannot catch it). Needs a corpus case (`utf8-bytes-passthru`)
and a byte-preservation sweep of read/expansion/heredoc paths.
## Closed
### D1 — IFS word-splitting applied to literal words (fixed in fork)
- **Found**: M0 corpus run (`ifs-split`), brush-shell-v0.4.0.
- **Repro**: `IFS=,; set -- a,b,c; echo $1` → bash `a b c`, nash `a`.
- **Root cause**: brush applied IFS splitting to *literal* words, not just
expansion results: under `IFS=,`, `echo a,b,c` printed `a b c` (bash:
`a,b,c`) and `set -- a,b,c` received 3 args (bash: 1). POSIX/bash split only
the results of parameter/command/arithmetic expansion. Mechanically,
`brush-core/src/expansion.rs` had a two-state `ExpansionPiece`
(`Splittable`/`Unsplittable`) conflating "field-splittable" with
"glob-active", and literal text was marked `Splittable`.
- **Fix**: added a third piece state `LiteralText` (never field-split, glob
chars active) and produce it for unquoted literal text, the no-expansion-chars
fast path, and retained-backslash escapes (`// nash:` markers in
`expansion.rs`). Corpus back to 100% (94/94); candidate for upstreaming.
- **Regression tests**: corpus `ifs-split` (+ `glob-all`, `glob-txt`,
`cond-pattern`, `case` guard the glob/pattern side); brush compat suite.
- **Fix fallout (caught by the brush compat suite, both fixed)**: (a) text
substituted by `${v:-word}`/`${v:+word}`/`${v:=word}` *is* an expansion result
and must split — restored via a boundary conversion in the
`ParameterExpansion` arm; (b) `compgen -W` splits its word-list string as
data — restored via `ExpanderOptions.field_split_literal_text`. Final suite:
1684 succeeded / 6 failed, failure set identical to the pristine-brush
baseline in this container (environmental only); 3 upstream known-fail IFS
tests now pass (markers flipped in `ifs.yaml`).