Merge nucleic/sleek-ember-seal-uady into dev
This commit is contained in:
@@ -0,0 +1,46 @@
|
||||
# Purpose-classifier energy and accelerator-residency check
|
||||
|
||||
This is the required per-model-version hardware spot check from
|
||||
`docs/PURPOSE_CLASSIFIER.md` §8.5. It is deliberately not CI: the numbers are meaningful
|
||||
only on named physical hardware with a stable power source, thermal state, and runtime
|
||||
placement.
|
||||
|
||||
## Common protocol
|
||||
|
||||
1. Record the artifact SHA-256, app/build commit, machine model, OS build, battery/AC
|
||||
state, ambient power mode, and runtime compute policy.
|
||||
2. Warm the model with 100 batch-one classifications, then classify the same fixed
|
||||
1,000-prompt sequence. Tokenization is included. Disable network activity and other
|
||||
foreground workloads.
|
||||
3. Run at least five trials per compute policy, alternating policy order. Report median
|
||||
wall time, p95 inference latency, average package power, and joules per inference.
|
||||
4. Capture the runtime placement evidence alongside the power trace. A fast CPU fallback
|
||||
is still a residency failure for `purpose-deep`.
|
||||
|
||||
## macOS
|
||||
|
||||
- Run the release Core ML artifact once with `.cpuAndNeuralEngine`, recording the
|
||||
`MLComputePlan` ANE operation share and the artifact's required floor.
|
||||
- Repeat with a diagnostic CPU-only configuration on the same Mac and power source.
|
||||
- Sample the 1,000-inference region with `sudo powermetrics --samplers cpu_power,gpu_power`
|
||||
at one-second intervals. Store the raw trace outside git and put its summarized table in
|
||||
the model metrics report.
|
||||
- The default ANE policy is accepted only when it uses less average power and fewer joules
|
||||
per inference than CPU-only. On AC, a `.all` GPU retry must separately meet its declared
|
||||
GPU operation-share floor before it can activate a deep model.
|
||||
|
||||
## Windows
|
||||
|
||||
- Record the Windows ML execution provider and assigned device after AOT compilation.
|
||||
Deep requires QNN NPU, OpenVINO NPU/GPU, NvTensorRT-RTX, or MIGraphX; bare CPU is only
|
||||
valid for lite.
|
||||
- On battery compare `MAX_EFFICIENCY` with an explicit CPU session. On AC compare the
|
||||
selected NPU, otherwise the validated vendor GPU EP, with CPU.
|
||||
- Capture the 1,000-inference window with HWiNFO sensor logging or the vendor's documented
|
||||
NPU/GPU telemetry. Include package/device average power and energy per inference in the
|
||||
model metrics report.
|
||||
- Verify that unplugging skips the discrete-GPU rung and that an idle session unloads
|
||||
without keeping the dGPU awake.
|
||||
|
||||
The report is incomplete if it gives latency without placement evidence and energy, or if
|
||||
it compares different prompt sequences between policies.
|
||||
Reference in New Issue
Block a user