Published July 2, 2026 | Version 1.4
Description
We study whether the spectral sensitivity of the per-token empirical Fisher Information Matrix (FIM), ωmax(Ft), can serve as a lightweight, model-agnostic runtime signal for anticipating hallucination during autoregressive decoding on consumer edge hardware. Across four pre-registered experiments on twelve open-weight models run locally via Ollama and MLX on an Apple M5 (32GB unified memory), we establish, in turn: cross-run stability of the signal, its predictive correlation with hallucination onset, its early-warning lead time, and the e!ectiveness of an intervention triggered by the resulting alarm. Concretely, ten of twelve models satisfy a stability gate (coe”cient of variation < 0.20); the signal correlates strongly with an entropy-based hallucination-onset label for non-thinking-mode decoding, most strongly for qwen3:32b (r = 0.962, p = 1.4 → 10→28, median r = 0.479 across nine non-thinking models) – a correlation we report as an upper bound given a noted oracle-circularity caveat between the label and the signal itself; a first-passage-time alarm flags onset at a median lead time of 27 tokens (IQR 12–39) with an 85.6% detection rate (t = 33.2, p = 2.2 → 10→114); and truncating generation once the alarm fires reduces post-onset token count by 66% (Cohen’s d = 1.95) without measurable quality degradation. Separately, we report and analyze a 244→ gap between the analytically derived and the empirically calibrated KL-divergence alarm threshold under 4-bit (Q4 K M) quantization on this hardware, corroborated by a multi-model robustness check (range 64–287→ across eight models) – a finding we frame not as a universal constant but as evidence for a general Calibration Necessity Protocol : safety-relevant inference-time thresholds must be empirically recalibrated per hardware/quantization configuration rather than transferred analytically. We report this drift, a full experimental and hardware configuration, and a set of explicit limitations – including a top-K = 20 logprob constraint, single-hardware-platform scope, and an open correction-mechanism gap – to support independent replication and extension.
Files
KA_FisherEarlyWarning_arxiv_preview.pdf
Files (538.8 kB)
Additional details
- J. Pennington and P. Worah, "The spectrum of the Fisher information matrix of a single- hidden-layer neural network," in Advances in Neural Information Processing Systems (NeurIPS), 2018. https://proceedings.neurips.cc/paper_files/paper/2018/file/18bb68e2b38e4a8ce7cf4f6b2625768c-Paper.pdf
- S. Kadavath et al., "Language models (mostly) know what they know," arXiv preprint arXiv:2207.05221, 2022. https://doi.org/10.48550/arXiv.2207.05221
- S. Farquhar et al., "Detecting hallucinations in large language models using semantic entropy," Nature, vol. 630, 2024. https://doi.org/10.1038/s41586-024-07421-0
- O. Shorinwa et al., "A survey on uncertainty quantification of large language models: taxonomy, open research challenges, and future directions," arXiv preprint, 2025. https://doi.org/10.48550/arXiv.2412.05563