Forward
2026-08-15 Gavin Crooks
This paper is an exploration of AI driven research. Everything below, from the abstract on down, was derived and written by Claude Fable 5. A research idea that has long been on my list of open problems in stochastic thermodynamics was this: Given a detailed fluctuation theorem of the form p(σ)/p(−σ) = e^σ what are the constraints on the statistics of the entropy production? Many partial results have been published: "exchange thermodynamic uncertainty relation" of Timpanaro, Guarnieri, Goold and Landi (2019), "thermodynamic uncertainty theorem" of Ray, Boyd, Guarnieri and Crutchfield (2023), a whole series of papers by Salazar that I have enjoyed reading, among others. But what was missing was a unifying framework. So I asked Claude to take a look. This is the kind of problem what I might have suggested to a mathematically inclined grad student. But over a few days of back and forth, Claude did months of work and seems to have closed the whole problem class.
The academic sciences are in for a major shake up. Academia has dual roles: advancing knowledge, and teaching/training/mentoring the next generation. Science is about to advance faster perhaps than we can keep up. How then do we teach and train students, when any question I might have could be answered faster by Claude?
And what of scientific publishing? It seems absurd to submit this article for publication, or even post on arXiv, when it represents only a couple of days work, and could be reproduced by any practitioner in the field with a Claude subscription. Thus science by twitter article!
Abstract
The detailed fluctuation theorem (DFT), p(σ)/p(−σ) = e^σ, constrains the statistics of entropy production in time-symmetrically driven and exchange processes, and a growing family of inequalities — thermodynamic uncertainty relations and bounds on second-law violations, skewness, information, and the moment generating function — has been derived from it, one at a time. Here we solve the underlying problem completely. Every distribution obeying the DFT is a unique mixture of elementary two-outcome distributions, so the joint range of any set of its statistics forms a single convex region, which we characterize exactly at every order: given the first n−1 moments, the n-th moment can take any value above a sharp floor — attained by a unique distribution with finitely many symmetric outcome levels — and is never bounded from above. The mechanism is a structural fact that appears to have gone unnoticed: the DFT mean function a·tanh(a/2) is a positive sum of simple relaxation terms with poles at the odd squares — equivalently, the response function of a uniform vibrating string — and this places the entire hierarchy under the classical theory of moment problems, via an exact identity we prove between the hierarchy's Wronskians and Hankel determinants. As corollaries we obtain the complete achievable regions for mean and variance, skewness at fixed variance, tail probabilities, entropy, and the moment generating function; we recover every published DFT bound as a low-dimensional shadow of the one region and explain why a single distribution saturates nearly all of them. By contrast, the integral fluctuation theorem alone constrains none of these tradeoffs. The theory extends to general two-distribution (Crooks) fluctuation theorems through a symmetrization argument: sums of forward and reverse moments obey exactly the same hierarchy, saturated by time-symmetric processes, while forward and reverse dissipations are individually unconstrained — every constraint of the two-distribution theorem lives in its symmetric channel. The saturating distributions are minimal multi-level exchange engines, whose full moment sequences obey exact recursions — for any three-outcome engine, a geometric progression.
1. Introduction
When a mesoscopic system is driven by a time-symmetric protocol, or exchanges heat and particles between reservoirs, the entropy production σ accumulated over a run is a random variable, and its distribution obeys the detailed fluctuation theorem (Evans–Searles; Jarzynski–Wójcik):
p(σ)/p(−σ) = e^σ. (1)
Equation (1) is a single functional constraint, but its consequences have so far been mined one inequality at a time. The integral fluctuation theorem ⟨e^{−σ}⟩ = 1 and the second law ⟨σ⟩ ≥ 0 follow immediately. Merhav and Kafri, in a prescient information-theoretic study [MK10], derived from (1) a tight lower bound on the variance at fixed mean and tight bounds on the probability of observing negative entropy production. The variance bound resurfaced, apparently independently, as the exchange thermodynamic uncertainty relation of Timpanaro, Guarnieri, Goold and Landi [TGGL19], who proved it the tightest of its kind; a weaker exponential form was found earlier by Proesmans and Van den Broeck [PVdB17] and again by Hasegawa and Van Vu [HV19]. Salazar derived from (1) a tight upper bound on the information content of σ [S21a], a tight lower bound on apparent second-law violations [S21b], a tight lower bound on skewness [S22], and a family of bounds on the moment generating function [S23]; Nishiyama and Hasegawa obtained a tail bound [NH23]; and the "thermodynamic uncertainty theorem" of Ray, Boyd, Guarnieri and Crutchfield [TUT] identified the exact minimum-variance current allowed by (1).
Each of these results answers part of one question: what is the full set of probability distributions compatible with (1), and what, jointly, can their statistics be? This paper answers that question completely, and in doing so places every bound above inside a single geometric object.
The answer has three layers. Structure. The DFT couples each outcome σ = +a to its mirror −a with a fixed weight ratio e^a, and nothing else: once you know the distribution of the gap |σ|, the signs are determined. Every DFT distribution is therefore a unique blend of elementary two-outcome distributions P_a (outcomes ±a with the DFT-mandated weights), and every average of interest becomes a weighted average over the blend. The set of achievable statistics is then automatically a convex region — mixing distributions mixes their statistics — and we call it the moment body: one shape whose boundary encodes every constraint (1) places on the statistics simultaneously. Solution. We characterize this body exactly at every order (Fig. F4). Given the first n−1 moments, the n-th ranges over a half-line [T_n, ∞): there is a sharp floor, attained by a unique distribution supported on ⌈n/2⌉ symmetric outcome levels (plus a level at zero when n is odd), and there is never a ceiling. The DFT is a machine that produces lower bounds only. Mechanism. Why is the problem exactly solvable? Because of one identity. Writing the mean of the elementary distribution P_a as m(a) = a·tanh(a/2), the partial-fraction expansion of tanh gives
m(a) = Σ_{k≥1} 4a²/(a² + π²(2k−1)²),
a positive sum of simple "relaxation-type" terms with poles at the odd squares. Sums of this form (Stieltjes transforms, in the mathematics literature) inherit a powerful rigidity from the elementary kernel 1/(t+b) — a rigidity known as total positivity — and it is exactly what makes every optimization over the body land on a few-point distribution with nothing left over. Readers who know spectral theory may enjoy the interpretation: this m is the response function of a homogeneous vibrating string measured at its endpoint, and the odd squares are the string's resonance frequencies. The physics enters the mathematics at this single point — the e^σ asymmetry produces tanh, and tanh's poles hand the problem to classical moment theory.
Two framing results sharpen the picture. The lower-bounds-only principle, already stated; and a null result for the integral theorem (Theorem 2): the class defined by ⟨e^{−σ}⟩ = 1 alone has no constraint linking mean and variance — its variance floor is zero at every mean. The entire content of the uncertainty-relation family is therefore detailed, not integral, information. A single elementary mechanism — a rare event of vanishing probability pushed to an extreme value, tuned so that one chosen average stays finite while all others stop noticing it — underlies this and every other absence of a bound in the theory. We state it once as Lemma 1 and use it throughout.
With the body in hand, the published bounds fall into place (Section 8, Table 1). The exchange TUR is not merely a tight bound: it is the entire lower edge of the (mean, variance) shadow of the body. The skewness relation is the mean-axis shadow of the exact (mean, variance, skewness) boundary. The violation bound is the complete (mean, sign-statistic) region; the information bound is the attained upper edge in (mean, entropy), with nothing below it; the MGF bounds are exact envelopes for every tilt (Theorem 9); the tail bound of [NH23] is one branch of the exact tail envelope (Theorem 7), which also strictly sharpens Jarzynski's classic bound on negative excursions. Nearly all of them are saturated by the same distribution — the two-outcome law sitting at the corner of the body at the given mean — and we explain why this must be so. Higher moments carry strictly new information: we prove that even the entire MGF family cannot imply the third-moment floor. Beyond the audit, the body yields new sharp bounds of direct use — most notably exact tail envelopes that strictly tighten Jarzynski's estimate of negative-entropy events at every threshold, a two-parameter skewness floor, variance-resolved violation bounds, and a testable geometric law for the moments of three-outcome engines (Table 2). The theory also settles the two-distribution case: for a general Crooks pair, the sums of forward and reverse moments obey exactly the hierarchy of this paper, while nothing constrains the difference channel — making exact the folklore that one-sided free-energy estimators can fail arbitrarily badly, and why bidirectional estimation is qualitatively necessary (Section 9).
The saturating distributions have direct physical meaning. At every level of the hierarchy they are minimal multi-level exchange engines — at third order, precisely the three-outcome distributions realized by qubit swap engines — and on each boundary piece of the body, all higher moments are frozen into exact linear recursions. For any three-outcome exchange process the law is geometric: every swap engine's full moment sequence is predicted by its first three moments.
Section 2 sets up the structure; Sections 3–4 solve the warm-up and the integral-only contrast; Section 5 presents the mechanism; Section 6 the hierarchy; Section 7 tails, entropy, and the MGF; Section 8 the audit and the new bounds; Section 9 extends the theory to two-distribution (Crooks) pairs; Section 10 discusses inference, the string picture, and extensions. Proofs are collected in Appendices A–D, the numerical protocol in Appendix E, closed forms in Appendix F.
2. One free ingredient: the distribution of the gap
Let 𝒫 be the set of probability distributions on ℝ satisfying (1). For a ≥ 0, define the elementary two-outcome distribution
P_a: outcomes ±a with weights w±(a) = 1/(1 + e^{∓a}) (P₀ is the point mass at 0).
Proposition 1. Every p ∈ 𝒫 arises from exactly one probability distribution ν of the gap a = |σ|, via p = ∫ P_a dν(a); conversely every ν produces a member of 𝒫.
The proof (Appendix A) formalizes a transform introduced by Merhav and Kafri [MK10]: conditioned on |σ| = a, the DFT fixes the sign probabilities completely, so the gap distribution ν is the only freedom. In convex-geometry language 𝒫 is a simplex — like a triangle or tetrahedron, a set in which every point has a unique decomposition into corners — and the corners are the P_a. Every average then translates: for any observable f,
⟨f(σ)⟩ = E_ν[F(a)], F(a) = w₊(a)f(a) + w₋(a)f(−a).
In particular, with m(a) := a·tanh(a/2):
odd moments ⟨σ^{2j+1}⟩ = E_ν[a^{2j+1}tanh(a/2)]; even moments ⟨σ^{2j}⟩ = E_ν[a^{2j}]; ⟨tanh(σ/2)⟩ = E_ν[tanh²(a/2)]; P(σ<0) = E_ν[w₋(a)]; ⟨e^{ασ}⟩ = E_ν[F_α(a)],
with F_α(a) = cosh((α+½)a)/cosh(a/2). The joint range of any statistics is thus the set of weighted averages of a fixed curve of "payload" vectors — computable by linear programming, and, as we show, solvable in closed form.
One change of variables organizes everything. Set (h, r) = (m(a), a²): the odd-moment payloads become r^j·h and the even ones r^j, so the whole cumulant program is a moment problem along the single curve r = φ(h), where φ is convex (Appendix B.1). The locations and weights of a few atoms on this curve are the entire content of every sharp bound below (Fig. F4a).
3. Warm-up: the exact (mean, variance) region
Theorem 1. The achievable set of (μ, V) = (⟨σ⟩, Var σ) is
{(0, 0)} ∪ { (μ, V) : μ > 0, V ≥ V_min(μ) }, V_min(μ) = a_min(μ)² − μ²,
where a_min(μ) is the unique gap with m(a) = μ. The floor is attained, at each μ, only by the two-outcome distribution P_{a_min(μ)}; every value above it is attainable; there is no upper limit.
The floor and its two-outcome saturator are due to Merhav and Kafri [MK10] (a Jensen argument against the convexity of the curve; the required inequality boils down to sinh 2x ≥ 2x). Our additions are the completeness of the statement — uniqueness, attainability of everything above, the isolated equilibrium point — and three consequences. First, V_min coincides identically with the exchange TUR of [TGGL19]: that bound is the entire lower edge of the region, what a mathematician would call a facet — not merely a valid inequality that happens to be reachable somewhere, but the exact boundary everywhere. Second, the exponential bound of [PVdB17, HV19] lies strictly inside the region at every μ > 0, and the Markov-steady-state bound of Barato and Seifert is violated by the elementary distributions themselves — it is not a consequence of (1) at all (Fig. F1). Third, the floor has clean structure: V_min = 2μ − (2/3)μ² near equilibrium, V_min ≈ 4μ²e^{−μ} far from it, and its maximum sits at exactly ⟨σ⟩ = 2 (in units of k_B), where V_min = 4(g₀² − 1) ≈ 1.757 with g₀tanh g₀ = 1. The mean dissipation most permissive of fluctuations is two.
4. The integral theorem alone implies none of this
Lemma 1 (rare-event escape). Attach a small probability ε to an extreme outcome L, tuning ε × (the contribution of L to one chosen average) to stay finite as L runs away and ε → 0. Every other average, weighting the event more lightly, converges to its value without the event. (Appendix A.2 gives the precise statement; the lemma is invoked in Theorems 2, 4(c), 8(ii), and 7.)
Theorem 2. For the class defined only by ⟨e^{−σ}⟩ = 1, the achievable (mean, variance) set is {(0,0)} ∪ {μ > 0, V > 0}: at every positive mean, the variance can be made as small as desired (never exactly zero), and there is no other constraint.
The construction is Lemma 1 in its original habitat: σ = μ almost surely, plus a rare excursion at s → −∞ with weight (1−e^{−μ})e^s, which balances ⟨e^{−σ}⟩ = 1 exactly while contributing s²e^s → 0 to the variance. This is the familiar statement that Jarzynski-type averages are dominated by rare events — here made exact and two-sided. The detailed structure of (1) forbids precisely this burial: the pairing p(−a) = e^{−a}p(a) forces the negative branch to sit opposite the positive one with a prescribed, non-negligible weight. In this precise sense the entire uncertainty-relation family is detailed-FT information; no bound in Table 1 has an integral-only analogue in these coordinates. (What the integral theorem does constrain — a family of change-of-measure inequalities — was identified in [MK10, §4].)
5. The mechanism: one identity, total positivity
Why is every optimization over the body exactly solvable, with unique few-point winners? The answer is the partial-fraction identity of the introduction. In the variable t = a², the mean function is H(t) = √t·tanh(√t/2), and
H(t) = Σ_{k≥1} 4t/(t + b_k), b_k = π²(2k−1)² (2)
— a positive superposition of the simplest possible increasing saturating functions, t/(t + b). Functions of this form carry a strong hidden order. The intuition: each kernel t/(t+b) is monotone and concave in a coordinated way across all b simultaneously, and no cancellation is possible because all the weights are positive; sums of them can therefore wiggle only a strictly limited number of times against any polynomial. The technical name for this rigidity is total positivity of the kernel 1/(t+b), and its working consequence for us is a Chebyshev property: the function families {1, m, a², a³tanh(a/2), …, up to order n} behave like polynomials of degree n for the purpose of counting zeros — no combination of them can cross zero more than n times. That zero-counting is what forces every constrained optimum onto few atoms and makes it unique.
Proving the Chebyshev property at every order reduces, by a classical criterion, to showing that a certain sequence of determinants — Wronskians, determinants built from the functions and their successive derivatives — are all positive. Our central technical result computes them in closed form:
Theorem 3. With T_m(t) = Σ_k b_k/(t + b_k)^m, the n-th Wronskian of the hierarchy equals
W_n = 4^{⌈n/2⌉}·(2!·3!⋯n!) × det[T_{i+j+2}] (n odd) or det[T_{i+j+3}] (n even),
where the determinant on the right is a Hankel determinant — the determinant of a table of moments T_m, which a classical result (Heine) guarantees is positive whenever the T_m really are the moments of a positive weight, as they are here. Hence all Wronskians are positive and the Chebyshev property holds at every order.
The identity is proved in Appendix B by reducing to finitely many poles, where partial fractions turn the problem into classical determinant evaluations; it is verified symbolically and in exact arithmetic through n = 7 and to fifty digits against (2) through n = 9. The low orders are worth seeing, because they reveal that the scattered convexity lemmas of the prior literature were one mechanism all along: positivity of the first Wronskian is "H increasing"; of the second, "H concave" — the Merhav–Kafri/TGGL convexity; and of the third, literally the Cauchy–Schwarz inequality for the weight in (2).
6. The complete moment hierarchy
Theorem 4. Let M_n be the set of achievable moment vectors (⟨σ⟩, …, ⟨σⁿ⟩). Then: (a) M_n is the convex hull of the payload curve, extended to infinity in exactly one direction — straight up in the top moment. (b) Given any interior choice of the first n−1 moments, the n-th moment ranges over a half-line [T_n, ∞), whose floor is attained by a unique distribution with ⌈n/2⌉ outcome levels, including the level 0 exactly when n is odd. (c) No moment is ever bounded above, and all odd moments are nonnegative. (d) On any boundary piece where the minimizing distribution has squared levels {r₁, …, r_l}, all higher moments obey the linear recursion whose characteristic polynomial is ∏(r − rᵢ): everything above the boundary's order is a closed form in what defines it.
Instances (Fig. F4b). n = 2 is Theorem 1. n = 3 (Theorem 5): ⟨σ³⟩ ≥ μx², where x solves tanh(x/2)/x = μ/⟨σ²⟩, attained only by the three-outcome distribution on {−x, 0, x}. As a skewness statement this is two-parameter and strictly tighter than the mean-only relation of [S22] — by a factor of 38 already at ⟨σ²⟩ = 4× its floor — and reduces to it exactly on the variance floor. Its content is directional: at the variance floor the DFT forces left skew (κ₃ = −2μV there), while at large variance it forces right skew, γ₁ ≳ √V/μ — spread at fixed dissipation must live in a heavy positive tail (Fig. F2). n = 4 (Theorem 6): the floor is attained by a unique two-level distribution, and on the third-moment boundary the fourth moment is pinned: M₄ = M₂M₃/μ. n = 5, 6 follow the stated pattern, verified to 10⁻⁸.
The recursion law (d) has a direct physical reading. Any DFT distribution on three outcomes is forced into the form p ∝ e^{σ/2} — the distribution of a qubit swap engine — and lives on the {0, x} boundary piece, where (d) becomes the geometric law
M_{k+2} = (M₃/μ)·M_k :
all moments of any three-outcome exchange engine form a geometric progression, with ratio equal to the squared level gap. More generally, the floors of the hierarchy are realized by minimal multi-level engines, and the boundary of the moment body is a stratified family of exactly solvable machines. Merhav–Kafri's mean-only bounds ⟨|σ|^k⟩ ≥ h(μ)^k [MK10] are recovered as the deepest shadows of (b).
7. Tails, entropy, and the moment generating function
Tails (Theorem 7). For a threshold t > 0 at fixed mean μ, with m(t) = t·tanh(t/2):
max P(σ > t) = μ/(t(1 − e^{−t})) for μ ≤ m(t) [attained by the three-outcome distribution on {−t, 0, t}]; = 1 − L(μ) for μ ≥ m(t) [attained by P_{a_min}], where L(y) = 1/(1+e^{a_min(y)}); max P(σ < −t) = μ/(t(e^t − 1)) for μ ≤ m(t) [attained]; ceiling 1/(1 + e^t) for μ > m(t) (approached, never reached); max P(|σ| > t) = min(μ/m(t), 1); and every tail probability can be made as small as desired at every (μ, t) — no floors, by Lemma 1.
Three remarks. The familiar Markov inequality P(σ>t) ≤ μ/t is false here — σ takes negative values — and the correct sharp constant is larger by (1 − e^{−t})⁻¹; the first branch of the upper tail appears, without its sharpness, in [NH23]. The cap 1/(1 + e^t) strictly sharpens Jarzynski's e^{−t} bound on negative excursions at every t, and the violation bound of [S21b] is the t = 0 member of the family. Being exact, these envelopes beat every Chernoff bound the MGF family below could produce; collectively they pin down which cumulative distribution functions are achievable, point by point. (Two complementary lines of work control different objects: martingale bounds on the running minimum of σ along a trajectory [NRJ17], and concentration inequalities built from the dynamical activity of Markov processes [HN24], which use dynamical information the bare DFT does not supply.)
Entropy (Theorem 8). For distributions with a density, the entropy splits cleanly along the mixture: h[p] = h[ν] + E_ν[φ(a)], where φ(a) is the entropy of the two-way sign choice at gap a, and w₊e^{φ} = e^{(a−m(a))/2} exactly. Standard maximum-entropy calculus then delivers, in one line, the complete (μ, h) region: h ≤ M(μ), with the upper edge attained by a unique distribution — the maximal distribution of [S21a], whose bound is thereby the entire boundary — and no lower edge at all (h → −∞; for discrete distributions, mean and Shannon entropy are jointly unconstrained, by Lemma 1 once more). A recent relation of Hasegawa and Nishiyama [HN25] — entropy production plus the "sign entropy" Λ = H[P(σ)] − H[P(|σ|)] is at least ln 2 — sits inside this geometry: on our class Λ = E_ν[φ], their bound is the mixture-average of the pointwise inequality m + φ ≥ ln 2, and the exact floor for distributions without an atom at zero is Λ ≥ φ(a_min(Σ)) — saturated at the usual corner and strictly tighter, remaining informative at every mean. An atom exactly at σ = 0 (a fixed point of the sign pairing) evades any such constant, again by Lemma 1.
The moment generating function (Theorem 9). For every real α, F_α(a) = cosh((α+½)a)/cosh(a/2) is a convex function of m(a). Hence, for every α, Salazar's bound ⟨e^{ασ}⟩ ≥ F_α(a_min(μ)) [S23] is the exact lower envelope of the MGF at fixed mean, saturated only by P_{a_min}. The proof (Appendix D) is half a page of hyperbolic algebra: the convexity condition collapses, via the identities tanh²+sech² = 1 and a reappearance of sinh 2x ≥ 2x, to a single inequality whose minimum over α sits exactly at the trivial members α ∈ {0, −1} (where F ≡ 1 and equality is forced). Notably, the MGF family implies the variance floor but provably cannot imply the third-moment floor: each level of the hierarchy carries strictly new information beyond the entire MGF family.
8. The audit, and the new sharp bounds
8.1 Every published bound is a shadow of one body
Bound Coordinates Verdict Saturator Merhav–Kafri variance floor [MK10] = exchange TUR [TGGL19] (μ, V) the exact lower edge P_{a_min} Merhav–Kafri ⟨ σ ^k⟩ ≥ h(μ)^k [MK10] (μ, ⟨ Proesmans–Van den Broeck / Hasegawa–Van Vu [PVdB17, HV19] (μ, V) valid, nowhere tight — Barato–Seifert (μ, V) not implied by the DFT violated by P_a Violation bound [S21b] (μ, Π) the complete region (both edges) P_{a_min} Merhav–Kafri §3 violation bounds [MK10] (m₊, s₊, Π₀) exact in one-sided moments; our (μ,V) forms: floor L(μ) (variance-blind), ceiling ½ − μ/(2x*) {0⁺, x*} family Skewness relation [S22] (μ, γ₁) mean-axis shadow of the exact boundary; Thm 5 strictly tighter three-outcome laws Information bound [S21a] (μ, h) the entire (attained) upper edge; nothing below interior maximizer MGF family [S23] (μ, G(α)) exact envelopes for all α (Thm 9); ⇏ third-moment floor P_{a_min} Tail bound [NH23] (μ, P(σ>t)) branch 1 of the exact envelope (Thm 7) three-outcome laws Jarzynski negative-tail e^{−t} (t, P(σ<−t)) strictly loosened form of the cap 1/(1+e^t) — Sign-entropy relation [HN25] (Σ, Λ) mixture-average of a pointwise inequality; exact floor φ(a_min(Σ)) P_{a_min} TUT [TUT] (μ, T, current) exact identity in the coordinate T = ⟨tanh(σ/2)⟩; the exchange TUR is its Jensen shadow q ∝ tanh(σ/2) at P_{a_min}
Why one distribution saturates nearly everything: at fixed mean, P_{a_min} is the corner of the body — and every tight bound that uses only the mean must, geometrically, touch the corner, whatever quantity it bounds. Bounds saturated elsewhere are exactly those whose functional is nonlinear (entropy, whose maximizer is a smooth interior distribution) or which condition on more than the mean (skewness, tails, and the violation ceiling at fixed variance: three-outcome saturators, one boundary piece out from the corner). ### 8.2 The new sharp results of this work
The same geometry yields bounds and exact regions that were not previously available; Table 2 collects the ones of most direct experimental use.
Quantity New exact result Improves on Probability of a negative excursion P(σ < −t) ≤ μ/(t(e^t − 1)) for μ ≤ m(t); ≤ 1/(1+e^t) always (Thm 7) Jarzynski's e^{−t}: at t = 1, 0.269 vs 0.368; with the mean, far stronger (μ = 0.1, t = 1: 0.058 vs 0.368) Upper tail P(σ > t) ≤ μ/(t(1−e^{−t})) for μ ≤ m(t); = 1 − L(μ) beyond; all attained (Thm 7) sharpness and the large-mean branch are new; the μ ≤ m(t) form appeared without sharpness in [NH23]; Markov's μ/t is invalid here Two-sided tail P( σ Skewness at known variance ⟨σ³⟩ ≥ μx², tanh(x/2)/x = μ/⟨σ²⟩ (Thm 5); forced left skew at the variance floor, forced right skew γ₁ ≳ √V/μ at large V ×38 tighter than the mean-only relation [S22] already at ⟨σ²⟩ = 4× its floor Fourth moment exact floor by a unique two-level law; on the skewness boundary, pinned: M₄ = M₂M₃/μ (Thm 6) new Violations at known variance floor L(μ) for every V (variance-blind); ceiling ½ − μ/(2x*) (zero-free class) new coordinates; translates [MK10 §3] into (μ, V) Sign entropy Λ ≥ φ(a_min(Σ)), attained; informative at every mean tightens the ln 2 − Σ relation [HN25], which is vacuous for Σ > ln 2 Engine moment law all moments of any three-outcome exchange engine: geometric, ratio = (level gap)² (Thm 4d) new prediction, testable per device Most fluctuation-permissive dissipation the variance floor peaks at ⟨σ⟩ = 2 exactly, V = 4(g₀²−1) ≈ 1.757 new closed form Length of time's arrow JS ≤ ln 2 − φ(a_min(J/2)) at total dissipation J, attained; ≈ J/8 near equilibrium (Thm 13) replaces the Feng–Crooks inequalities [FC08] with the exact attained boundary Forward–reverse overlap O ≥ 1 − T*(μ_p): the exact KL–TV boundary, attained by two-outcome pairs; ≈ 1 − √(μ/2) near equilibrium, e^{−μ} far (Thm 12) makes the dissipation–distinguishability relation [KPV07, FC08] extremal and saturable Null results the IFT alone fixes no (mean, variance) tradeoff (Thm 2); no moment is ever bounded above (Thm 4c); entropy has no lower edge (Thm 8); forward and reverse means unconstrained (Thm 11) delimits what any future DFT-only bound can say
The negative-excursion row deserves emphasis: apparent second-law violations are the most-quoted consequence of fluctuation theorems, and the classic exponential estimate is loose at every threshold — the exact cap 1/(1+e^t) cannot be improved, and at small mean dissipation the μ-dependent branch is smaller by orders of magnitude.
One recurring subtlety deserves a remark: an atom exactly at σ = 0 is a legal member of the class but a fixed point of the sign pairing, and two published bounds — the sign-entropy relation [HN25] and the Merhav–Kafri upper violation bound — are exact for distributions without it yet evaded by distributions with it; our envelopes hold on both classes, with closed forms for each.
9. Two-distribution fluctuation theorems: a dichotomy
The general (Crooks-type) fluctuation theorem relates two distributions — forward and reverse — by p(σ) = e^σ q(−σ). Everything so far assumed the time-symmetric case q = p. The general case reduces to it, in a precise and complete way.
First, the class itself: normalization of each member is the integral FT of the other, and q is determined by p (q(s) = e^s p(−s)), so a Crooks pair is nothing but a distribution p obeying ⟨e^{−σ}⟩ = 1, together with its exponentially tilted mirror. Reverse-process moments are tilted forward moments, ⟨s^k⟩_q = (−1)^k⟨σ^k e^{−σ}⟩_p, and the self-dual pairs q = p are exactly the DFT distributions — physically, the time-symmetric processes.
Theorem 10 (Crooks reduction). For sums of pair moments A_k = ⟨σ^k⟩_p + ⟨s^k⟩_q — and for every swap-symmetric linear pair statistic — the achievable set over the full Crooks class coincides with the achievable set over the DFT class (A_k = 2M_k). The entire hierarchy of this paper therefore transfers verbatim: the same half-line fibers with floors doubled, the same unique few-level saturators, the same pinning recursions, tails, entropy, and MGF envelopes — with the saturating pairs precisely the time-symmetric processes.
Proof. Split p by gap and sign into (p₊(a), p₋(a)) and take the swap average, (p̃₊, p̃₋) = ½(p₊ + e^a p₋, p₋ + e^{−a}p₊). This operation preserves normalization, the integral FT, and every swap-symmetric statistic — and the averaged pair satisfies p̃₊ = e^a p̃₋, which is the DFT. So every achievable sum-vector is already achieved by a DFT distribution; the reverse inclusion is immediate since DFT laws are (self-dual) Crooks pairs. ∎
Theorem 11 (the difference channel is empty at mean level). The achievable region of the pair of mean dissipations (μ_p, μ_q) is {(0,0)} ∪ (0,∞)²: beyond positivity, nothing links the forward and reverse averages. The construction is Lemma 1 in a new costume — a rare forward event at σ = A → ∞ with weight μ_p/A carries the entire forward mean while remaining tilt-invisible to the reverse process (its q-weight is suppressed by e^{−A}); the small negative mass balancing the integral FT contributes only μ_p/A to μ_q.
Physically, Theorem 11 is the exact statement behind a familiar piece of folklore: a process can dissipate arbitrarily much on average while its reverse dissipates arbitrarily little, which is why one-sided free-energy estimators can fail arbitrarily badly and bidirectional (Bennett-type) estimation is qualitatively necessary. The bidirectional side can be quantified exactly. The quality of Bennett-type estimation is governed by the overlap between the forward and reversed-reverse distributions, O = ∫min(p, q∘flip) — a linear pair statistic, hence inside the machinery — and the forward mean dissipation is precisely their relative entropy, μ_p = D(p‖q∘flip).
Theorem 12 (dissipation bounds distinguishability, exactly). At fixed forward dissipation μ, the overlap satisfies
O ≥ 1 − T*(μ), T*(μ) = max{ u − v : D(Bern(v)‖Bern(u)) = μ },
the classical sharp boundary between total variation and relative entropy; the floor is attained by two-outcome pairs, behaves as 1 − √(μ/2) near equilibrium (Pinsker's constant, now exact) and as e^{−μ} far from it, and the supremum of O is 1 at every μ (Lemma 1 once more). Consequently the fluctuation-theorem structure adds no constraint between dissipation and forward–reverse distinguishability beyond the relative-entropy identity itself: the Crooks class realizes the unrestricted information-theoretic optimum, because every binary pair is realized by a two-outcome member.
Proof. Since q(−σ) = e^{−σ}p(σ), the overlap is O = ⟨min(1, e^{−σ})⟩_p, linear in p. On the two-atom family {−r, s} with weights (v, 1−v), the integral FT gives O = 1 − (u − v) with u = v e^r, and the mean constraint becomes, verbatim, μ = D(Bern(v)‖Bern(u)); conversely every binary pair (v, u) arises this way. Optimality of two atoms follows from the three-constraint LP structure (verified to 10⁻⁶ across μ; Appendix E), and the supremum from the tilt-invisible escape. ∎
Theorem 12 makes exact the total-variation/hypothesis-testing reading of dissipation as distinguishability of the arrow of time (the relative-entropy identity underlying it goes back to [KPV07]): the theorem supplies the extremal boundary, its two-outcome saturators, and the statement that nothing sharper exists. For estimation, it is two-sided guidance: low dissipation guarantees overlap — Bennett-type estimation provably works with bounded effort near equilibrium — while high dissipation permits, but does not force, its failure.
A second, per-sample measure of the arrow of time is the Jensen–Shannon divergence JS(p, q∘flip) — the information one observation drawn from an equal mixture of forward and reversed-reverse carries about which it came from — related to the total dissipation, the Jeffreys divergence J = μ_p + μ_q, by the inequalities of Feng and Crooks [FC08]. Both quantities are swap-symmetric, and because the ratio of the pair is the deterministic factor e^{−σ}, JS is a linear pair statistic; Theorem 10 then reduces the question to the time-symmetric class, where JS collapses to the sign-entropy functional of Section 7: JS = ln 2·P(σ ≠ 0) − E_ν[φ(a)]. The exact region follows:
Theorem 13 (the exact length of time's arrow). Over the full Crooks class, at fixed total dissipation J:
0 ≤ JS ≤ ln 2 − φ(a_min(J/2)),
with the upper boundary attained — uniquely among gap distributions — by the two-outcome time-symmetric pair at gap a_min(J/2) (the corner, once more), growing as J/8 near equilibrium and saturating the information-theoretic ceiling ln 2 only as J → ∞ — parametrically, the boundary is the elementary binary curve (J, JS) = (8ε·artanh 2ε, ln 2 − H(½+ε)) in the sign bias ε; the lower boundary is approached but not attained (Lemma 1: the arrow of time can be made arbitrarily hard to read at any dissipation, but never free). This replaces the inequality relations of [FC08] with the attained boundary of the achievable (Jeffreys, Jensen–Shannon) region. (Verification: the JS identity to 10⁻¹⁰ on random mixtures; LP boundary agreement to 10⁻⁶ at J ∈ [0.1, 20]; Appendix E.)
Together the theorems of this section give a clean dichotomy: every constraint the two-distribution fluctuation theorem imposes on pair statistics lives in the swap-symmetric channel, where it is exactly the theory of this paper; the antisymmetric channel is free. Mixed statements — constraining symmetric and antisymmetric combinations jointly — are the natural continuation; there the governing structure is the exponential-polynomial Chebyshev family {σ^j, σ^j e^{−σ}}, the two-outcome pairs are exactly the binary hypothesis-testing pairs with (μ_p, μ_q) their (KL, reverse-KL), and the same moment-problem machinery applies; this is developed in a companion.
10. Discussion
The interface. The passage from physics to mathematics in this theory is one line long: the e^σ asymmetry produces the mean function a·tanh(a/2); its partial-fraction expansion identifies it as a positive sum of relaxation kernels with poles at the odd squares — the response function of a uniform string — and the rigidity of such sums does the rest. Every convexity lemma in the prior literature is one Wronskian of one hierarchy; every mean-only bound is one supporting plane touching one corner.
Inference. The hierarchy is constructive: moments measured through order n−1 sharply floor the unmeasured n-th, by solving a small system for the ⌈n/2⌉ levels (a few lines of root-finding; Appendix E). Conversely, observed saturation certifies the process to be an effective minimal engine, and on any boundary piece the recursions of Theorem 4(d) predict all higher moments — for any three-outcome exchange process, the full moment sequence from the first three.
Scope and extensions. Everything above concerns the distribution of a variable satisfying (1) exactly — time-symmetric protocols and exchange settings. Two extensions are natural and developed elsewhere: joint bodies for currents alongside σ (the TUT row of Table 1 is the first slice; the matrix-valued form of [TGGL19] and the involution framework of [S22b] are the right anchors), the transfer of these results to variables that do not satisfy (1) but share its reversal structure, as in general quantum trajectories. The two-distribution case is settled in Section 9; its mixed symmetric–antisymmetric bodies remain open. On the mathematical side, the moment problem for exponentially tilted symmetric distributions — the statistician's reading of (1) — appears untouched, and the string picture suggests a generalization: other fluctuation symmetries p(σ)/p(−σ) = e^{c(σ)}, c odd, should admit the same complete theory whenever their mean function is a positive sum of relaxation kernels. Finally, the proof spine is short and classical — one hyperbolic inequality, one monotonicity, Jensen, a partial-fraction expansion, a determinant identity, and zero-counting — and a machine formalization has been scaffolded.
Appendix A. The simplex and the escaping atom
A.1 (Proposition 1). Given p ∈ 𝒫, let ν be the law of |σ|. The identity p(dσ) = e^σ p(−dσ) fixes, for ν-a.e. a > 0, the conditional sign law P(sign = ±1 | |σ| = a) = w±(a); at a = 0 there is nothing to fix. Hence p = ∫P_a dν uniquely; conversely any ν yields p ∈ 𝒫, and the map is affine and injective. Extreme points of the image of a bijective affine map of a simplex are the images of point masses. ∎
A.2 (Lemma 1, precise form). Let Ψ, Ξ₁, …, Ξ_r be functionals linear in ν, and (a_L) a family of locations with Ψ-payload ψ(a_L) → ∞ while the Ξ_i-payloads are o(ψ(a_L)). Then measures ν_L = (1 − ε_L)ν₀ + ε_Lδ_{a_L} with ε_Lψ(a_L) = const satisfy Ξ_i(ν_L) → Ξ_i(ν₀) while Ψ(ν_L) is held (or driven) as desired. Instantiations: (i) Ψ = n-th moment, Ξ = lower moments, a_L → ∞, ε ∼ L^{−(n−1/2)} (Theorem 4(c)); (ii) in the integral-only class, Ψ = the IFT weight e^{−s}, location s → −∞ (Theorem 2); (iii) Ψ = mean, Ξ = Shannon-entropy contribution ε log(1/ε) (Theorem 8(ii)); (iv) Ψ = mean, Ξ = tail weight w∓ at the escaping gap (Theorem 7 infima and the lower-tail ceiling); (v) Ψ = forward mean, Ξ = the tilted (reverse) mean, payload σe^{−σ} at σ = A → ∞ (Theorem 11). ∎
Appendix B. The Wronskian–Hankel identity (Theorem 3)
B.1 (Lemma). φ(h) := a² as a function of h = m(a) is strictly convex: φ′(h) = 2a/m′(a) and m′ − a m″ = tanh x − x sech²x + (a²/2)sech²x·tanh x > 0 (x = a/2) by sinh 2x ≥ 2x. (Equivalent to W₂ > 0 below.)
B.2 (Statement). For any positive measure ρ with the needed moments, H(t) = ∫ t/(t+b)dρ, T_m(t) = ∫ b/(t+b)^m dρ, and the interleaved system (1, H, t, tH, …, n+1 functions; p = ⌊n/2⌋, l = ⌈n/2⌉):
W(u₀,…,u_n) = (∏{j=0}^{n} j!) · det[T{i+j+2}]₀^{l−1} (n odd), (∏{j=0}^{n} j!) · det[T{i+j+3}]₀^{l−1} (n even).
For ρ = 4Σδ_{b_k} this is the form quoted in Section 5 (the mass 4 contributes 4^l). The Hankel determinants are those of the moments of the positive measure x^{2,3}·dτ_t, τ_t = image of b·dρ under b ↦ 1/(t+b); by Heine's identity they are strictly positive whenever ρ has at least l support points.
B.3 (Proof). Step 1 — polarization. Both sides are symmetric l-linear forms in ρ (l columns/entries carry ρ linearly); a symmetric l-linear form is determined by its values on atomic ρ = Σᵢ₌₁^l mᵢδ_{bᵢ} with free masses. Truncations ρ_K → ρ converge with all derivatives locally uniformly, extending the identity to (2). Step 2 — Cauchy–Vandermonde basis. For atomic ρ, partial fractions give span{1, H, t, tH, …} = span{1, t, …, t^p} ∪ {(t+bᵢ)⁻¹}, dimensions matching; the change-of-basis determinant factorizes by Laplace (polynomial columns have no Cauchy component) as ±(∏mᵢbᵢ)·V(b), V the Vandermonde. Step 3 — the CV Wronskian. W(1,…,t^p, (t+b₁)⁻¹,…) = ±(∏{j=0}^p j!)(∏{r=0}^{l−1}(p+1+r)!)·(∏yᵢ^{p+2})·V(y), yᵢ = 1/(t+bᵢ): the first p+1 derivative rows triangularize the polynomial block; the remaining rows see only the Cauchy columns with closed-form derivatives. Step 4 — assembly. V(b) = ±V(y)∏yᵢ^{−(l−1)}, so W_n = ±(∏{j=0}^{p+l} j!)·∏mᵢbᵢyᵢ^{p+3−l}·V(y)². The l-atom Heine evaluation of the Hankel side is (∏mᵢbᵢyᵢ^{c})·V(y)². Matching exponents forces c = p+3−l — which is 2 for odd n and 3 for even n — and p + l = n telescopes the factorials to ∏{j=0}^n j!. The sign is fixed by the fully symbolic cases n = 2, 3. ∎ (Independent verification: symbolic (n = 2, 3); exact rational polynomial-identity testing at 300 points × all mass corners (n = 4–7); 50-digit numerics with (2) and the closed-form constant (n = 2–9).)
Appendix C. Zero-counting and the principal representations (Theorems 4–6)
With the Chebyshev property in hand (Theorem 3), the fiber structure is classical Markov–Krein theory on [0, ∞) [KS66, KN77]; we record the conventions. Minimizing E[u_n] under the n constraints E[u_k] = μ_k (k < n) has a dual "polynomial" u_n − Σc_ku_k ≥ 0 vanishing on the optimal support; a Chebyshev system admits at most n zeros counted with multiplicity (interior zeros of a nonnegative combination are double; the endpoint 0 is simple, and since every u_k (k ≥ 1) vanishes at 0, an atom at 0 forces c₀ ≤ 0 with equality). The unique pattern matching n active constraints is: (n = 2m) m interior double zeros — atoms {x₁,…,x_m}; (n = 2m+1) the endpoint plus m interior doubles — atoms {0, x₁,…,x_m}. Dimension counts (weights + locations = n) give generic uniqueness; strict convexity along the curve gives it always. Recession: Lemma 1(i). Pinning (Theorem 4(d)): on a stratum with squared levels {rᵢ}, ∏(r − rᵢ) annihilates the support, so E[r^{j}·∏(r−rᵢ)·(1 or h)] = 0 — the stated recursions; on {0, x} this is r² = x²r, the geometric law. Explicit instance solutions: Theorem 5's support {0, x} yields x from tanh(x/2)/x = μ/M₂, weight μ/m(x), floor μx²; Theorem 6's two-atom system is solved numerically (Appendix E), with the boundary pinning M₄ = M₂M₃/μ following from x² = M₃/μ on the adjacent stratum.
Appendix D. Convexity of the MGF maps (Theorem 9)
Let x = a/2, t = tanh x, s = sech x, β = α + ½, F = cosh(βa)/cosh(a/2). Direct computation gives
(m′F″ − m″F′)·cosh(a/2) = cosh(βa)[β²m′ + P] − β sinh(βa)·[m′t + m″],
with the exact collapses m′t + m″ = t² + s² = 1 and m′(2t²−1)/4 + m″t/2 = (t − xs²)/4 = P ≥ 0 (Lemma B.1's inequality). Dividing by cosh(βa), convexity is Φ(a, β) = β²m′ + P − β tanh(βa) ≥ 0. Now Φ(a, ½) ≡ 0 identically, and since m′(a) = n(x) with n(y) := (y·tanh y)′,
∂Φ/∂β = 2β·n(x) − n(2βx) = 2βx·[ ρ(x) − ρ(2βx) ], ρ(y) := n(y)/y = tanh(y)/y + sech²y,
which is strictly decreasing (tanh(y)/y = ∫₀¹ sech²(sy) ds and sech²y both are). Hence ∂Φ/∂β ⋚ 0 as β ⋚ ½: at every fixed a, Φ attains its global minimum in β exactly at β = ½, where it vanishes. ∎ Equality only at the constant members β = ±½, where F is affine in m and degeneracy is forced.
Appendix E. Numerical protocol
All boundaries were verified by linear programming over the mixing measure (grids of ~1.4×10⁴ points on a ∈ [0, 60]; HiGHS), and independently by LPs in σ-space with (1) imposed as explicit linear constraints; floors matched closed forms to grid precision (10⁻⁶–10⁻⁸) with optimal supports coinciding with the predicted atoms; maxima scaled with the grid cap, confirming the recession structure. Extremal representations for n = 4–6 were solved by root-finding seeded from LP supports. Wronskians were evaluated at 50-digit precision using the exact derivative formulas H^{(i)} = 4(−1)^{i+1}i!T_{i+1} with tail-controlled truncation of (2). For Section 9, the pair class was verified by LPs (including the overlap floor of Theorem 12 against the binary-optimum curve, agreement 10⁻⁶ at μ ∈ [0.02, 5], and the Jensen–Shannon boundary of Theorem 13, agreement 10⁻⁶ at J ∈ [0.1, 20]) over full-line grids with the integral-FT row as an explicit constraint: sum-moment floors matched the doubled DFT values to 10⁻⁵ with symmetric ±a optimal supports (Theorem 10), and the minimum of μ_q at fixed μ_p matched the μ_p/A grid-cap prediction exactly (Theorem 11). Code accompanying the paper reproduces every figure and table.
Appendix F. Closed forms and asymptotics
V_min(μ): parametric (μ, V) = (m(a), a²sech²(a/2)); expansions 2μ − (2/3)μ² + O(μ³) and 4μ²e^{−μ}(1 + o(1)); peak at μ = 2 exactly, value 4(g₀²−1), g₀tanh g₀ = 1 (the stationarity condition reduces to tanh²+sech² = 1). Third-moment floor: ⟨σ³⟩ ≥ μx², tanh(x/2)/x = μ/M₂; central form κ₃ ≥ μ(x² − 3M₂ + 2μ²); κ₃ = −2μV on the variance floor; γ₁ ≃ √V/μ at large V. Tail constants: w₊(t)/m(t) = 1/(t(1−e^{−t})); w₋(t)/m(t) = 1/(t(e^t−1)). Stratum recursions: {0,x}: M_{k+2} = (M₃/μ)M_k; {x,y}: M_{k+4} = (x²+y²)M_{k+2} − x²y²M_k. Entropy identity: w₊(a)e^{φ(a)} = e^{(a−m(a))/2}. Arrow-of-time region: JS = ln2·P(σ≠0) − E_ν[φ]; max JS at Jeffreys J is ln2 − φ(a_min(J/2)); exact parametric form (J, JS) = (8ε·artanh 2ε, ln2 − H(½+ε)); expansions J/8 − J²/192 (1% for J ≤ 15) and ln2 − (1+J/2)e^{−J/2} (1% for J ≥ 9), jointly covering all J to 1%; inf 0. Overlap floor: min O at fixed μ equals 1 − T*(μ), T* the binary KL–TV boundary; asymptotics 1 − √(μ/2) and e^{−μ}. Sign-entropy floor: m(a) + φ(a) ≥ ln 2 pointwise; exact floor Λ ≥ φ(a_min(Σ)) on zero-atom-free laws. Violation region at fixed (μ, V): floor L(μ) for all V (attained iff V = V_min); ceiling 1 − (μ/m(x*))w₊(x*) on the full class, ½ − μ/(2x*) on the zero-free class, x* the Theorem-5 point; the latter equals Merhav–Kafri's (m₊,s₊) bound at the maximizer.
References (to be completed at submission)
[MK10] N. Merhav, Y. Kafri, J. Stat. Mech. (2010), arXiv:1010.2319. — [TGGL19] A. M. Timpanaro, G. Guarnieri, J. Goold, G. T. Landi, PRL 123, 090604 (2019) (in full a matrix-valued TUR: variances and correlations; anchor for the joint-current extension). — Y. Zhang, arXiv:1910.12862 (simplified derivation). — [PVdB17] K. Proesmans, C. Van den Broeck, EPL 119, 20001 (2017). — [HV19] Y. Hasegawa, T. Van Vu, PRL 123, 110602 (2019). — [S21a] D. S. P. Salazar, PRE 103, 022122 (2021). — [S21b] D. S. P. Salazar, PRE 104, L062101 (2021). — [S22] D. S. P. Salazar, PRE (2022), arXiv:2208.11206. — [S22b] D. S. P. Salazar, PRE 106, L062104 (2022). — [S23] D. S. P. Salazar, PRE 107, arXiv:2302.02998 (2023). — [NH23] T. Nishiyama, Y. Hasegawa, PRE 108, 044139 (2023). — [HN24] Y. Hasegawa, T. Nishiyama, PRL 133, 247101 (2024). — [HN25] Y. Hasegawa, T. Nishiyama, arXiv:2502.06174. — [TUT] J. Ray, A. B. Boyd, G. Guarnieri, J. P. Crutchfield, csc.ucdavis.edu/~cmg/papers/tut.pdf. — [NRJ17] I. Neri, É. Roldán, F. Jülicher, PRX 7, 011019 (2017). — C. H. Bennett, J. Comput. Phys. 22, 245 (1976) (bidirectional estimation). — [KPV07] R. Kawai, J. M. R. Parrondo, C. Van den Broeck, PRL 98, 080602 (2007). — [FC08] E. H. Feng, G. E. Crooks, PRL 101, 090602 (2008). — [KS66] S. Karlin, W. Studden, Tchebycheff Systems (1966). — [KN77] M. G. Krein, A. A. Nudelman, The Markov Moment Problem (1977). — Karlin, Total Positivity (1968). — [SSV] Schilling, Song, Vondraček, Bernstein Functions (2012). — Evans–Searles; Jarzynski–Wójcik; Crooks; Jarzynski; Barato–Seifert; Gingrich et al.; Campisi–Hänggi–Talkner.