A digestion of the proof of Sendov’s conjecture

· What's new ·

24 min read Original article ↗

This post concerns the following conjecture of Sendov, as well as its strengthening by Phelps–Rodriguez:

Conjecture 1 (Sendov’s conjecture) Let {n \geq 2}, and let {p : {\bf C} \rightarrow {\bf C}} be a degree {n} polynomial with all zeroes in the unit disk. Then for every zero {a} of {p}, there exists a critical point {\zeta} of {p} with {|\zeta-a| \leq 1}.
Conjecture 2 (Phelps–Rodriguez conjecture) Let {n \geq 2}, and let {p : {\bf C} \rightarrow {\bf C}} be a degree {n} polynomial with all zeroes in the unit disk. Then for every zero {a} of {p}, there exists a critical point {\zeta} of {p} with {|\zeta-a| < 1}, unless {a} is on the unit circle and {p} is a scalar multiple of {z^n - a^n}.

By applying a rotation around the origin, we can normalize {a} to be a real number with {0 \leq a \leq 1}.

From the work of Rubinstein, both conjectures were already established in the {a=1} case, so one can restrict to the {0 \leq a < 1} case. Both of these conjectures then follow from

Conjecture 3 (Sendov’s conjecture in interior) Let {n \geq 2}. Let {p : {\bf C} \rightarrow {\bf C}} be a degree {n} polynomial with all zeroes in the unit disk. Then if {0 \leq a < 1} is a zero of {p}, there exists a critical point {\zeta} of {p} with {|\zeta-a| < 1}.

All three of these conjectures were established for {n \leq 8} (in a sequence of papers culminating in this paper of Brown and Xiang) and for sufficiently large {n} (in a paper of myself, which in turn built upon several partial results in this setting). This left the case of intermediate {n} to be settled. My arguments used some qualitative ingredients (most notably analytic continuation) and as such did not easily lend themselves to quantifying the threshold of {n} above which the argument was valid.

Recently, Lech Mazur was able to use an AI tool to resolve Sendov’s conjecture for all {n \geq 2}, with the proof verified in Lean. However, the AI-generated proof was not human-digested to be in the form of a publication-ready preprint; and it has taken me several days (with heavy AI assistance) to perform such a digestion, to place the proof in proper context with previous literature and to simplify and streamline the argument to highlight the main ideas. (Note: the above chat log only represents a portion of the digestion work: the rest was performed with pen and paper, or using some further AI agents.) The same arguments also give a new proof of Rubinstein’s theorem, which I also give below the fold.

One consequence of this digestion is that the argument in fact demonstrates Conjecture 3, and thus resolves both the Sendov conjecture and the Phelps–Rodriguez conjecture in full generality.

The proof ends up being remarkably elementary. No complex analysis is used other than the fundamental theorem of algebra (and very basic facts about Möbius transformations); and the deepest inequality used as input is the Maclaurin inequality (and we only need a special case of that inequality which can be derived from the arithmetic mean-harmonic mean inequality and an induction argument).

Using an AI agent, I have been able to formalize the entire argument in Lean, extended to {n \geq 2} by some minor modifications to the proof. This formalization is more streamlined than the original formalization (it has about 15,000 lines of code, compared with around 90,000 for the original proof).

We now prove Conjecture 3. The {n \leq 4} cases have long been known but need to be treated separately; a short proof using the machinery developed here is provided at the end of the post. Suppose now that we have a counterexample for some {n \geq 5}, thus one can find a degree {n} polynomial {p} with zeroes

\displaystyle  a, z_1, \dots,z_{n-1}

for some {0 \leq a < 1} and {z_j}, {j=1,\dots,n-1} in the closed unit disk, whose critical points all lie a distance at least {1} from {a}. We use {O()} notation here in the non-asymptotic sense, thus {X = O(Y)} means that {|X| \leq C Y} for some absolute constant {C} (independent of {n}). We will also use the notation {O_{\leq}(Y)} to denote a quantity that is bounded in magnitude by {Y}.

To capture the fact that the critical points lie at a distance at least {1} from {a}, we write these critical points as

\displaystyle  a - \frac{1}{q_1}, \dots, a - \frac{1}{q_{n-1}}

for some (non-zero) {q_1,\dots,q_{n-1}} in the closed unit disk.

Example 4 If {p(z) = z^n - 1} and {a=1}, then {z_1,\dots,z_{n-1}} are the non-trivial {n^{th}} roots of unity, while the {q_1,\dots,q_{n-1}} are all equal to {1}. Strictly speaking this is not actually a counterexample to Conjecture 3, because {a} is not strictly less than one; nevertheless this is an important motivating near-counterexample for the arguments below.
Example 5 A generalization of the previous example was studied in Section 4 of my paper. Here one took

\displaystyle  p(z) = (z + \frac{c_2}{n})^{n-m} P(z) - (a + \frac{c_2}{n})^{n-m} P(a)

where {n} was an asymptotic parameter going to infinity,

\displaystyle  P(z) = (z-\lambda_1) \dots (z - \lambda_m)

was a low-degree polynomial for some {m=O(1)},

\displaystyle  a = 1 - \frac{c_1}{n},

and {c_1,c_2 > 0} were constants. This polynomial has a zero at {a}, {n-m-1} critical points at {-\frac{c_2}{n}}, and {m} additional critical points near {\lambda_1,\dots,\lambda_m}. If all the critical points were at distance at least one from {a}, one would have

\displaystyle  c_2 \geq c_1

and

\displaystyle  |1-\lambda_j| \geq 1 - o(1)

while if all the zeroes were in the unit disk, the calculations in my paper showed that

\displaystyle  c_2 - c_1 - c_2 \cos \theta + \sum_{j=1}^m \log |\frac{1-\lambda_j}{e^{i\theta}-\lambda_j}| \leq o(1) \ \ \ \ \ (1)

Here {o(1)} denotes a quantity that goes to zero as {n \rightarrow \infty}. If one ignores the {o(1)} errors, one can show that these conditions are only simultaneously feasible if {c_1=c_2} and all the {\lambda_j} vanish, but the argument was somewhat subtle (I had to proceed by inspecting the second Fourier coefficient of (1)). This illustrates the fact that the regime {a = 1 - O(1/n)} is particularly delicate.

We now have two sets of points in the closed unit disk: {z_1,\dots,z_{n-1}} and {q_1,\dots,q_{n-1}}. They “communicate” with each other through the polynomial {p} and its first derivative {p'}, both of which can be expressed in terms of either set of points (as well as {n} and {a}). Indeed, if we normalize {p} to be monic, then we can factor {p} in terms of the zeroes as

\displaystyle  p(z) = (z-a) \prod_{j=1}^{n-1} (z-z_j) \ \ \ \ \ (2)

and thus upon differentiating

\displaystyle  p'(z) = \left(\prod_{j=1}^{n-1} (z-z_j)\right) \left(1 + (z-a) \sum_{j=1}^{n-1} \frac{1}{z-z_j}\right). \ \ \ \ \ (3)

Here and in the sequel we adopt the convention of removing singularities when dealing with expressions that involve multiplication by both {\frac{1}{z-z_j}} and {z-z_j}, by cancelling such terms first in the event that {z=z_j}.

In a similar vein, {p'} can be factored

\displaystyle  p'(z) = n \prod_{j=1}^{n-1} \left(z - a + \frac{1}{q_j}\right) \ \ \ \ \ (4)

and thus on integrating (and using {p(a)=0})

\displaystyle  p(z) = (z-a) \int_0^1 n \prod_{j=1}^{n-1} \left(t(z - a) + \frac{1}{q_j}\right)\ dt. \ \ \ \ \ (5)

It is convenient to rule out the easy case {a=0} right away. In this case we see from (3), (4) that

\displaystyle  p'(0) = \prod_{j=1}^{n-1} (-z_j) = n \prod_{j=1}^{n-1} \frac{1}{q_j}

which is absurd since the first product has magnitude at most one, and the second product has magnitude at least one. Thus we can assume henceforth that {a>0}.

By inspecting {p} or {p'} at various natural locations, we can thus obtain a number of identities relating the {z_j} to the {q_j}. We record the ones that we actually need here:

Lemma 6 Let {F} denote the function

\displaystyle  F(t) := \prod_{j=1}^{n-1} (1 - atq_j). \ \ \ \ \ (6)

  • (i) (Centroid identity) We have

    \displaystyle  \frac{1}{n} \left(a + \sum_{j=1}^{n-1} z_j\right) = \frac{1}{n-1} \sum_{j=1}^{n-1} \left(a - \frac{1}{q_j}\right). \ \ \ \ \ (7)

    That is to say, the centroid of the zeroes equals the centroid of the critical values.
  • (ii) (Polar identity) We have

    \displaystyle  \prod_{j=1}^{n-1} \frac{1-az_j}{a-z_j} = \int_0^1 \prod_{j=1}^{n-1} (t (1-a^2) q_j + a)\ dt. \ \ \ \ \ (8)

  • (iii) (First origin identity) We have

    \displaystyle  (-1)^{n-1} \prod_{j=1}^{n-1} z_j = \frac{n}{\prod_{j=1}^{n-1} q_j} \int_0^1 F(t)\ dt. \ \ \ \ \ (9)

  • (iv) (Second origin identity) We have

    \displaystyle  (-1)^{n-1} \prod_{j=1}^{n-1} z_j \left( 1 + a \sum_{j=1}^{n-1} \frac{1}{z_j} \right) = \frac{n}{\prod_{j=1}^{n-1} q_j} F(1). \ \ \ \ \ (10)

    (Again, we are using the convention of removing singularities to deal with the case where some of the {z_j} vanish.)

Proof: For (i), we inspect the behavior of {p'(z)} as {z \rightarrow \infty}. From (2) we have

\displaystyle  p(z) = z^n - \left( a + \sum_{j=1}^{n-1} z_j \right) z^{n-1} + O(z^{n-2})

and thus on differentiating term by term

\displaystyle  p'(z) = n z^{n-1} - (n-1) \left( a + \sum_{j=1}^{n-1} z_j \right) z^{n-2} + O(z^{n-3}).

Meanwhile, from (4) we have

\displaystyle  p'(z) = n z^{n-1} - n \left(\sum_{j=1}^{n-1} \left(a - \frac{1}{q_j}\right)\right) z^{n-2} + O(z^{n-3}).

Comparing coefficients, we obtain the claim.

For (ii), we consider the expression {p(1/a) / p'(a)}. On the one hand, from (2), (3) one has

\displaystyle  \frac{p(1/a)}{p'(a)} = \frac{(1/a - a) \prod_{j=1}^{n-1} (1/a - z_j)}{\prod_{j=1}^{n-1} (a-z_j)}.

(Note from hypothesis that {a} cannot be a critical point, so the denominator is non-zero.) On the other hand, from (4), (5) one has

\displaystyle  \frac{p(1/a)}{p'(a)} = \frac{(1/a - a) \int_0^1 n \prod_{j=1}^{n-1} \left(t\left(\frac{1}{a} - a\right) + \frac{1}{q_j}\right)\ dt}{n \prod_{j=1}^{n-1} \frac{1}{q_j}}.

Equating the two identities, we obtain (ii) after some algebra.

For (iii), we evaluate {p(0)}. From (2) we have

\displaystyle  p(0) = -a (-1)^{n-1} \prod_{j=1}^{n-1} z_j

while from (5) we have

\displaystyle  p(0) = -a \int_0^1 n \prod_{j=1}^{n-1} \left(-at + \frac{1}{q_j}\right)\ dt.

Equating the two identities, we obtain (iii) after some algebra using (6).

For (iv), we similarly evaluate {p'(0)}. From (3) we have

\displaystyle  p'(0) = (-1)^{n-1} \left(\prod_{j=1}^{n-1} z_j\right) \left(1 + a \sum_{j=1}^{n-1} \frac{1}{z_j}\right)

while from (4) one has

\displaystyle  p'(0) = n \prod_{j=1}^{n-1} \left(- a + \frac{1}{q_j}\right).

Equating the two identities, we obtain (iv) after some algebra using (6). \Box

Remarkably, the polynomial {p} will play no further role in the argument: the identities in (i)-(iv), together with the hypotheses that {0 \leq a < 1} and {z_j, q_j} lie in the closed unit disk, will be sufficient by themselves to obtain a contradiction.

Example 7 Continuing the example in Example 4, in (i) both sides vanish. In (ii), both sides are equal to one. For (iii) and (iv), we have {F(t) = (1-t)^{n-1}}, with both sides of (iii) equal to one, and both sides of (iv) equal to zero.
Remark 8 The centroid identity is extremely classical, going back to this 1948 paper of Popoviciu. The comparison of the polynomial at a location {a} and at the polar inversion {1/a} of that location across the closed unit disk is a familiar trick in the literature; see, e.g., Lemma 5 and Theorem 8 of Dégot. The specific form of the polar identity is implicit in the first part of Section 5 of Mazur’s AI-generated proof, while the origin identities are extracted from equation (6.3) of that proof. The first origin identity is also very close to Theorem 6 of Dégot, while the second origin identity is similar to some identities appearing in the proof of Lemma 6 of Dégot, as well as the work of Mir–Nazir–Wani and (in the {a=1} case) Rubinstein. The work of Meir–Sharma and Mir–Nazir–Wani also contain several further identities relating the {z_j} to the {q_j}; see in particular Lemma 15 below. Variants of (5) also appear in Proposition 10 of Miller.
Remark 9 The first origin identity (9) is already strong enough to handle asymptotically all examples of the form in Example 5, except in the endpoint case where {c_1, c_2} vanish and the {\lambda_j} are all {0}. Indeed, as the {z_j, q_j} are in the closed unit disk, (9) implies that

\displaystyle  n \left|\int_0^1 F(t)\ dt\right| \leq 1.

On the other hand, routine calculations (omitted here) show that

\displaystyle  n \int_0^1 F(t)\ dt = 1 + \frac{c_2 + \sum_{j=1}^{m} \frac{-\lambda_j}{1-\lambda_j}}{n} + O\left(\frac{1}{n^2}\right)

leading asymptotically to the constraint

\displaystyle  c_2 + \mathrm{Re} \sum_{j=1}^{m} \frac{-\lambda_j}{1-\lambda_j} \leq 0.

But all terms here are non-negative (since {|1-\lambda_j| \geq 1}), so this forces a contradiction unless {c_2} (and hence also {c_1}) and the {\lambda_j} all vanish.

As mentioned in Example 5, the most delicate regime occurs when {a = 1 - O(1/n)}. It is convenient to introduce the normalized version

\displaystyle  \alpha := \frac{n-1}{2} (1-a^2), \ \ \ \ \ (11)

of {a}, thus {0 < \alpha < \frac{n-1}{2}}, and the case {a = 1 - O(1/n)} corresponds to {\alpha = O(1)}. Informally, {\alpha} measures how close {a} is to {1} (at the scale of {O(1/n)}).

A key role in the argument will be played by the mean

\displaystyle  x + iy := \frac{1}{n-1} \sum_{j=1}^{n-1} q_j \ \ \ \ \ (12)

of the {q_j}, particularly the real part {x}. As the {q_j} all lie in the unit disk, the mean {x+iy} does also, so that

\displaystyle  -1 \leq x \leq 1

and

\displaystyle  |y| \leq \sqrt{1-x^2}. \ \ \ \ \ (13)

On the other hand, in the example in Example 4, {x} is equal to the extremal value of {1}, and {y=0}. In Example 5, we have {x = 1 - O(1/n)} (and {y = O(1/n)}).

It will be convenient to work with the quadratic polynomial

\displaystyle  \beta(t) := 1 - 2atx + a^2 t^2 = 1 - x^2 + (x - at)^2 \ \ \ \ \ (14)

with a particular emphasis on the value at {t=1}:

\displaystyle  \begin{array}{rl} \beta(1) &= 1 - 2ax + a^2 \\ &= (1-a)^2 + 2a(1-x) = 1 - x^2 + (x-a)^2. \end{array} \ \ \ \ \ (15)

One should primarily think of {\beta(1)} as a measure of how close {x} is to {1}. Clearly we have

\displaystyle  \beta(t) > 0

for all {0 \leq t \leq 1} (note that {at} is strictly less than {1}).

The arguments will revolve around the relationship between {\alpha} and {\beta(1)}. Specifically, we will establish the following two inequalities below the fold. The first inequality, which we call the “polar inequality”, comes in three forms:

Proposition 10 (Polar inequality)
  • (i) (Raw polar inequality) We have

    \displaystyle  1 \leq \int_0^1 (a^2 + 2 a t x (1-a^2) + t^2 (1-a^2)^2)^{\frac{n-1}{2}}\ dt. \ \ \ \ \ (16)

  • (ii) (Polar inequality in {\alpha}, {\beta(1)} form) We have

    \displaystyle  1 < \int_0^1 \exp ( \alpha (-1 + (2-\beta(1)) t))\ dt. \ \ \ \ \ (17)

  • (iii) (Simplified polar inequality) We have

    \displaystyle  0 < \beta(1) < \min\left( \frac{\alpha}{3+\alpha}, 1 - \frac{\log \alpha}{\alpha} \right). \ \ \ \ \ (18)

    In particular, since {1 - \beta(1) = 2a(x-a/2)}, one has

    \displaystyle  x > a/2 > 0.

It will be the inequality (18) that we use in practice, but it will be derived from (17), which in turn is a consequence of (16), which will follow from the polar identity (8) together with the fact that the {z_j} and {q_j} lie in the unit disk. The bound (18) is only slightly weaker than (17); see the (Gemini-generated) image below.

I was not able to find an exact duplicate of the above polar inequalities in past literature, but the paper of Dégot contains several similar inequalities. The inequality (16) was extracted from (5.1) of Mazur’s AI-generated proof; the subsequent bounds (17), (18) arose from my attempts to simplify the arguments after that point.

The second inequality, which is more difficult, also will come in several forms:

Proposition 11 (Origin inequality) Let {n \geq 5}.
  • (i) (Raw origin inequality) We have

    \displaystyle  2 \alpha + ax \leq \frac{1-x^2}{2(n-1)} + a^2 n(n-1) \int_0^1 t \beta(t)^{\frac{n-2}{2}}\ dt. \ \ \ \ \ (19)

  • (ii) ({\alpha} bound) We have

    \displaystyle  \alpha \leq 17. \ \ \ \ \ (20)

  • (iii) (Origin inequality in {\alpha}, {\beta(1)} form) We have

    \displaystyle  \begin{array}{rl} 1 \leq & \frac{\beta(1)}{2\alpha(1-\beta(1))} + \frac{\beta(1)}{4\alpha} + \frac{1}{2(n-1)} + \frac{\beta(1)}{4\alpha(n-1)} \\ & + \frac{a^4 n (n-1) (n-2) \beta(1)}{4\alpha} \int_0^1 t^3 \beta(t)^{\frac{n-4}{2}}\ dt. \end{array} \ \ \ \ \ (21)

Part (i) (which was extracted with some effort from Section 6 of the original AI-generated argument) will be deduced from the first and second origin identities (9), (10), as well as the centroid identity (7). Part (ii) will follow from (i) and the polar inequality (18), while part (iii) is an elementary consequence of (i).

As it turns out, the last three terms in (21) are asymptotically negligible as {n \rightarrow \infty}. Dropping those terms gives a competing feasibility region for {\alpha} and {\beta(1)} which is disjoint from the one coming from the polar inequality (17) (or (18)):

This already suggests that one can use this approach to recover my previous result on Sendov’s conjecture holding for all sufficiently large {n}. In fact, even with the three error terms in (21) added, there is enough room between the two inequalities (18), (21) to obtain a contradiction for all {n \geq 5} (using the additional bound {\alpha \leq 17} to control these errors), although showing this for medium-sized {n} (such as {5 \leq n \leq 200}) requires a certain amount of computer assistance.

For fixed {\alpha}, the right-hand side of (21) is monotone increasing in {\beta(1)} (or equivalently, monotone decreasing in {x}). In view of (18), we can thus replace {\beta(1)} by {\frac{\alpha}{3+\alpha}} in this inequality, so that {ax = 1 - \frac{\alpha}{n-1} - \frac{\beta(1)}{2}} is replaced by

\displaystyle c(\alpha) := 1 - \frac{\alpha}{n-1} - \frac{\alpha}{2(3+\alpha)},

and {\beta(t) = 1 - 2axt + a^2t^2} replaced by {1 - 2c(\alpha) t + a^2 t^2}. The inequality (21) then becomes an inequality involving only {\alpha} and {n}:

\displaystyle  \begin{array}{rl} 1 \leq & \frac{1}{6} + \frac{1}{4(3+\alpha)} + \frac{1}{2(n-1)} + \frac{1}{4(n-1)(3+\alpha)} \\ & + \frac{a^4 n (n-1) (n-2)}{4(3+\alpha)} \int_0^1 t^3 (1 - 2c(\alpha) t + a^2 t^2)^{\frac{n-4}{2}}\ dt. \end{array} \ \ \ \ \ (22)

We also note that the bounds {0 \leq c(\alpha) \leq xa} force the constraint

\displaystyle  c(\alpha)^2 \leq a^2 = 1 - \frac{2\alpha}{n-1}.

This prevents {\alpha} from getting too close to the upper limit {\frac{n-1}{2}} (or {a} getting too close to zero).

We can now eliminate all large degrees, e.g., {n > 200}, as follows. The quadratic {1 - 2c(\alpha) t + a^2 t^2} attains its minimum at {t =c(\alpha)/a^2}. For {t \leq c(\alpha)/a^2} we have

\displaystyle  1 - 2c(\alpha) t + a^2 t^2 \leq 1 - c(\alpha) t \leq \exp( - c(\alpha) t)

while for {c(\alpha) / a^2 < t \leq 1} (if this region is non-vacuous) we can bound the quadratic by its value {\frac{\alpha}{3+\alpha}} at {t=1}. Thus

\displaystyle  \begin{array}{rl} & \int_0^1 t^3 (1 - 2c(\alpha) t + a^2 t^2)^{\frac{n-4}{2}}\ dt \\ \leq & \int_0^\infty t^3 \exp\left( - \frac{n-4}{2} c(\alpha) t\right)\ dt + \int_0^1 t^3 \left(\frac{\alpha}{3+\alpha}\right)^{\frac{n-4}{2}}\ dt. \end{array}

Evaluating these expressions, we arrive at

\displaystyle  \begin{array}{rl} 1 \leq & \frac{1}{6} + \frac{1}{4(3+\alpha)} + \frac{1}{2(n-1)} + \frac{1}{4(n-1)(3+\alpha)} \\ & + \frac{24 a^4 n (n-1) (n-2)}{(3+\alpha) (n-4)^4 c(\alpha)^4} + \frac{a^4 n (n-1) (n-2)}{16(3+\alpha)} \left(\frac{\alpha}{3+\alpha}\right)^{\frac{n-4}{2}}. \end{array}

Since {\alpha \leq 17}, we have {\frac{\alpha}{3+\alpha} \leq \frac{17}{20}}. Next, we claim that {(3+\alpha) c(\alpha)^4 \geq 1}. As {c(\alpha)} is monotone increasing in {n}, it suffices to do this when {n=201}. Here one can directly compute that

\displaystyle  \frac{d}{d\alpha} ((3+\alpha) c(\alpha)^4) = - c(\alpha)^3 \frac{5(3+\alpha)^2 - 103(3+\alpha) + 900}{200(3+\alpha)} < 0

since the discriminant {-7391} of the numerator is negative, we conclude that

\displaystyle  (3+\alpha) c(\alpha)^4 \geq (3+17) c(17)^4 \geq 1.1529\dots > 1

as desired.

Dropping some {\alpha} and {a} terms, we conclude that

\displaystyle  \begin{array}{rl} 1 \leq & \frac{1}{6} + \frac{1}{12} + \frac{7}{12(n-1)} + \frac{24 n (n-1) (n-2)}{(n-4)^4} \\ & + \frac{n (n-1) (n-2)}{48} \left(\frac{17}{20}\right)^{\frac{n-4}{2}}. \end{array}

Every term on the right-hand side can be seen to be decreasing in {n} for {n \geq 201}. Thus the right-hand side can be bounded by

\displaystyle  \begin{array}{rl} & \frac{1}{6} + \frac{1}{12} + \frac{7}{12 \times 200} + \frac{24 \cdot 201 \cdot 200 \cdot 199}{197^4} \\ & + \frac{201 \cdot 200 \cdot 199}{48} \left(\frac{17}{20}\right)^{\frac{201-4}{2}} \leq 0.399, \end{array}

giving the desired contradiction.

The remaining range to handle is when

\displaystyle  5 \leq n \leq 200; \quad 0 \leq \alpha \leq 17; \quad c(\alpha)^2 \leq 1 - \frac{2\alpha}{n-1}.

It turns out that (22) remains infeasible in this range. This can be illustrated numerically without much difficulty: see this applet. For instance, in the most delicate case {n=53}, the right-hand side of (22) only gets as large as {0.853} (and in particular stays below {1}) throughout the range {0 \leq \alpha \leq 17}:

I have also verified this bound in Lean.

— 1. The polar inequality —

We begin with a proof of Proposition 10.

As is well known, the Möbius transform {z \mapsto \frac{a-z}{1-az}} maps the closed unit disk to itself. In particular, we have

\displaystyle  \left|\frac{1-az_j}{a-z_j}\right| \geq 1

for all of the zeroes {z_j}. Inserting this into the polar identity (8) and using the triangle inequality, we conclude the lower bound

\displaystyle  \int_0^1 \prod_{j=1}^{n-1} |t (1-a^2) q_j + a|\ dt \geq 1. \ \ \ \ \ (23)

We now convert this bound to a bound involving the quantity {x} in (12). From the arithmetic mean-geometric mean inequality we have

\displaystyle  \prod_{j=1}^{n-1} |t (1-a^2) q_j + a| \leq \left(\frac{1}{n-1} \sum_{j=1}^{n-1} |t (1-a^2) q_j + a|^2\right)^{\frac{n-1}{2}} \ \ \ \ \ (24)

and from (12) we have

\displaystyle  \begin{array}{rl} \sum_{j=1}^{n-1} |t (1-a^2) q_j + a|^2 = & (n-1) a^2 + 2 (n-1) a t (1-a^2) x \\ & + t^2 (1-a^2)^2 \sum_{j=1}^{n-1} |q_j|^2. \end{array}

Since {|q_j|^2 \leq 1}, we thus have

\displaystyle  \frac{1}{n-1} \sum_{j=1}^{n-1} |t (1-a^2) q_j + a|^2 \leq a^2 + 2 a t x (1-a^2) + t^2 (1-a^2)^2 \ \ \ \ \ (25)

giving the raw polar inequality (16).

Bounding {t^2} by {t} and using the quantities {\alpha,\beta(1)} from (11), (15), we observe that

\displaystyle  a^2 + 2 a t x (1-a^2) + t^2 (1-a^2)^2 \leq 1 + \frac{2}{n-1} \alpha (-1 + (2-\beta(1)) t).

Using the basic inequality {1 + y \leq \exp(y)}, we thus have

\displaystyle  (a^2 + 2 a t x (1-a^2) + t^2 (1-a^2)^2)^{\frac{n-1}{2}} \leq \exp ( \alpha (-1 + (2-\beta(1)) t)),

with strict inequality for {t < 1}. From (16) we conclude (17). This also implies {\beta(1) < 1}, since otherwise the integrand is always bounded by {1}, which is absurd.

On evaluating the integral in (17), we obtain

\displaystyle 1 < \frac{e^{\alpha (1 - \beta(1))}- e^{-\alpha}}{\alpha (2 - \beta(1))}

and thus

\displaystyle  e^{\alpha (1 - \beta(1))} > \alpha (2 - \beta(1)) + e^{-\alpha} \geq \alpha,

so on taking logarithms we obtain

\displaystyle  1 - \beta(1) > \frac{\log \alpha}{\alpha}.

It remains to establish the bound

\displaystyle  \beta(1) < \frac{\alpha}{3+\alpha}. \ \ \ \ \ (26)

Here we use an AI-generated argument. One can directly calculate

\displaystyle  \int_0^1 \exp ( \alpha (-1 + (2-\beta(1)) t))\ dt = e^{-u} \frac{\sinh h}{h}

where {u = \alpha \beta(1)/2} and {h = \alpha (2-\beta(1))/2}. If we can show that

\displaystyle  \log \frac{\sinh h}{h} \leq \sqrt{h^2+9} - 3 \ \ \ \ \ (27)

for all {h > 0}, then taking logarithms in (17) yields

\displaystyle  u < \log \frac{\sinh h}{h} \leq \sqrt{h^2+9} - 3

from which (26) will follow by routine algebra.

Both sides of (27) vanish at {h=0}. Taking derivatives, it suffices to show that

\displaystyle  \coth h - \frac{1}{h} \leq \frac{h}{\sqrt{h^2+9}}

which rearranges to

\displaystyle  h^4 \sinh^2 h - (h^2+9) (h \cosh h - \sinh h)^2 \geq 0.

To expand the left-hand side, we use the double angle formulae {\sinh^2 h = \frac{\cosh 2h - 1}{2}} and

\displaystyle  (h \cosh h - \sinh h)^2 = \frac{(h^2+1) \cosh 2h + (h^2-1)}{2} - h \sinh 2h

to rewrite it as

\displaystyle  \frac{h^4 (\cosh 2h - 1)}{2} - (h^2+9) \left( \frac{(h^2+1) \cosh 2h + (h^2-1)}{2} - h \sinh 2h \right).

Collecting the coefficient of {h^{2k}} for {k \geq 3} and extracting a common factor of {\frac{2^{2k-3}}{(2k)!}}, one is left with

\displaystyle  \begin{array}{rl} & N (N-1) (N-2) - 10 N (N-1) + 36 N - 36 \\ &= (N-1) (N^2 - 12N + 36) = (N-1) (N-6)^2 \end{array} \ \ \ \ \ (28)

where {N := 2k}. (The remaining coefficients, which also receive contributions from the polynomial terms, all vanish.) Thus the left-hand side has the Taylor expansion

\displaystyle  \sum_{k=4}^\infty \frac{2^{2k-3} (2k-1) (2k-6)^2}{(2k)!} h^{2k},

in which every coefficient is non-negative, giving the claim.

Remark 12 As the image in the introduction suggests, the bound (18) is only slightly weaker than (17). For small {\alpha}, one can perform Taylor approximation on the latter bound to obtain

\displaystyle  \beta(1) < \frac{\alpha}{3} - \frac{\alpha^2}{9} + \frac{19 \alpha^3}{540} - \frac{17 \alpha^4}{1620} + O(\alpha^5)

while the former bound is

\displaystyle  \beta(1) < \frac{\alpha}{3} - \frac{\alpha^2}{9} + \frac{\alpha^3}{27} - \frac{\alpha^4}{81} + O(\alpha^5).

Note that {\frac{19}{540} = 0.0351\dots} is slightly smaller than {\frac{1}{27} = 0.0370\dots}.

Relating to this, the constant {9} in (27) cannot be improved.

— 2. The origin inequality —

Now we turn to the proof of Proposition 11, which is more difficult and revolves around an analysis of the function {F} defined in (6). We begin with a heuristic analysis. Inserting the approximation {1 - \varepsilon \approx \exp(-\varepsilon)} for small {\varepsilon} into (6) and using (12), we are led to the approximation

\displaystyle  F(t) \approx \exp( - (n-1) a t (x+iy) ), \ \ \ \ \ (29)

at least when {t} is small (which turns out to be the dominant regime in applications). This suggests a relation

\displaystyle  \int_0^1 F(t)\ dt \approx \frac{1 - F(1)}{(n-1) a (x+iy)} \ \ \ \ \ (30)

between the two expressions involving {F} in the origin identities in Lemma 6. Substituting in this approximation, we obtain some (slightly complicated) approximation for the sum {\sum_{j=1}^{n-1} \frac{1}{z_j}} in terms of {a}, {n}, {x}, {y}, and the product {J := \prod_{j=1}^{n-1} z_j q_j}.

As {z_j, q_j} lie in the closed unit disk, the product {J} does also. However, past experience with the Sendov conjecture has taught us that the worst cases tend to be when {z_j, q_j} lie very close to the boundary of the disk, so that {|J|} is close to one. For instance, in Example 4 all the {z_j} and {q_j} lie on the unit circle, and {J = (-1)^{n+1}}. See Remark 3 of Dégot or Theorem 1.10(ii) of my own paper for other places where this heuristic is noted. To simplify the discussion, let us assume for now that {|J|} is exactly one, so that {z_j, q_j} all lie on the unit circle. This leads in particular to the inversion identities

\displaystyle  \frac{1}{z_j} = \overline{z_j}; \quad \frac{1}{q_j} = \overline{q_j}. \ \ \ \ \ (31)

The centroid identity in Lemma 6(i) relates the sum of the {z_j} with the sum of the {1/q_j}. Using (31), this gives a similar identity relating the sum of the {1/z_j} with the sum of the {q_j}. The latter sum is of course just {(n-1) (x+iy)}. This combines well with the previous approximation, thus giving an approximate identity relating {x}, {y} to {J}, {a}, and {n}. As it turns out, the roles of {y} and {J} are minor and can be quickly eliminated for the purposes of obtaining useful bounds, leading eventually to the relation in Proposition 11.

We turn to the details. To make the approximation (30) more precise, we note that {F(0)=1}, and hence by the fundamental theorem of calculus

\displaystyle  1 = F(1) - \int_0^1 F'(t)\ dt.

The heuristic (29) predicts that {F'(t) \approx -(n-1) a(x+iy) F(t)}, which would give (30). If we actually differentiate (6) carefully, we obtain the exact identity

\displaystyle  F'(t) = - (n-1) a(x+iy) F(t) - \sum_{j=1}^{n-1} a^2 t q_j^2 \prod_{k \neq j} (1 - atq_k).

Bounding {|q_j^2| \leq 1}, we write this

\displaystyle  F'(t) = - (n-1) a(x+iy) F(t) + O_{\leq} \left( a^2 t \sum_{j=1}^{n-1} \prod_{k \neq j} |1 - atq_k| \right).

When faced with a similar expression in (24), we used the arithmetic mean-geometric mean inequality. Here, the analogous tool is Maclaurin’s inequality, which gives

\displaystyle  \frac{1}{n-1} \sum_{j=1}^{n-1} \prod_{k \neq j} |1 - atq_k| \leq \left(\frac{1}{n-1} \sum_{j=1}^{n-1} |1 - at q_j|\right)^{n-2},

and hence by Cauchy–Schwarz

\displaystyle  \frac{1}{n-1} \sum_{j=1}^{n-1} \prod_{k \neq j} |1 - atq_k| \leq \left(\frac{1}{n-1} \sum_{j=1}^{n-1} |1 - at q_j|^2\right)^{\frac{n-2}{2}}.

Repeating the calculations used to show (25), we have

\displaystyle  \frac{1}{n-1} \sum_{j=1}^{n-1} |1 - at q_j|^2 \leq 1 - 2 a x t + a^2 t^2 = \beta(t)

and so we obtain the bound

\displaystyle  F'(t) = - (n-1) a(x+iy) F(t) + O_{\leq} \left( a^2 t (n-1) \beta(t)^{\frac{n-2}{2}} \right).

Integrating this, we obtain a rigorous analogue of (30),

\displaystyle  \begin{array}{rl} 1 - F(1) = & (n-1) a(x+iy) \int_0^1 F(t)\ dt \\ & + O_{\leq} \left( \int_0^1 a^2 t (n-1) \beta(t)^{\frac{n-2}{2}}\ dt \right) \end{array}

and thus by the triangle inequality

\displaystyle  \begin{array}{rl} 1 \leq & \left|F(1) + (n-1) a(x+iy) \int_0^1 F(t)\ dt\right| \\ & + \int_0^1 a^2 t (n-1) \beta(t)^{\frac{n-2}{2}}\ dt. \end{array} \ \ \ \ \ (32)

From the first and second origin identities (9), (10) we have

\displaystyle  \begin{array}{rl} & F(1) + (n-1) a(x+iy) \int_0^1 F(t)\ dt \\ &= \frac{(-1)^{n-1} J}{n} \left(1 + a \sum_{j=1}^{n-1} \frac{1}{z_j} + (n-1) a(x+iy) \right). \end{array} \ \ \ \ \ (33)

The next step is thus to estimate {\sum_{j=1}^{n-1} \frac{1}{z_j}}. When {|J|=1}, then all the {z_j, q_j} were on the unit circle and we could use (31) (and the centroid identity) to proceed. Now, we are no longer assuming {|J|} to equal {1}, but we can still adapt the previous arguments with a loss proportional to {1 - |J|^2}. The key lemma is

Lemma 13 (Defect lemma) Let {w_1,\dots,w_N} be some points in the closed unit disk. Then

\displaystyle  \prod_{j=1}^N |w_j| \times \sum_{j=1}^N \left|\frac{1}{w_j} - \overline{w_j}\right| \leq 1 - \prod_{j=1}^N |w_j|^2.

Proof: By a limiting argument we may assume that none of the {w_j} vanish. If we write {|w_j| = e^{-a_j}} for some {a_j \geq 0}, then we can calculate that

\displaystyle  \left|\frac{1}{w_j} - \overline{w_j}\right| = 2 \sinh a_j

and

\displaystyle  1 - \prod_{j=1}^N |w_j|^2 = \prod_{j=1}^N |w_j| \times 2 \sinh \sum_{j=1}^N a_j.

Thus the desired inequality reduces to the superadditivity property

\displaystyle  \sinh \sum_{j=1}^N a_j \geq \sum_{j=1}^N \sinh a_j.

But from the sinh addition formula {\sinh(a+b) = \sinh a \cosh b + \sinh b \cosh a} we have

\displaystyle  \sinh(a+b) \geq \sinh a + \sinh b

for all non-negative {a,b} (this also follows from the convex nature of {\sinh} together with {\sinh 0 = 0}), and the claim follows by induction. \Box

We remark that the lemma can also be proven by direct induction, without an appeal to hyperbolic trigonometry.

From taking complex conjugates of the centroid identity (7) and performing some algebra, we have

\displaystyle  J \sum_{j=1}^{n-1} \left(\frac{n-1}{n} \overline{z_j} + \frac{1}{\overline{q_j}}\right) = a J \frac{(n-1)^2}{n}.

Using the defect lemma (applied to the points {z_1,\dots,z_{n-1},\overline{q_1},\dots,\overline{q_{n-1}}}) and the triangle inequality we conclude that

\displaystyle  J \sum_{j=1}^{n-1} \left(\frac{n-1}{n} \frac{1}{z_j} + q_j\right) = a J \frac{(n-1)^2}{n} + O_{\leq}\left( 1 - |J|^2 \right),

where as before we are removing singularities when some of the {z_j} vanish. Applying (12) and some algebraic manipulation, we arrive at

\displaystyle  J \sum_{j=1}^{n-1} \frac{1}{z_j} = a J (n-1) - n J (x+iy) + O_{\leq}\left( \frac{n}{n-1} (1 - |J|^2) \right)

Substituting this back into (33), we conclude that

\displaystyle  \begin{array}{rl} & \left|F(1) + (n-1) a(x+iy) \int_0^1 F(t)\ dt\right| \\ &= \left| \frac{|J|}{n} + \frac{a^2 |J|(n-1)}{n} - a|J| (x+iy) + \frac{a (n-1) (x+iy) |J|}{n} \right| \\ & \quad + O_{\leq}\left( \frac{a}{n-1} (1-|J|^2) \right) \end{array}

and hence after some algebra and the triangle inequality

\displaystyle  \begin{array}{rl} & \left|F(1) + (n-1) a(x+iy) \int_0^1 F(t)\ dt\right| \\ &\leq |J| \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| + \frac{a}{n-1} (1-|J|^2). \end{array}

Inserting this into (32), we obtain

\displaystyle  \begin{array}{rl} 1 \leq & |J| \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| + \frac{a}{n-1} (1-|J|^2) \\ & + \int_0^1 a^2 t (n-1) \beta(t)^{\frac{n-2}{2}}\ dt. \end{array} \ \ \ \ \ (34)

We can simplify (34) by reducing to the {|J|=1} case. Indeed, we shall show that

\displaystyle  \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| \geq 2 \frac{a}{n-1} \ \ \ \ \ (35)

which implies that the right-hand side of (34) is non-decreasing in {|J|} in the range {0 \leq |J| \leq 1}. Thus we may replace {|J|} by {1} in (34) to conclude that

\displaystyle  1 \leq \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| + \int_0^1 a^2 t (n-1) \beta(t)^{\frac{n-2}{2}}\ dt. \ \ \ \ \ (36)

Let us now verify (35). Using {|x+iy| \leq 1} and the triangle inequality, we can lower bound

\displaystyle  \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| \geq \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a}{n}.

Inserting this into (35) and clearing denominators, we reduce after some algebra to

\displaystyle  (n-1)^2 a^2 - (3n-1) a + n-1 > 0.

But as a quadratic polynomial in {a}, the left-hand side has discriminant {(3n-1)^2 - 4(n-1)^3}, which one can check to be negative for sufficiently large {n} (in fact {n \geq 5} suffices), giving the claim (35).

Next we eliminate the role of the imaginary term {iy}. Observe for any complex number {s+it} with positive real part that

\displaystyle  |s+it| \leq s + \frac{t^2}{2s}

as can be seen by squaring both sides. The expression {\frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n}} has real part

\displaystyle  \frac{a^2 (n-1)}{n} + \frac{1-ax}{n}

which lies between {\frac{a^2 (n-1)}{n}} and {1} (in particular, it is positive), and imaginary part of magnitude at most

\displaystyle  \frac{a}{n} (1 - x^2)^{1/2}

by (13). We conclude that

\displaystyle  \left| \frac{a^2 (n-1)}{n} + \frac{1}{n} - \frac{a(x+iy)}{n} \right| \leq \frac{a^2 (n-1)}{n} + \frac{1-ax}{n} + \frac{1-x^2}{2n(n-1)}.

The right-hand side can be rearranged using the quantity {\alpha} from (11) as

\displaystyle  1 - \frac{2\alpha}{n} - \frac{ax}{n} + \frac{1-x^2}{2n(n-1)},

so the bound (36) gives (19).

— 2.1. Upper bound on {\alpha}

Now we can prove (20). Suppose for contradiction that {\alpha > 17}; since {\alpha \leq \frac{n-1}{2}}, this implies that {n \geq 36}. Crudely discarding the {ax} term in (19) and bounding {\frac{1-x^2}{2(n-1)}} by {\frac{1}{2(n-1)}}, we have

\displaystyle  2 \alpha \leq \frac{1}{2(n-1)} + a^2 n(n-1) \int_0^1 t \beta(t)^{\frac{n-2}{2}}\ dt.

The quadratic polynomial {\beta(t)} equals {1} at {t=0} and attains its minimum at {t = x/a} with value {1-x^2}. By convexity, we thus have

\displaystyle  \beta(t) \leq 1 - axt \leq \exp(-axt)

for {0 \leq t \leq x/a} and

\displaystyle  \beta(t) \leq \beta(1)

for {x/a \leq t \leq 1} (this latter statement is vacuous if {x/a > 1}). Since {\int_0^1 t\ dt = \frac{1}{2}}, we can therefore crudely bound

\displaystyle  \int_0^1 t (1 - 2 a x t + a^2 t^2)^{\frac{n-2}{2}}\ dt \leq \int_0^\infty t \exp\left( - \frac{n-2}{2} axt\right)\ dt + \frac{1}{2} \beta(1)^{\frac{n-2}{2}}

\displaystyle  = \frac{4}{(n-2)^2 a^2 x^2} + \frac{1}{2} \beta(1)^{\frac{n-2}{2}},

and hence

\displaystyle  2 \alpha \leq \frac{1}{2(n-1)} + \frac{4n(n-1)}{(n-2)^2 x^2} + \frac{a^2 n(n-1)}{2} \beta(1)^{\frac{n-2}{2}}.

From (15) we have {x^2 \geq 1-\beta(1)}, thus by (18) one has

\displaystyle  \frac{1}{x^2} \leq \frac{\alpha}{\log \alpha}.

From another application of (18) one has

\displaystyle  \beta(1) \leq \exp( - (1 - \beta(1)) ) \leq \alpha^{-1/\alpha}.

We conclude that

\displaystyle  2 \alpha \leq \frac{1}{2(n-1)} + \frac{4n(n-1)}{(n-2)^2 \log \alpha} \alpha + \frac{a^2 n(n-1)}{2} \alpha^{-\frac{n-2}{2\alpha}}.

It is now convenient to introduce the quantity {u := \frac{a^2}{1-a^2}}, thus {0 < u < \infty} with

\displaystyle  a^2 = \frac{u}{1+u}

and

\displaystyle  \frac{n-2}{2 \alpha} = 1 + u - \frac{1}{2\alpha}.

Inserting these bounds and dividing by {\alpha}, we conclude

\displaystyle  2 \leq \frac{1}{2(n-1) \alpha} + \frac{4n(n-1)}{(n-2)^2 \log \alpha} + \frac{u}{2(1+u) \alpha^2} n(n-1) \alpha^{-u} \alpha^{\frac{1}{2\alpha}}.

Since {\alpha = \frac{n-1}{2(1+u)}}, we obtain

\displaystyle  2 \leq \frac{1}{2(n-1) \alpha} + \frac{4n(n-1)}{(n-2)^2 \log \alpha} + \frac{2n}{n-1} u (1+u) \alpha^{-u} \alpha^{\frac{1}{2\alpha}}.

Since {n \geq 36} and {\alpha \geq 17}, we conclude that

\displaystyle  2 \leq \frac{1}{2 \times 35 \times 17} + \frac{4 \times 36 \times 35}{(34)^2 \log 17} + \frac{2 \times 36}{35} u (1+u) 17^{-u} 17^{\frac{1}{2 \times 17}}.

Routine calculus shows that {u(1+u) 17^{-u}} has a maximum of at most {0.1825}, and that the right-hand side here is at most {1.948}, giving the required contradiction. This proves (20).

— 2.2. A simplified estimate —

Now we show (21). Note from (11) that

\displaystyle  1-a^2 = \frac{2\alpha}{n-1} \ \ \ \ \ (37)

while from (15) we have

\displaystyle  1 - x^2 \leq \beta(1) \ \ \ \ \ (38)

and hence also

\displaystyle  \frac{1}{x^2} \leq 1 + \frac{\beta(1)}{1-\beta(1)}. \ \ \ \ \ (39)

From (15) we have

\displaystyle  \beta(t) = (1 - axt)^2 + a^2 t^2 (1-x^2).

By the mean value theorem (noting that {1-axt} is non-negative) we thus have

\displaystyle  \beta(t)^{\frac{n-2}{2}} \leq \left((1 - axt)^2\right)^{\frac{n-2}{2}} + a^2 t^2 (1-x^2) \frac{n-2}{2} \beta(t)^{\frac{n-4}{2}}.

From the standard beta function identity

\displaystyle  \int_0^{1/ax} t (1 - axt)^{n-2}\ dt = \frac{1}{a^2 x^2 n(n-1)}

(and the fact that {1/ax \geq 1}) we can thus replace (19) by

\displaystyle  \begin{array}{rl} 2\alpha + ax \leq & \frac{1-x^2}{2(n-1)} + \frac{1}{x^2} \\ & + \frac{a^4 n (n-1) (n-2) (1-x^2)}{2} \int_0^1 t^3 \beta(t)^{\frac{n-4}{2}}\ dt. \end{array}

From (15) we have

\displaystyle ax = 1 - \frac{\beta(1)}{2} - \frac{1-a^2}{2}.

Thus by (37), (38), (39)

\displaystyle  \begin{array}{rl} 2\alpha \leq & \frac{\beta(1)}{1-\beta(1)} + \frac{\beta(1)}{2} + \frac{\alpha}{n-1} + \frac{\beta(1)}{2(n-1)} \\ & + \frac{a^4 n (n-1) (n-2) \beta(1)}{2} \int_0^1 t^3 \beta(t)^{\frac{n-4}{2}}\ dt. \end{array}

Dividing by the positive quantity {2\alpha} gives the claim.

— 3. Rubinstein’s theorem —

We now adapt the arguments to give a proof of Rubinstein’s theorem that the Phelps–Rodriguez conjecture holds in the {a=1} case, i.e.,

Theorem 14 (Rubinstein’s theorem) Let {n \geq 2}, and let {p : {\bf C} \rightarrow {\bf C}} be a degree {n} polynomial with all zeroes in the unit disk. If {p(1)=0}, then there exists a critical point {\zeta} of {p} with {|\zeta-1| < 1}, unless {p} is a scalar multiple of {z^n - 1}.

Taking contrapositives, we may assume that the critical points {\zeta_j} are of the form {1 - 1/q_j} for some {q_j} in the closed unit disk, and normalize {p} to be monic; our task is to show that {p(z) = z^n-1}.

The polar identity (8), based on calculating {p(1/a) / p'(a)} degenerates to a triviality when {a=1}, but we have the following usable substitute, valid for any choice of {a}, first observed in equation (3.2) of Meir–Sharma:

Lemma 15 (Meir–Sharma identity) If {p(a)=0} and the critical points are of the form {a - 1/q_j} then all the zeroes {z_j} are not equal to {a}, and

\displaystyle  \sum_{j=1}^{n-1} q_j = 2 \sum_{j=1}^{n-1} \frac{1}{a-z_j}.

Proof: By hypothesis, {a} is not a critical point of {p}, so {p'(a) \neq 0} and {z_j \neq a} for all {j}. Instead of computing {p(1/a)/p'(a)}, we instead consider the expression {p''(a) / p'(a)}. On the one hand, from (4) we have

\displaystyle  p'(a) = n \prod_{j=1}^{n-1} \frac{1}{q_j}

while from differentiating (4) we have

\displaystyle  p''(a) = n \prod_{j=1}^{n-1} \frac{1}{q_j} \times \sum_{j=1}^{n-1} q_j.

Meanwhile, from (3) we have

\displaystyle  p'(a) = \prod_{j=1}^{n-1} (a-z_j)

and from differentiating (3) we have

\displaystyle  p''(a) = \left(\prod_{j=1}^{n-1} (a-z_j)\right) \times \left( \sum_{j=1}^{n-1} \frac{1}{a-z_j} + \sum_{j=1}^{n-1} \frac{1}{a-z_j} \right).

Using these identities to compute {p''(a) / p'(a)} in two different ways gives the claim. \Box

Now take {a=1}. Since {z_j} lie in the closed unit disk, {\frac{1}{1-z_j}} has real part at least {1/2}, while {\mathrm{Re} q_j} is at most {1}. Thus, the only way that the above identity can hold is if {\mathrm{Re} q_j = 1} for all {j}, hence {q_j = 1} for all {j}. Thus all critical points are at the origin, which forces {p(z) = z^n - c} for some {c}. Since {p(1)=0}, we conclude that {c=1}, giving the claim.

— 4. The {n \leq 4} cases —

We now prove the {n \leq 4} cases of Conjecture 3. The starting point is (23). Using the triangle inequality and {|q_j| \leq 1}, this implies that

\displaystyle  1 \leq \int_0^1 (a+(1-a^2) t)^{n-1}\ dt.

(This also follows from (16) and {x \leq 1}.) From Hölder’s inequality and {n \leq 4} we conclude that

\displaystyle  1 \leq \int_0^1 (a+(1-a^2) t)^3\ dt.

The right-hand side can be computed to equal

\displaystyle  1 - \frac{(a-1)^2}{4} (a^2 (1-a)^2 + 3(1-a^2) + 2a),

which is obviously less than {1} for {0 \leq a < 1}, giving the contradiction.

Remark 16 The same argument also works for {n=5}, but breaks down for higher {n}.

— 5. Further directions —

The Sendov and Phelps–Rodriguez conjectures are now resolved, but several related conjectures remain open. The following strengthening of Sendov’s conjecture, by Borcea, is open for any {q < \infty}:

Conjecture 17 (Borcea conjecture) Let {1 \leq q < \infty} and {n \geq 2}, and let {p} be a degree {n} polynomial with zeroes {z_1,\dots,z_n} satisfying {\frac{1}{n} \sum_{j=1}^n |z_j|^q \leq 1}. Then for every zero {a} of {p}, there exists a critical point {\zeta} of {p} with {|\zeta - a| \leq 1}.

Sendov’s conjecture is the limiting case {q=\infty} of this conjecture. There has been relatively little progress on this conjecture: the cases {(q,n) = (1,3), (2,4)} were established by Khavinson, Pereira, Putinar, Saff, and Shimorin, and in this previous paper we reported the negative result that AlphaEvolve failed to find a counterexample to the conjecture. The proof methods here do not seem to extend easily; all the identities relating zeroes and critical points continue to hold, but now that the {z_j} are only constrained to the unit disk in an averaged moment sense, all of the inequalities developed above now fail.

Another strengthening of Sendov’s conjecture that remains open is Schmeisser’s conjecture:

Conjecture 18 (Schmeisser’s conjecture) Let {n \geq 2}, and let {p} be a degree {n} polynomial with all zeroes in the closed unit disk. Then for any {a} in the convex hull of the zeroes of {p}, there exists a critical point {\zeta} of {p} with {|\zeta - a| \leq 1}.

Schmeisser proved several special cases of this conjecture, and AlphaEvolve again failed to find a counterexample, but there has not been much further progress. Here, the {z_j} are now back in the closed unit disk, but we no longer have {p(a)=0}, again rendering most of the previous identities invalid. But perhaps some modification of the arguments here can make some progress on this conjecture.

A common generalization of the Borcea and Schmeisser conjectures was proposed in Conjecture 2.4 of this paper of Zhang.

Another well known variant of Sendov’s conjecture is Smale’s problem:

Conjecture 19 (Smale’s problem) Let {n \geq 2}, and let {p} be a degree {n} polynomial. Then for any zero {a} of {p}, there exists a critical point {\zeta} of {p} with {|p(\zeta)| \leq (1-\frac{1}{n}) |\zeta-a| |p'(a)|}.

The constant {1-\frac{1}{n}} is best possible, as can be seen by the example {p(z) = z^n-z} and {a=0}. Using the Koebe one-quarter theorem, Smale proved this conjecture with {1 - \frac{1}{n}} replaced by {4}. Some slight improvements of this bound have been obtained over the years; for instance for {n \geq 8}, the improved bound of {4 - \frac{2.263}{\sqrt{n}}} was obtained by Crane. Again, AlphaEvolve failed to find a counterexample to this conjecture. This problem does not seem to have a direct relationship with Sendov’s conjecture, and there is no useful normalization of the zeroes and critical points that is confined to the unit disk. Nevertheless there may be some hope of making progress on this conjecture, perhaps working first in the asymptotic regime {n \rightarrow \infty}.

Needless to say, I did try some desultory attempts to use AI tools to attack these questions, but without much notable success.

One potential way forward is to find further proofs of Sendov’s conjecture that utilize other techniques that might be more broadly applicable to this larger family of problems. The proof here is remarkable in that the zeroes and critical points are treated almost as independent mathematical objects, communicating with each other only very narrowly through four identities in which one only inspects the underlying polynomial (and its derivative) at a small number of points. It could be that an approach focusing on more global features of the polynomial may lead to new proofs of Sendov’s conjecture, and perhaps also of its generalizations.