You are currently browsing the monthly archive for January 2018.
This is the first official “research” thread of the Polymath15 project to upper bound the de Bruijn-Newman constant . Discussion of the project of a non-research nature can continue for now in the existing proposal thread. Progress will be summarised at this Polymath wiki page.
The proposal naturally splits into at least three separate (but loosely related) topics:
- Numerical computation of the entire functions , with the ultimate aim of establishing zero-free regions of the form for various .
- Improved understanding of the dynamics of the zeroes of .
- Establishing the zero-free nature of when and is sufficiently large depending on and .
Below the fold, I will present each of these topics in turn, to initiate further discussion in each of them. (I thought about splitting this post into three to have three separate discussions, but given the current volume of comments, I think we should be able to manage for now having all the comments in a single post. If this changes then of course we can split up some of the discussion later.)
To begin with, let me present some formulae for computing (inspired by similar computations in the Ki-Kim-Lee paper) which may be useful. The initial definition of is
where
is a variant of the Jacobi theta function. We observe that in fact extends analytically to the strip
as has positive real part on this strip. One can use the Poisson summation formula to verify that is even, (see this previous post for details). This lets us obtain a number of other formulae for . Most obviously, one can unfold the integral to obtain
In my previous paper with Brad, we used this representation, combined with Fubini’s theorem to swap the sum and integral, to obtain a useful series representation for in the case. Unfortunately this is not possible in the case because expressions such as diverge as approaches . Nevertheless we can still perform the following contour integration manipulation. Let be fixed. The function decays super-exponentially fast (much faster than , in particular) as with ; as is even, we also have this decay as with (this is despite each of the summands in having much slower decay in this direction – there is considerable cancellation!). Hence by the Cauchy integral formula we have
Splitting the horizontal line from to at and using the even nature of , we thus have
Using the functional equation , we thus have the representation
where
where is the oscillatory integral
The formula (2) is valid for any . Naively one would think that it would be simplest to take ; however, when and is large (with bounded), it seems asymptotically better to take closer to , in particular something like seems to be a reasonably good choice. This is because the integrand in (3) becomes significantly less oscillatory and also much lower in amplitude; the term in (3) now generates a factor roughly comparable to (which, as we will see below, is the main term in the decay asymptotics for ), while the term still exhibits a reasonable amount of decay as . We will use the representation (2) in the asymptotic analysis of below, but it may also be a useful representation to use for numerical purposes.
Building on the interest expressed in the comments to this previous post, I am now formally proposing to initiate a “Polymath project” on the topic of obtaining new upper bounds on the de Bruijn-Newman constant . The purpose of this post is to describe the proposal and discuss the scope and parameters of the project.
De Bruijn introduced a family of entire functions for each real number , defined by the formula
where is the super-exponentially decaying function
As discussed in this previous post, the Riemann hypothesis is equivalent to the assertion that all the zeroes of are real.
De Bruijn and Newman showed that there existed a real constant – the de Bruijn-Newman constant – such that has all zeroes real whenever , and at least one non-real zero when . In particular, the Riemann hypothesis is equivalent to the upper bound . In the opposite direction, several lower bounds on have been obtained over the years, most recently in my paper with Brad Rodgers where we showed that , a conjecture of Newman.
As for upper bounds, de Bruijn showed back in 1950 that . The only progress since then has been the work of Ki, Kim and Lee in 2009, who improved this slightly to . The primary proposed aim of this Polymath project is to obtain further explicit improvements to the upper bound of . Of course, if we could lower the upper bound all the way to zero, this would solve the Riemann hypothesis, but I do not view this as a realistic outcome of this project; rather, the upper bounds that one could plausibly obtain by known methods and numerics would be comparable in achievement to the various numerical verifications of the Riemann hypothesis that exist in the literature (e.g., that the first non-trivial zeroes of the zeta function lie on the critical line, for various large explicit values of ).
In addition to the primary goal, one could envisage some related secondary goals of the project, such as a better understanding (both analytic and numerical) of the functions (or of similar functions), and of the dynamics of the zeroes of these functions. Perhaps further potential goals could emerge in the discussion to this post.
I think there is a plausible plan of attack on this project that proceeds as follows. Firstly, there are results going back to the original work of de Bruijn that demonstrate that the zeroes of become attracted to the real line as increases; in particular, if one defines to be the supremum of the imaginary parts of all the zeroes of , then it is known that this quantity obeys the differential inequality
whenever is positive; furthermore, once for some , then for all . I hope to explain this in a future post (it is basically due to the attraction that a zero off the real axis has to its complex conjugate). As a corollary of this inequality, we have the upper bound
for any real number . For instance, because all the non-trivial zeroes of the Riemann zeta function lie in the critical strip , one has , which when inserted into (2) gives . The inequality (1) also gives for all . If we could find some explicit between and where we can improve this upper bound on by an explicit constant, this would lead to a new upper bound on .
Secondly, the work of Ki, Kim and Lee (based on an analysis of the various terms appearing in the expression for ) shows that for any positive , all but finitely many of the zeroes of are real (in contrast with the situation, where it is still an open question as to whether the proportion of non-trivial zeroes of the zeta function on the critical line is asymptotically equal to ). As a key step in this analysis, Ki, Kim, and Lee show that for any and , there exists a such that all the zeroes of with real part at least , have imaginary part at most . Ki, Kim and Lee do not explicitly compute how depends on and , but it looks like this bound could be made effective.
If so, this suggests a possible strategy to get a new upper bound on :
- Select a good choice of parameters .
- By refining the Ki-Kim-Lee analysis, find an explicit such that all zeroes of with real part at least have imaginary part at most .
- By a numerical computation (e.g. using the argument principle), also verify that zeroes of with real part between and have imaginary part at most .
- Combining these facts, we obtain that ; hopefully, one can insert this into (2) and get a new upper bound for .
Of course, there may also be alternate strategies to upper bound , and I would imagine this would also be a legitimate topic of discussion for this project.
One appealing thing about the above strategy for the purposes of a polymath project is that it naturally splits the project into several interacting but reasonably independent parts: an analytic part in which one tries to refine the Ki-Kim-Lee analysis (based on explicitly upper and lower bounding various terms in a certain series expansion for – I may detail this later in a subsequent post); a numerical part in which one controls the zeroes of in a certain finite range; and perhaps also a dynamical part where one sees if there is any way to improve the inequality (2). For instance, the numerical “team” might, over time, be able to produce zero-free regions for with an increasingly large value of , while in parallel the analytic “team” might produce increasingly smaller values of beyond which they can control zeroes, and eventually the two bounds would meet up and we obtain a new bound on . This factoring of the problem into smaller parts was also a feature of the successful Polymath8 project on bounded gaps between primes.
The project also resembles Polymath8 in another aspect: that there is an obvious way to numerically measure progress, by seeing how the upper bound for decreases over time (and presumably there will also be another metric of progress regarding how well we can control in terms of and ). However, in Polymath8 the final measure of progress (the upper bound on gaps between primes) was a natural number, and thus could not decrease indefinitely. Here, the bound will be a real number, and there is a possibility that one may end up having an infinite descent in which progress slows down over time, with refinements to increasingly less significant digits of the bound as the project progresses. Because of this, I think it makes sense to follow recent Polymath projects and place an expiration date for the project, for instance one year after the launch date, in which we will agree to end the project and (if the project was successful enough) write up the results, unless there is consensus at that time to extend the project. (In retrospect, we should probably have imposed similar sunset dates on older Polymath projects, some of which have now been inactive for years, but that is perhaps a discussion for another time.)
Some Polymath projects have been known for a breakneck pace, making it hard for some participants to keep up. It’s hard to control these things, but I am envisaging a relatively leisurely project here, perhaps taking the full year mentioned above. It may well be that as the project matures we will largely be waiting for the results of lengthy numerical calculations to come in, for instance. Of course, as with previous projects, we would maintain some wiki pages (and possibly some other resources, such as a code repository) to keep track of progress and also to summarise what we have learned so far. For instance, as was done with some previous Polymath projects, we could begin with some “online reading seminars” where we go through some relevant piece of literature (most obviously the Ki-Kim-Lee paper, but there may be other resources that become relevant, e.g. one could imagine the literature on numerical verification of RH to be of value).
One could also imagine some incidental outcomes of this project, such as a more efficient way to numerically establish zero free regions for various analytic functions of interest; in particular, the project may well end up focusing on some other aspect of mathematics than the specific questions posed here.
Anyway, I would be interested to hear in the comments below from others who might be interested in participating, or at least observing, this project, particularly if they have suggestions regarding the scope and direction of the project, and on organisational structure (e.g. if one should start with reading seminars, or some initial numerical exploration of the functions , etc..) One could also begin some preliminary discussion of the actual mathematics of the project itself, though (in line with the leisurely pace I was hoping for), I expect that the main burst of mathematical activity would happen later, once the project is formally launched (with wiki page resources, blog posts dedicated to specific aspects of the project, etc.).
In this post we assume the Riemann hypothesis and the simplicity of zeroes, thus the zeroes of in the critical strip take the form for some real number ordinates . From the Riemann-von Mangoldt formula, one has the asymptotic
as ; in particular, the spacing should behave like on the average. However, it can happen that some gaps are unusually small compared to other nearby gaps. For the sake of concreteness, let us define a Lehmer pair to be a pair of adjacent ordinates such that
The specific value of constant is not particularly important here; anything larger than would suffice. An example of such a pair would be the classical pair
discovered by Lehmer. It follows easily from the main results of Csordas, Smith, and Varga that if an infinite number of Lehmer pairs (in the above sense) existed, then the de Bruijn-Newman constant is non-negative. This implication is now redundant in view of the unconditional results of this recent paper of Rodgers and myself; however, the question of whether an infinite number of Lehmer pairs exist remain open.
In this post, I sketch an argument that Brad and I came up with (as initially suggested by Odlyzko) the GUE hypothesis implies the existence of infinitely many Lehmer pairs. We argue probabilistically: pick a sufficiently large number , pick at random from to (so that the average gap size is close to ), and prove that the Lehmer pair condition (1) occurs with positive probability.
Introduce the renormalised ordinates for , and let be a small absolute constant (independent of ). It will then suffice to show that
(say) with probability , since the contribution of those outside of can be absorbed by the factor with probability .
As one consequence of the GUE hypothesis, we have with probability . Thus, if , then has density . Applying the Hardy-Littlewood maximal inequality, we see that with probability , we have
which implies in particular that
for all . This implies in particular that
and so it will suffice to show that
(say) with probability .
By the GUE hypothesis (and the fact that is independent of ), it suffices to show that a Dyson sine process , normalised so that is the first positive point in the process, obeys the inequality
with probability . However, if we let be a moderately large constant (and assume small depending on ), one can show using -point correlation functions for the Dyson sine process (and the fact that the Dyson kernel equals to second order at the origin) that
for any natural number , where denotes the number of elements of the process in . For instance, the expression can be written in terms of the three-point correlation function as
which can easily be estimated to be (since in this region), and similarly for the other estimates claimed above.
Since for natural numbers , the quantity is only positive when , we see from the first three estimates that the event that occurs with probability . In particular, by Markov’s inequality we have the conditional probabilities
and thus, if is large enough, and small enough, it will be true with probability that
and
and simultaneously that
for all natural numbers . This implies in particular that
and
for all , which gives (2) for small enough.
Remark 1 The above argument needed the GUE hypothesis for correlations up to fourth order (in order to establish (3)). It might be possible to reduce the number of correlations needed, but I do not see how to obtain the claim just using pair correlations only.
Brad Rodgers and I have uploaded to the arXiv our paper “The De Bruijn-Newman constant is non-negative“. This paper affirms a conjecture of Newman regarding to the extent to which the Riemann hypothesis, if true, is only “barely so”. To describe the conjecture, let us begin with the Riemann xi function
where is the Gamma function and is the Riemann zeta function. Initially, this function is only defined for , but, as was already known to Riemann, we can manipulate it into a form that extends to the entire complex plane as follows. Firstly, in view of the standard identity , we can write
and hence
By a rescaling, one may write
and similarly
and thus (after applying Fubini’s theorem)
We’ll make the change of variables to obtain
If we introduce the mild renormalisation
of , we then conclude (at least for ) that
which one can verify to be rapidly decreasing both as and as , with the decrease as faster than any exponential. In particular extends holomorphically to the upper half plane.
If we normalize the Fourier transform of a (Schwartz) function as , it is well known that the Gaussian is its own Fourier transform. The creation operator interacts with the Fourier transform by the identity
Since , this implies that the function
is its own Fourier transform. (One can view the polynomial as a renormalised version of the fourth Hermite polynomial.) Taking a suitable linear combination of this with , we conclude that
is also its own Fourier transform. Rescaling by and then multiplying by , we conclude that the Fourier transform of
is
and hence by the Poisson summation formula (using symmetry and vanishing at to unfold the summation in (2) to the integers rather than the natural numbers) we obtain the functional equation
which implies that and are even functions (in particular, now extends to an entire function). From this symmetry we can also rewrite (1) as
which now gives a convergent expression for the entire function for all complex . As is even and real-valued on , is even and also obeys the functional equation , which is equivalent to the usual functional equation for the Riemann zeta function. The Riemann hypothesis is equivalent to the claim that all the zeroes of are real.
De Bruijn introduced the family of deformations of , defined for all and by the formula
From a PDE perspective, one can view as the evolution of under the backwards heat equation . As with , the are all even entire functions that obey the functional equation , and one can ask an analogue of the Riemann hypothesis for each such , namely whether all the zeroes of are real. De Bruijn showed that these hypotheses were monotone in : if had all real zeroes for some , then would also have all zeroes real for any . Newman later sharpened this claim by showing the existence of a finite number , now known as the de Bruijn-Newman constant, with the property that had all zeroes real if and only if . Thus, the Riemann hypothesis is equivalent to the inequality . Newman then conjectured the complementary bound ; in his words, this conjecture asserted that if the Riemann hypothesis is true, then it is only “barely so”, in that the reality of all the zeroes is destroyed by applying heat flow for even an arbitrarily small amount of time. Over time, a significant amount of evidence was established in favour of this conjecture; most recently, in 2011, Saouter, Gourdon, and Demichel showed that .
In this paper we finish off the proof of Newman’s conjecture, that is we show that . The proof is by contradiction, assuming that (which among other things, implies the truth of the Riemann hypothesis), and using the properties of backwards heat evolution to reach a contradiction.
Very roughly, the argument proceeds as follows. As observed by Csordas, Smith, and Varga (and also discussed in this previous blog post, the backwards heat evolution of the introduces a nice ODE dynamics on the zeroes of , namely that they solve the ODE
for all (one has to interpret the sum in a principal value sense as it is not absolutely convergent, but let us ignore this technicality for the current discussion). Intuitively, this ODE is asserting that the zeroes repel each other, somewhat like positively charged particles (but note that the dynamics is first-order, as opposed to the second-order laws of Newtonian mechanics). Formally, a steady state (or equilibrium) of this dynamics is reached when the are arranged in an arithmetic progression. (Note for instance that for any positive , the functions obey the same backwards heat equation as , and their zeroes are on a fixed arithmetic progression .) The strategy is to then show that the dynamics from time to time creates a convergence to local equilibrium, in which the zeroes locally resemble an arithmetic progression at time . This will be in contradiction with known results on pair correlation of zeroes (or on related statistics, such as the fluctuations on gaps between zeroes), such as the results of Montgomery (actually for technical reasons it is slightly more convenient for us to use related results of Conrey, Ghosh, Goldston, Gonek, and Heath-Brown). Another way of thinking about this is that even very slight deviations from local equilibrium (such as a small number of gaps that are slightly smaller than the average spacing) will almost immediately lead to zeroes colliding with each other and leaving the real line as one evolves backwards in time (i.e., under the forward heat flow). This is a refinement of the strategy used in previous lower bounds on , in which “Lehmer pairs” (pairs of zeroes of the zeta function that were unusually close to each other) were used to limit the extent to which the evolution continued backwards in time while keeping all zeroes real.
How does one obtain this convergence to local equilibrium? We proceed by broad analogy with the “local relaxation flow” method of Erdos, Schlein, and Yau in random matrix theory, in which one combines some initial control on zeroes (which, in the case of the Erdos-Schlein-Yau method, is referred to with terms such as “local semicircular law”) with convexity properties of a relevant Hamiltonian that can be used to force the zeroes towards equilibrium.
We first discuss the initial control on zeroes. For , we have the classical Riemann-von Mangoldt formula, which asserts that the number of zeroes in the interval is as . (We have a factor of here instead of the more familiar due to the way is normalised.) This implies for instance that for a fixed , the number of zeroes in the interval is . Actually, because we get to assume the Riemann hypothesis, we can sharpen this to , a result of Littlewood (see this previous blog post for a proof). Ideally, we would like to obtain similar control for the other , , as well. Unfortunately we were only able to obtain the weaker claims that the number of zeroes of in is , and that the number of zeroes in is , that is to say we only get good control on the distribution of zeroes at scales rather than at scales . Ultimately this is because we were only able to get control (and in particular, lower bounds) on with high precision when (whereas has good estimates as soon as is larger than (say) ). This control is obtained by the expressing in terms of some contour integrals and using the method of steepest descent (actually it is slightly simpler to rely instead on the Stirling approximation for the Gamma function, which can be proven in turn by steepest descent methods). Fortunately, it turns out that this weaker control is still (barely) enough for the rest of our argument to go through.
Once one has the initial control on zeroes, we now need to force convergence to local equilibrium by exploiting convexity of a Hamiltonian. Here, the relevant Hamiltonian is
ignoring for now the rather important technical issue that this sum is not actually absolutely convergent. (Because of this, we will need to truncate and renormalise the Hamiltonian in a number of ways which we will not detail here.) The ODE (3) is formally the gradient flow for this Hamiltonian. Furthermore, this Hamiltonian is a convex function of the (because is a convex function on ). We therefore expect the Hamiltonian to be a decreasing function of time, and that the derivative should be an increasing function of time. As time passes, the derivative of the Hamiltonian would then be expected to converge to zero, which should imply convergence to local equilibrium.
Formally, the derivative of the above Hamiltonian is
Again, there is the important technical issue that this quantity is infinite; but it turns out that if we renormalise the Hamiltonian appropriately, then the energy will also become suitably renormalised, and in particular will vanish when the are arranged in an arithmetic progression, and be positive otherwise. One can also formally calculate the derivative of to be a somewhat complicated but manifestly non-negative quantity (a sum of squares); see this previous blog post for analogous computations in the case of heat flow on polynomials. After flowing from time to time , and using some crude initial bounds on and in this region (coming from the Riemann-von Mangoldt type formulae mentioned above and some further manipulations), we can eventually show that the (renormalisation of the) energy at time zero is small, which forces the to locally resemble an arithmetic progression, which gives the required convergence to local equilibrium.
There are a number of technicalities involved in making the above sketch of argument rigorous (for instance, justifying interchanges of derivatives and infinite sums turns out to be a little bit delicate). I will highlight here one particular technical point. One of the ways in which we make expressions such as the energy finite is to truncate the indices to an interval to create a truncated energy . In typical situations, we would then expect to be decreasing, which will greatly help in bounding (in particular it would allow one to control by time-averaged quantities such as , which can in turn be controlled using variants of (4)). However, there are boundary effects at both ends of that could in principle add a large amount of energy into , which is bad news as it could conceivably make undesirably large even if integrated energies such as remain adequately controlled. As it turns out, such boundary effects are negligible as long as there is a large gap between adjacent zeroes at boundary of – it is only narrow gaps that can rapidly transmit energy across the boundary of . Now, narrow gaps can certainly exist (indeed, the GUE hypothesis predicts these happen a positive fraction of the time); but the pigeonhole principle (together with the Riemann-von Mangoldt formula) can allow us to pick the endpoints of the interval so that no narrow gaps appear at the boundary of for any given time . However, there was a technical problem: this argument did not allow one to find a single interval that avoided gaps for all times simultaneously – the pigeonhole principle could produce a different interval for each time ! Since the number of times was uncountable, this was a serious issue. (In physical terms, the problem was that there might be very fast “longitudinal waves” in the dynamics that, at each time, cause some gaps between zeroes to be highly compressed, but the specific gap that was narrow changed very rapidly with time. Such waves could, in principle, import a huge amount of energy into by time .) To resolve this, we borrowed a PDE trick of Bourgain’s, in which the pigeonhole principle was coupled with local conservation laws. More specifically, we use the phenomenon that very narrow gaps take a nontrivial amount of time to expand back to a reasonable size (this can be seen by comparing the evolution of this gap with solutions of the scalar ODE , which represents the fastest at which a gap such as can expand). Thus, if a gap is reasonably large at some time , it will also stay reasonably large at slightly earlier times for some moderately small . This lets one locate an interval that has manageable boundary effects during the times in , so in particular is basically non-increasing in this time interval. Unfortunately, this interval is a little bit too short to cover all of ; however it turns out that one can iterate the above construction and find a nested sequence of intervals , with each non-increasing in a different time interval , and with all of the time intervals covering . This turns out to be enough (together with the obvious fact that is monotone in ) to still control for some reasonably sized interval , as required for the rest of the arguments.
ADDED LATER: the following analogy (involving functions with just two zeroes, rather than an infinite number of zeroes) may help clarify the relation between this result and the Riemann hypothesis (and in particular why this result does not make the Riemann hypothesis any easier to prove, in fact it confirms the delicate nature of that hypothesis). Suppose one had a quadratic polynomial of the form , where was an unknown real constant. Suppose that one was for some reason interested in the analogue of the “Riemann hypothesis” for , namely that all the zeroes of are real. A priori, there are three scenarios:
- (Riemann hypothesis false) , and has zeroes off the real axis.
- (Riemann hypothesis true, but barely so) , and both zeroes of are on the real axis; however, any slight perturbation of in the positive direction would move zeroes off the real axis.
- (Riemann hypothesis true, with room to spare) , and both zeroes of are on the real axis. Furthermore, any slight perturbation of will also have both zeroes on the real axis.
The analogue of our result in this case is that , thus ruling out the third of the three scenarios here. In this simple example in which only two zeroes are involved, one can think of the inequality as asserting that if the zeroes of are real, then they must be repeated. In our result (in which there are an infinity of zeroes, that become increasingly dense near infinity), and in view of the convergence to local equilibrium properties of (3), the analogous assertion is that if the zeroes of are real, then they do not behave locally as if they were in arithmetic progression.
The Polymath14 online collaboration has uploaded to the arXiv its paper “Homogeneous length functions on groups“, submitted to Algebra & Number Theory. The paper completely classifies homogeneous length functions on an arbitrary group , that is to say non-negative functions that obey the symmetry condition , the non-degeneracy condition , the triangle inequality , and the homogeneity condition . It turns out that these norms can only arise from pulling back the norm of a Banach space by an isometric embedding of the group. Among other things, this shows that can only support a homogeneous length function if and only if it is abelian and torsion free, thus giving a metric description of this property.
The proof is based on repeated use of the homogeneous length function axioms, combined with elementary identities of commutators, to obtain increasingly good bounds on quantities such as , until one can show that such norms have to vanish. See the previous post for a full proof. The result is robust in that it allows for some loss in the triangle inequality and homogeneity condition, allowing for some new results on “quasinorms” on groups that relate to quasihomomorphisms.
As there are now a large number of comments on the previous post on this project, this post will also serve as the new thread for any final discussion of this project as it winds down.
Recent Comments