Archive for hybrid Monte Carlo

Bayesian decision-theory for data privacy [surfin’ the Oce’n, 30 April, INRIA Paris]

Posted in Statistics, University life with tags , , , , , , , , , , , , , , , , , on April 23, 2025 by xi'an

Abstract

The scientific and economic value of data continues to grow alongside technology advances. New hardware and software developments enable, but often require, larger and more complex datasets to function effectively. As the importance of input data to these systems becomes increasingly recognized, so too does the loss of privacy for data providers. In this context, data privacy emerges as a critical issue for fields such as statistics and machine learning, as well as for scientific and industrial endeavours that rely on sensitive data. We propose a framework for measuring privacy from a Bayesian decision-theoretic perspective. This framework enables the creation of new, purpose-driven privacy principles that are rigorously justified, while also allowing for the assessment of existing privacy definitions through decision theory. We pay particular attention to the privacy of deterministic algorithms, which are overlooked by current privacy standards, and to the privacy of N Monte Carlo samples drawn from an invariant distribution as N goes to infinity. We show that Probabilistic Differential Privacy is a special case of our framework and provide some new interpretations for Differential Privacy as a result.

keep meetings hybrid

Posted in Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on September 30, 2022 by xi'an

I was reading the latest ISBA Bulletin and the tribune by ISBA President Sudipto Banerjee celebrating the return to the physical ISBA World meeting, along with worries about participants who caught COVID there. (Unfortunately, one good friend of mine experienced symptoms that went beyond the mild cold-like ones I zoomed through a few days ago.) This particular issue of creating a COVID cluster [during coffee breaks?!] provides [me with] one further argument for my supporting hybrid and multimodal meetings on a general basis. Which should [imho] appear in the proposals for the 2026 and 2028 World Meetings (deadline on 31 October)…(The 2024 meeting in Venezia will certainly involve hybridicity! As will BayesComp in Levi.) Discussing the topic with others in some scientific committees recently made me realise this was not such a shared perspective, from reasons varying from worrying about balancing the budget, to zoom fatigue, to the added value of informal interactions. Still, there also are reasons for hybridising our meetings, from reduced travel impact, to more inclusiveness,  on geographical, diversity, affordability, seniority grounds. Holding hybrid conferences with multiple regional mirrors allows for a potentially higher degree of interaction and local input.  And a minimal organisational effort.

congrats, Dr. Clarté!

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , on October 9, 2021 by xi'an

Grégoire Clarté, whom I co-supervised with Robin Ryder, successfully defended his PhD thesis last Wednesday! On sign language classification, ABC-Gibbs and collective non-linear MCMC. Congrats to the now Dr.Clarté for this achievement and all the best for his coming Nordic adventure, as he is starting a postdoc at the University of Helsinki, with Aki Vehtari and others. It was quite fun to work with Grégoire along these years. And discussing on an unlimited number of unrelated topics, incl. fantasy books, teas, cooking and the role of conferences and travel in academic life! The defence itself proved a challenge as four members of the jury, incl. myself, were “present remotely” and frequently interrupted him for gaps in the Teams transmission, which nonetheless broadcasted perfectly the honks of the permanent traffic jam in Porte Dauphine… (And alas could not share a celebratory cup with him!)

convergence of MCMC

Posted in Statistics with tags , , , , , , , , , on June 16, 2017 by xi'an

Michael Betancourt just posted on arXiv an historical  review piece on the convergence of MCMC, with a physical perspective.

“The success of these of Markov chain Monte Carlo, however, contributed to its own demise.”

The discourse proceeds through augmented [reality!] versions of MCMC algorithms taking advantage of the shape and nature of the target distribution, like Langevin diffusions [which cannot be simulated directly and exactly at the same time] in statistics and molecular dynamics in physics. (Which reminded me of the two parallel threads at the ICMS workshop we had a few years ago.) Merging into hybrid Monte Carlo, morphing into Hamiltonian Monte Carlo under the quills of Radford Neal and David MacKay in the 1990’s. It is a short entry (and so is this post), with some background already well-known to the community, but it nonetheless provides a perspective and references rarely mentioned in statistics.

slice sampling revisited

Posted in Books, pictures, Statistics with tags , , , , , , , , on April 15, 2016 by xi'an

Figure 1 (c.) Neal, 2003Thanks to an X validated question, I re-read Radford Neal’s 2003 Slice sampling paper. Which is an Annals of Statistics discussion paper, and rightly so. While I was involved in the editorial processing of this massive paper (!), I had only vague memories left about it. Slice sampling has this appealing feature of being the equivalent of random walk Metropolis-Hastings for Gibbs sampling, without the drawback of setting a scale for the moves.

“These slice sampling methods can adaptively change the scale of changes made, which makes them easier to tune than Metropolis methods and also avoids problems that arise when the appropriate scale of changes varies over the distribution  (…) Slice sampling methods that improve sampling by suppressing random walks can also be constructed.” (p.706)

One major theme in the paper is fighting random walk behaviour, of which Radford is a strong proponent. Even at the present time, I am a bit surprised by this feature as component-wise slice sampling is exhibiting clear features of a random walk, exploring the subgraph of the target by random vertical and horizontal moves. Hence facing the potential drawback of backtracking to previously visited places.

“A Markov chain consisting solely of overrelaxed updates might not be ergodic.” (p.729)

Overrelaxation is presented as a mean to avoid the random walk behaviour by removing rejections. The proposal is actually deterministic projecting the current value to the “other side” of the approximate slice. If it stays within the slice it is accepted. This “reflection principle” [in that it takes the symmetric wrt the centre of the slice] is also connected with antithetic sampling in that it induces rather negative correlation between the successive simulations. The last methodological section covers reflective slice sampling, which appears as a slice version of Hamiltonian Monte Carlo (HMC). Given the difficulty in implementing exact HMC (reflected in the later literature), it is no wonder that Radford proposes an approximation scheme that is valid if somewhat involved.

“We can show invariance of this distribution by showing (…) detailed balance, which for a uniform distribution reduces to showing that the probability density for x¹ to be selected as the next state, given that the current state is x0, is the same as the probability density for x⁰ to be the next state, given that x¹ is the current state, for any states x⁰ and x¹ within [the slice] S.” (p.718)

In direct connection with the X validated question there is a whole section of the paper on implementing single-variable slice sampling that I had completely forgotten, with a collection of practical implementations when the slice

S={x; u < f(x) }

cannot be computed in an exact manner. Like the “stepping out” procedure. The resulting set (interval) where the uniform simulation in x takes place may well miss some connected component(s) of the slice. This quote may sound like a strange argument in that the move may well leave a part of the slice off and still satisfy this condition. Not really since it states that it must hold for any pair of states within S… The very positive side of this section is to allow for slice sampling in cases where the inversion of u < f(x) is intractable. Hence with a strong practical implication. The multivariate extension of the approximation procedure is more (potentially) fraught with danger in that it may fell victim to a curse of dimension, in that the box for the uniform simulation of x may be much too large when compared with the true slice (or slice of the slice). I had more of a memory of the “trail of crumbs” idea, mostly because of the name I am afraid!, which links with delayed rejection, as indicated in the paper, but seems awfully delicate to calibrate.