Archive for murmuration

Nature tidbits

Posted in Books, Kids, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 10, 2024 by xi'an

I was quite glad to have taken this 28 March issue of Nature for the train ride to Marseille earlier this month, as it proved full of great entries. Starting with a tribune from a PhD candidate from Ghana unable to speak at a tropical ecology conference in Lisbon for being denied a Schengen visa by the Dutch embassy in Accra, despite providing a serious amount of supporting material. This is a general issue related to international conferences that I plan to include in my presentation in a round table about the future of conference at the ISBA 2024 conference next month. Researchers and students from low-income countries are regularly victims of arbitrary rejections by consular and embassy officers, with little leeway from inviting universities and academic societies for counteracting such arbitrariness. An example came last year when, as part of our Data Science for Social Good program at Warwick, one computer student from Pakistan got denied a visa, while another one with an almost identical profile did receive her visa.

In News in focus, the development of a revolutionary (CAR-T) cancer therapy by  a Munbai company at 1/10 the cost of similar products in the US, with several impressive innovations (and a clinical trial over 33 people). The coverage of Michel Talagrand’s 2024 Abel prize, the deteriorating state of scientific research in Russia, a review on how the Big Bang got its name (hint: it was on BBC 3), and remembering India’s Chipko tree-hugging women of Western Himalayas of 50 years ago. The development of signs for scientific terms to add to the Indian and American sign languages.

There is also an almost sleuthing story on the debates in the ecology community about the theory of “Mother Tree”, introduced by Suzanne Simard, about trees communicating among themselves by underground fungal networks, and thus creating a “wood wide web“.  With mature trees favouring their own kin, a theory that critics in the field deem incompatible with evolutionary theory and more crucially missing supporting evidence. (Arguments from Simard that “the European male society hates the mother tree” concept do not help towards a rational denate. The Guardian now has a podcast about this debate.)

And another story on the (racial and gender) bias of AI image generators. With the conclusion that removing such biases goes through open sourcing. And regulation like the EU’s AI Act. that requires technical documentation on the training datasets. Plus a retrospective on the theory of bird-flight origins, with John Ostrom proposing in 1974 that the evolution to flight was ground-up rather than tree-down.

A cover-paper on IBM error correcting code for quantum computing that only requires each qubit to connect with six others (just like birds in murmuration only keep track of seven neighbours) for an error threshold less than 1%. (With the property that 12 logical qubits can be preserved for 10⁶ cyclces using 288 qubits in total.)

Mostly Monte Carlo Xminas

Posted in Kids, pictures, Statistics, University life with tags , , , , , , , , , , , , , on December 7, 2023 by xi'an

The next and last of 2023 occurrence of our monthly series of Parisian seminars on the theory and practice of Monte Carlo in statistics and data science, in conjunction with our ERC OCEAN project , will be on Friday 15 December. The next seminars will be on 15 January, 12 February, and 98 March.

4pm/16h CEST: SVBMC: Fast post-processing Bayesian inference with noisy evaluations of the likelihood

Grégoire Clarté – University of Helsinki, University of Edinburgh

In many cases, the exact likelihood is unavailable, and can only be accessed through a noisy and expensive process – for example, in Plasma Physics. Furthermore, Bayesian inference often comes in at a second moment, for example after running an optimization algorithm to find a MAP estimate. To tackle both these issues, we introduce Sparse Variational Bayesian Monte Carlo (SVBMC), a method for fast “post-processes” Bayesian inference for models with black-box and noisy likelihoods. SVBMC reuses all existing target density evaluations – for example, from previous optimizations or partial Markov Chain Monte Carlo runs – to build a sparse Gaussian process (GP) surrogate model of the log posterior density. Uncertain regions of the surrogate are then refined via active learning as needed. Our work builds on the Variational Bayesian Monte Carlo (VBMC) framework for sample-efficient inference, with several novel contributions. First, we make VBMC scalable to a large number of pre-existing evaluations via sparse GP regression, deriving novel Bayesian quadrature formulae and acquisition functions for active learning with sparse GPs. Second, we introduce noise shaping, a general technique to induce the sparse GP approximation to focus on high posterior density regions. Third, we prove theoretical results in support of the SVBMC refinement procedure. We validate our method on a variety of challenging synthetic scenarios and real-world applications. We find that SVBMC consistently builds good posterior approximations by post-processing of existing model evaluations from different sources, often requiring only a small number of additional density evaluations.

5pm/17h CEST: Variance reduction using control variates and importance sampling for applications in computational statistical physics

Urbain Vaes – INRIA, CERMICS

The scaling of the mobility coefficient associated with two-dimensional Langevin dynamics in a periodic potential as the friction vanishes is not well understood. Theoretical results are lacking, and numerical calculation of the mobility in the underdamped regime is challenging. In the first part of this talk, I will present a new variance reduction approach based on control variates for efficiently estimating the mobility of Langevin-type dynamics, together with numerical experiments illustrating the performance of the approach.

In the second part of this talk, we study an importance sampling approach for calculating averages with respect to multimodal probability distributions. Traditional Markov chain Monte Carlo methods to this end, which are based on time averages along a realization of a Markov process ergodic with respect to the target probability distribution, are usually plagued by a large variance due to the metastability of the process. The estimator we study is based on an ergodic average along a realization of an overdamped Langevin process for a modified potential. We obtain an explicit expression for the optimal biasing potential in dimension 1 and propose a general numerical approach for approximating the optimal potential in the multi-dimensional setting.

Ocean’s four!

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , on October 25, 2022 by xi'an

Fantastic news! The ERC-Synergy¹ proposal we submitted last year with Michael Jordan, Éric Moulines, and Gareth Roberts has been selected by the ERC (which explains for the trips to Brussels last month). Its acronym is OCEAN [hence the whale pictured by a murmuration of starlings!], which stands for On intelligenCE And Networks​: Mathematical and Algorithmic Foundations for Multi-Agent Decision-Making​. Here is the abstract, which will presumably turn public today along with the official announcement from the ERC:

Until recently, most of the major advances in machine learning and decision making have focused on a centralized paradigm in which data are aggregated at a central location to train models and/or decide on actions. This paradigm faces serious flaws in many real-world cases. In particular, centralized learning risks exposing user privacy, makes inefficient use of communication resources, creates data processing bottlenecks, and may lead to concentration of economic and political power. It thus appears most timely to develop the theory and practice of a new form of machine learning that targets heterogeneous, massively decentralized networks, involving self-interested agents who expect to receive value (or rewards, incentive) for their participation in data exchanges.

OCEAN will develop statistical and algorithmic foundations for systems involving multiple incentive-driven learning and decision-making agents, including uncertainty quantification at the agent’s level. OCEAN will study the interaction of learning with market constraints (scarcity, fairness), connecting adaptive microeconomics and market-aware machine learning.

OCEAN builds on a decade of joint advances in stochastic optimization, probabilistic machine learning, statistical inference, Bayesian assessment of uncertainty, computation, game theory, and information science, with PIs having complementary and internationally recognized skills in these domains. OCEAN will shed a new light on the value and handling data in a competitive, potentially antagonistic, multi-agent environment, and develop new theories and methods to address these pressing challenges. OCEAN requires a fundamental departure from standard approaches and leads to major scientific interdisciplinary endeavors that will transform statistical learning in the long term while opening up exciting and novel areas of research.

Since the ERC support in this grant mostly goes to PhD and postdoctoral positions, watch out for calls in the coming months or contact us at any time.

Continue reading →