Archive for neural network

the Harvard and Brown school of computer science

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on October 2, 2025 by xi'an

 “In the late 1980s, LeCun, then a researcher at AT&T Bell Labs, developed a powerful neural network that learned to recognise handwritten zip codes by training on thousands of examples. A parallel development soon unfolded at Harvard and Brown. In 1995, Zhu and a team of researchers there started developing probability-based methods that could learn to recognise patterns and textures (…) and even generate new examples of that pattern. These were not neural networks: members of the “Harvard-Brown school”, as Zhu called his team, cast vision as a problem of statistics and relied on methods such as “Bayesian inference” and “Markov random fields”. The two schools spoke different mathematical languages and had philosophical disagreements. But they shared an underlying logic – that data, rather than hand-coded instructions, could supply the infrastructure for machines to grasp the world and reproduce its patterns – that exists in today’s AI systems such as ChatGPT.” 

next OWABI webinar [30 Jan]

Posted in pictures, Statistics, Uncategorized with tags , , , , , , , , , , , , on January 19, 2025 by xi'an


The next One World Approximate Bayesian Inference (OWABI) Seminar is scheduled on Thursday, the 30th January at 11am UK time, with the speaker being Paul Bürkner (TU Dortmund University),

Amortized Mixture and Multilevel Models

Abstract: Probabilistic mixture and multilevel models are central building blocks in Bayesian data analysis. However, they remain challenging to estimate and evaluate, especially when the involved likelihoods or priors are analytically intractable. Recent developments in generative deep learning and simulation-based inference have shown promising results in scaling up Bayesian inference through amortization. Against this background, we have developed specialized neural inference frameworks for estimating Bayesian mixture and multilevel models. The involved neural architectures are closely mirroring the probabilistic symmetries and conditional (in-)dependencies assumed by these models. This not only speeds up neural network training, but also enables amortized inference for new datasets of varying number of groups and sample sizes.

Keywords: Amortized Bayesian Inference; Neural Posterior Estimation; Probabilistic Factorization

NobAIl prAIzes

Posted in Books, pictures, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on October 11, 2024 by xi'an

I am quite surprised that the NobAIl committee did not select ChatGPT‘s Sam Altman for their literature prize—since he made a reality of monkeys typing at random on a typewriter, hence bound to a.s. produce past and future masterpieces—, now betting on Twitter’s Dorsey, Williams and Stone for their peace prize— for achieving worldwide harmony if within each single-minded community in only 140 characters—, and Amazon’s Jeff Bezos for their Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel—for produing the ultimate monopoly of a single worldwide convenience store—, given their picks for physics—with spin glass Ising models that have always sounded to me like the worst possible illustration for MCMC techniques— and chemistry—for a deep learning predictor of protein structures, built by large teams and numerous CPU hours, achieving high success rates if not perfection, but prediction is not explanation, reminding me of Nietzsche’s “physics, too, is only an interpretation and exegesis of the world (to suit us, if I may say so!) and not a world-explanation”—… Keeping their sharp focus on AI’s, corporate funded research, and male recipients… (As coïncidences come, I am currently reading a book of the literature 2024 Nobel recipient, Han Kang, The Vegetarian, that I find stunning and immensely original!)

veniSBA²

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , on July 4, 2024 by xi'an

After another morning cycle of 2Xing Porte della Libertà (under a light and pleasant rain) and swimming in Sant’ Alviso (in too warm a water), I did not make it for the beginning of the Bayesian deep learning session, breakfast oblige!, and cumulated with different percolation events (ie, meeting friend after friend on my way to the classroom), I could not get enough of the session to report anything even barely useful!

As I did not rush fast enough to Andrew’s Foundation lecture (another sequence of percolations!), I had to stand in the back of the packed main amphitheatre (and former sorting hall of the Venice slaughterhouse!), Guido Cazzavillan’s Aula Magna, while he talked a fresco about some holes in Bayesian data analysis (the analysis, not the book!), those being [verbatim]

  1. the usual rules of conditional probability fail in the quantum realm,
  2. flat or weak priors lead to terrible inferences about things we care about,
  3. subjective priors are incoherent,
  4. Bayesian decision picks the wrong model,
  5. Bayes factors fail in the presence of flat or weak priors,
  6. for Cantorian reasons we need to check our models, but this destroys the coherence of Bayesian inference.

After lunch, I attended the (mostly sequential) simulation based inference (renamed from ABC!) session with a composite likelihood proposal by Lorenzo Rimella, that uses marginals to approximate the likelihood of a hidden Markov SIS epidemic model by composite likelihood towards getting more efficient if inexact versions. Then [1WABC webinar co-organiser] Umberto Picchini on surrogates for likelihood and posterior functions, with sequential improvements (w/o ABC and w/o neural networks). Called “Sequential mixture posterior and likelihood estimation”, using mixtures of experts when the weights are functions of the observed or simulated y. With adapting the number of components in the mixture. Comparing favourably with normalising flows. And Wentao Li on correcting by ABC for composite likelihood as in Ruli et al. (2016). Where a posterior distribution given composite scores (seen as [summary] statistics) is employed but requires a convergent estimator of the unknown parameter.

 No congratulation today to our PhD student who managed to fall in a canal (but survived)..!

6th Workshop on Sequential Monte Carlo Methods

Posted in Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , on May 16, 2024 by xi'an

Very glad to be back to an SMC workshop as it has been nine years since my attending SMC 2015 in Malakoff! The more for the workshop taking place in Edinburgh and at the Bayes Centre. It is one of these places where I feel somewhat returning to familiar grounds with accumulated memories. Like my last visit there when I had a tea with Mike Titterington…

The overall pace of the workshop was quite nice, with long breaks for informal discussions (and time for ‘oggin’!) and interesting poster late afternoons, helped by the small number of them at each instance, incl. one on reversible jump HMC. Here are a few scribbled entries about some talks along the first two days.

After my opening talk (!), Joaquín Míguez talked about the impact of a sequential (Euler-Marayama) discretisation scheme for stochastic differential equations on Bayesian filtering with control of the approximation effect. Axel Finke (in a joint work with Adrien Corenflos, now an ERC Ocean postdoc in Warwick) built a sequence of particle filter algorithms targeting good performances (high expected jumping distance) against both large dimensions and high time horizon, exploiting gradient shift MALA-like, as well as prior impact, with the conclusion that their jack-of-all-trades solutions, Particle­-MALA and Particle­-mGRAD, enjoyed this resistance in nearly normal models. Interesting reminder of the auxiliary particle trick and good insights on using the smoothing target, even when accounting for the computing time, but too many versions for a single talk without checking against the preprint.

The SMC sampler-like algorithm involves propagating N “seed” particles z(i), with a mutation mechanism consisting of the generation of N integrator snippets 𝗓:=(z,ψ⁢(z),ψ²⁢(z),…) started at every seed particle z(i), resulting in N×(T+1) particles which are then whittled down to a set of N seed particles using a standard resampling scheme. Andrieu et al., 2024

Christophe Andrieu talked about Monte Carlo sampling with integrator snippets, starting with recycling solutions for the leapfrog integrator HMC and unfolding Hamiltonians for moving more easily. With snippets representing discretised paths along the level sets being used as particles, picking zero, one, or more particles along each path, since importance weights are connection with multinomial HMC

This relatively small algorithmic modification of the conditional particle filter, which we call the conditional backward sampling particle filter has a dramatically improved performance over the conditional particle filter. Karjalainen et al., 2024

Anthony Lee looked at mixing times for backward sampling SMC (CBPF/ancestor sampling) cf Lee et al. (2020), where the backward step consists in computing the weight of a randomly drawn backward or ancestral history. Improving on earlier results to reach mixing time O(log T) and complexity O(T log T) (with T the time horizon). Thanks to maximal coupling and boundedness assumptions on the prior and likelihood functions.

Neil Chada presented a work on Bayesian multilevel Monte Carlo on deep networks. À la Giles, with a telescoping identity. Always puzzling to envision a prior on all parameters of a neural network. Achieving a computational cost inverse to the order of the MSE, at best. With a useful reminder that pushing the size of the NN to infinity results in a (poor) Gaussian process prior (Sell et al., 2023).

On my first evening, I stopped with a friend in my favourite Blonde [restaurant], as in almost every other visit to Edinburgh, enjoyable as always, but I also found the huge offer of Asian minimarkets in the area too tempting to resist, between Indian, Korean, and Chinese products. (Although with a disappointing hojicha!). As I could not reach any new Munro by train or bus within a reasonable time range I resorted to the nearer Pentland Hills, with a stop by Rosslyn Chapel (mostly of Da Vinci Code fame!, if classic enough). And some delays in finding a bus getting there (misled by google map!) and a trail (misled by my poor map reading skills) up the actual hills. The mist did not help either.