Archive for data encryption

Data science ethics [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 5, 2025 by xi'an

Data science ethics (concepts, techniques and cautionary tales), by David Martens, was published in 2022 by Oxford University Press. The book is inspired by the author’s  course on Data Science and ethics he has been teaching at the University of Antwerp. (With a link to his slides.) The 255p book proceeds by decomposing the ethics of data science into its different steps: data gathering (Chap. 2), data preprocessing (Chap. 3), modelling (Chap. 4), evaluation (Chap. 5), and deployment (Chap. 6). Following the `FAT Flow Framework´, where FAT stands for fairness, accountability, and transparency.

Do not expect much maths, stats, or anything quantitative: this book is mostly about concepts, even though some (mostly well-known) illustrations are provided. Chapter 2 includes a description of encryption (with homomorphic encryption treated in Chapter 4). And somewhat improbably quantum computing. Differential privacy gets a few pages with not a single formula (until Chapter 4, again).

Chapter 3 covers k-anonymity, record linkage, reidentification, (through cautionary tales) and discrimination through biases in the (learning) dataset.  Chapter 4 is defining ε differential privacy with the Laplace randomization as a possible implementation and with no critical stance on the limitations of the concept. The computation limitations of homomorphic encryption are more clearly pointed out. Federated learning is only quickly mentioned. The section about measuring fairness and reducing bias implies that some prior knowledge is available about whom is potentially discriminated and which covariates to add to the model. The last section on explicability of predictions is worthwhile in signalling the difficulty with most (black box) AI but the example opposing an SVM model to a logistic model is not tremendously convincing in that neither model is true.

Chapter 5 addresses the crucial challenge of ethical evaluation in a rather verbose and vague manner. For instance, with no instruction on how to resist adversarial attacks. Or criticising p-hacking and multiple testing while missing the elephant in the room (p-values!). Drifting from the topic when discussing the misdeeds of Diederik Stapel. Most of the same goes about Chapter 6 and its take on ethical deployment, when going through examples such as Google’s policies in China. Or general musing on the impact of AI on societal inequalities (with mentions of companies and CEOs who have since then back-pedalled on their ethical engagement). These chapters are lacking in tools and (more) practical recommendations.

One interesting aspect of the book is the attention paid to the EU(ropean) aspect of these concerns, through the GDPR (General DAta Protection Regulations) adopted by the European Parliament in 2016. (There is also a brief mention of China’s regulations, but no details beyond a reference. Maybe the Chinese edition differs.)

document the attacks on USA’s evidence infrastructure [reposted]

Posted in Statistics, University life with tags , , , , , , , , , on February 23, 2025 by xi'an


WASHINGTON, D.C., February 14, 2025 — The Data Foundation today announced the launch of SAFE-Track (Secure Anonymous Federal Evidence, Data and Analysis Tracking), a new secure, encrypted portal that enables confidential reporting of impacts from recent changes to federal evidence and data activities. This initiative responds to reports of significant disruptions affecting private sector research firms, federal contractors, and ongoing evaluation projects across government.

“We are hearing initial reports from our private sector partners in the data and evaluation communities about widespread effects on America’s evidence infrastructure, from mass layoffs at research firms to the premature termination of critical studies,” said Dr. Nick Hart, President and CEO of the Data Foundation. “Through SAFE-Track, we aim to systematically understand these impacts and inform policymakers while protecting the privacy and security of those sharing information.”

SAFE-Track (www.safe-track.org) provides:

  • Complete anonymity for submissions

  • End-to-end encryption of all data

  • No requirement for email or personal identification

  • Option for secure follow-up communication through anonymous conversation codes

  • Protections against collection of personally identifiable information

The Data Foundation is particularly interested in documenting:

  • Business impacts on private sector firms and contractors

  • Effects on ongoing research and evaluation projects

  • Implications for evidence-based policymaking

  • Challenges in implementing the Evidence Act’s requirements

  • Real-time impacts on government effectiveness and efficiency

“This secure platform will help the Data Foundation gather information about how recent changes affect the data and evaluation community’s ability to measure government performance and ensure taxpayer dollars are spent effectively,” added Hart. “Understanding these impacts is essential for supporting the bipartisan commitment to evidence-based policymaking established by the Foundations for Evidence-Based Policymaking Act, which President Trump signed into law in 2019.”

The Data Foundation emphasizes that submissions should avoid including personally identifiable information or confidential business information unless respondents specifically choose to be identified. The Foundation will use aggregated findings to inform Congress, the White House, and other stakeholders about the scope and scale of impacts while maintaining strict confidentiality of individual responses.

Privacy-preserving Computing [book review]

Posted in Books, Statistics with tags , , , , , , , , , , , , , , on May 13, 2024 by xi'an

Privacy-preserving Computing for Big Data Analytics and AI, by Kai Chen and Qiang Yang, is a rather short 2024 CUP book translated from the 2022 Chinese version (by the authors).  It covers secret sharing, homomorphic encryption, oblivious transfer, garbled circuit, differential privacy, trusted execution environment, federated learning, privacy-preserving computing platforms, and case studies. The style is survey-like, meaning it often is too light for my liking, with too many lists of versions and extensions, and more importantly lacking in detail to rely (solely) on it for a course. At several times standing closer to a Wikipedia level introduction to a topic. For instance, the chapter on homomorphic encryption [Chap.5] does not connect with the (presumably narrow) picture I have of this method. And the chapter on differential privacy [Chap.6] does not get much further than Laplace and Gaussian randomization, as in eg the stochastic gradient perturbation of Abadi et al. (2016) the privacy requirement is hardly discussed. The chapter on federated leaning [Chap.8] is longer if not much more detailed, being based on a entire book on Federated learning whose Qiang Yang is the primary author. (With all figures in that chapter being reproduced from said book.)  The next chapter [Chap.9] describes to some extent several computing platforms that can be used for privacy purposes, such as FATE, CryptDB, MesaTEE, Conclave, and PrivPy, while the final one goes through case studies from different areas, but without enough depth to be truly formative for neophyte readers and students. Overall, too light for my liking.

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]

really random generators [again!]

Posted in Books, Statistics with tags , , , , , , , , , on March 2, 2020 by xi'an

A pointer sent me to Chemistry World and an article therein about “really random numbers“. Or “truly” random numbers. Or “exactly” random numbers. Not particularly different from the (in)famous lava lamp generator!

“Cronin’s team has developed a robot that can automatically grow crystals in a 10 by 10 array of vials, take photographs of them, and use measurements of their size, orientation, and colour to generate strings of random numbers. The researchers analysed the numbers generated from crystals grown in three solutions – including a solution of copper sulfate – and found that they all passed statistical tests for the quality of their randomness.” Chemistry World, Tom Metcalfe, 18 February 2020

The validation of this truly random generator is thus exactly the same as a (“bad”) pseudo-random generator, namely that in the law of large number sense, it fits the predicted behaviour. And thus the difference between them cannot be statistical, but rather cryptographic:

“…we considered the encryption capability of this random number generator versus that of a frequently used pseudorandom number generator, the Mersenne Twister.” Lee et al., Matter, February 10, 2020

Meaning that the knowledge of the starting point and of the deterministic transform for the Mersenne Twister makes it feasible to decipher, which is not the case for a physical and non-reproducible generator as the one advocated. One unclear aspect of the proposed generator is the time required to produce 10⁶, even though the authors mention that “the bit-generation rate is significantly lower than that in other methods”.