April 2020
S	M	T	W	T	F	S

			1	2	3	4
5	6	7	8	9	10	11
12	13	14	15	16	17	18
19	20	21	22	23	24	25
26	27	28	29	30

Fri, 25 December 2020

Earthquake Detection with Crowd-sourced Data

Have you ever wanted to hear what an earthquake sounds like? Today on the show we have Omkar Ranadive, Computer Science Masters student at NorthWestern University, who collaborates with Suzan van der Lee, an Earth and Planetary Sciences professor at Northwestern University, on the crowd-sourcing project Earthquake Detective.

Email Links:
Suzan: suzan@earth.northwestern.edu
Omkar: omkar.ranadive@u.northwestern.edu

Works Mentioned:

Paper: Applying Machine Learning to Crowd-sourced Data from Earthquake Detective
https://arxiv.org/abs/2011.04740
by Omkar Ranadive, Suzan van der Lee, Vivan Tang, and Kevin Chao
Github: https://github.com/Omkar-Ranadive/Earthquake-Detective
Earthquake Detective: https://www.zooniverse.org/projects/vivitang/earthquake-detective

Thanks to our sponsors!

Brilliant.org Is an awesome platform with interesting courses, like Quantum Computing! There is something for you and surely something for the whole family! Get 20% off Brilliant Premium at http://brilliant.com/dataskeptic

Direct download: earthquake-detection-with-crowd-sourced-data.mp3
Category:general -- posted at: 8:21am PDT

Tue, 22 December 2020

Byzantine Fault Tolerant Consensus

Byzantine fault tolerance (BFT) is a desirable property in a distributed computing environment. BFT means the system can survive the loss of nodes and nodes becoming unreliable. There are many different protocols for achieving BFT, though not all options can scale to large network sizes.

Ted Yin joins us to explain BFT, survey the wide variety of protocols, and share details about HotStuff.

Direct download: byzantine-fault-tolerant-consensus.mp3
Category:general -- posted at: 5:00am PDT

Fri, 11 December 2020

Alpha Fold

Kyle shared some initial reactions to the announcement about Alpha Fold 2's celebrated performance in the CASP14 prediction. By many accounts, this exciting result means protein folding is now a solved problem.

Thanks to our sponsors!

Brilliant is a great last-minute gift idea! Give access to 60 + interactive courses including Quantum Computing and Group Theory. There's something for everyone at Brilliant. They have award-winning courses, taught by teachers, researchers and professionals from MIT, Caltech, Duke, Microsoft, Google and many more. Check them out at brilliant.org/dataskeptic to take advantage of 20% off a Premium memebership.
Betterhelp is an online professional counseling platform. Start communicating with a licensed professional in under 24 hours! It's safe, private and convenient. From online messages to phone and video calls, there is something for everyone. Get 10% off your first month at betterhelp.com/dataskeptic

Direct download: alpha-fold.mp3
Category:general -- posted at: 9:45am PDT

Fri, 4 December 2020

Arrow's Impossibility Theorem

Above all, everyone wants voting to be fair. What does fair mean and how can we measure it? Kenneth Arrow posited a simple set of conditions that one would certainly desire in a voting system. For example, unanimity - if everyone picks candidate A, then A should win!

Yet surprisingly, under a few basic assumptions, this theorem demonstrates that no voting system exists which can satisfy all the criteria.

This episode is a discussion about the structure of the proof and some of its implications.

Works Mentioned

A Difficulty in the Concept of Social Welfare by Kenneth J. Arrow

Three Brief Proofs of Arrows Impossibility Theorem by John Geanakoplos

Thank you to our sponsors!

Better Help is much more affordable than traditional offline counseling, and financial aid is available! Get started in less than 24 hours. Data Skeptic listeners get 10% off your first month when you visit: betterhelp.com/dataskeptic

Let Springboard School of Data jumpstart your data career! With 100% online and remote schooling, supported by a vast network of professional mentors with a tuition-back guarantee, you can't go wrong. Up to twenty $500 scholarships will be awarded to Data Skeptic listeners. Check them out at springboard.com/dataskeptic and enroll using code: DATASK

Direct download: arrows-impossibility-theorem.mp3
Category:general -- posted at: 8:39am PDT

Fri, 27 November 2020

Face Mask Sentiment Analysis

As the COVID-19 pandemic continues, the public (or at least those with Twitter accounts) are sharing their personal opinions about mask-wearing via Twitter. What does this data tell us about public opinion? How does it vary by demographic? What, if anything, can make people change their minds?

Today we speak to, Neil Yeung and Jonathan Lai, Undergraduate students in the Department of Computer Science at the University of Rochester, and Professor of Computer Science, Jiebo-Luoto to discuss their recent paper. Face Off: Polarized Public Opinions on Personal Face Mask Usage during the COVID-19 Pandemic.

Works Mentioned
https://arxiv.org/abs/2011.00336

Emails:
Neil Yeung
nyeung@u.rochester.edu

Jonathan Lia
jlai11@u.rochester.edu

Jiebo Luo
jluo@cs.rochester.edu

Thanks to our sponsors!

Springboard School of Data offers a comprehensive career program encompassing data science, analytics, engineering, and Machine Learning. All courses are online and tailored to fit the lifestyle of working professionals. Up to 20 Data Skeptic listeners will receive $500 scholarships. Apply today at springboard.com/datasketpic
Check out Brilliant's group theory course to learn about object-oriented design! Brilliant is great for learning something new or to get an easy-to-look-at review of something you already know. Check them out a Brilliant.org/dataskeptic to get 20% off of a year of Brilliant Premium!

Direct download: face-mask-sentiment-analysis.mp3
Category:general -- posted at: 10:56am PDT

Fri, 20 November 2020

Counting Briberies in Elections

Niclas Boehmer, second year PhD student at Berlin Institute of Technology, comes on today to discuss the computational complexity of bribery in elections through the paper “On the Robustness of Winners: Counting Briberies in Elections.”

Links Mentioned:
https://www.akt.tu-berlin.de/menue/team/boehmer_niclas/

Works Mentioned:
“On the Robustness of Winners: Counting Briberies in Elections.” by Niclas Boehmer, Robert Bredereck, Piotr Faliszewski. Rolf Niedermier

Thanks to our sponsors:

Springboard School of Data: Springboard is a comprehensive end-to-end online data career program. Create a portfolio of projects to spring your career into action. Learn more about how you can be one of twenty $500 scholarship recipients at springboard.com/dataskeptic. This opportunity is exclusive to Data Skeptic listeners. (Enroll with code: DATASK)

Nord VPN: Protect your home internet connection with unlimited bandwidth. Data Skeptic Listeners-- take advantage of their Black Friday offer: purchase a 2-year plan, get 4 additional months free. nordvpn.com/dataskeptic (Use coupon code DATASKEPTIC)

Direct download: counting-briberies-in-elections.mp3
Category:general -- posted at: 8:26am PDT

Fri, 13 November 2020

Sybil Attacks on Federated Learning

Clement Fung, a Societal Computing PhD student at Carnegie Mellon University, discusses his research in security of machine learning systems and a defense against targeted sybil-based poisoning called FoolsGold.

Works Mentioned:
The Limitations of Federated Learning in Sybil Settings

Twitter:

@clemfung

Website:
https://clementfung.github.io/

Thanks to our sponsors:

Brilliant - Online learning platform. Check out Geometry Fundamentals! Visit Brilliant.org/dataskeptic for 20% off Brilliant Premium!

BetterHelp - Convenient, professional, and affordable online counseling. Take 10% off your first month at betterhelp.com/dataskeptic

Direct download: sybil-attacks-on-federated-learning.mp3
Category:general -- posted at: 10:25am PDT

Fri, 6 November 2020

Differential Privacy at the US Census

Simson Garfinkel, Senior Computer Scientist for Confidentiality and Data Access at the US Census Bureau, discusses his work modernizing the Census Bureau disclosure avoidance system from private to public disclosure avoidance techniques using differential privacy. Some of the discussion revolves around the topics in the paper Randomness Concerns When Deploying Differential Privacy.

WORKS MENTIONED:

“Calibrating Noise to Sensitivity in Private Data Analysis” by Cynthia Dwork, Frank McSherry, Kobbi Nissim, Adam Smith
"Issues Encountered Deploying Differential Privacy" by Simson L Garfinkel, John M Abowd, and Sarah Powazek
"Randomness Concerns When Deploying Differential Privacy" by Simson L. Garfinkel and Philip Leclerc

Check out: https://simson.net/page/Differential_privacy

Thank you to our sponsor, BetterHelp. Professional and confidential in-app counseling for everyone. Save 10% on your first month of services with www.betterhelp.com/dataskeptic

Direct download: differential-privacy-at-the-us-census.mp3
Category:general -- posted at: 8:13am PDT

Thu, 29 October 2020

Distributed Consensus

Computer Science research fellow of Cambridge University, Heidi Howard discusses Paxos, Raft, and distributed consensus in distributed systems alongside with her work “Paxos vs. Raft: Have we reached consensus on distributed consensus?”

She goes into detail about the leaders in Paxos and Raft and how The Raft Consensus Algorithm actually inspired her to pursue her PhD.

Paxos vs Raft paper: https://arxiv.org/abs/2004.05074

Leslie Lamport paper “part-time Parliament”
https://lamport.azurewebsites.net/pubs/lamport-paxos.pdf

Leslie Lamport paper "Paxos Made Simple"
https://lamport.azurewebsites.net/pubs/paxos-simple.pdf

Twitter : @heidiann360

Thank you to our sponsor Monday.com! Their apps challenge is still accepting submissions! find more information at monday.com/dataskeptic

Direct download: distributed-consensus.mp3
Category:general -- posted at: 10:36pm PDT

Fri, 23 October 2020

ACID Compliance

Linhda joins Kyle today to talk through A.C.I.D. Compliance (atomicity, consistency, isolation, and durability). The presence of these four components can ensure that a database’s transaction is completed in a timely manner. Kyle uses examples such as google sheets, bank transactions, and even the game rummy cube.

Thanks to this week's sponsors:

Monday.com - Their Apps Challenge is underway and available at monday.com/dataskeptic
Brilliant - Check out their Quantum Computing Course, I highly recommend it! Other interesting topics I’ve seen are Neural Networks and Logic. Check them out at Brilliant.org/dataskeptic

Direct download: acid-compliance.mp3
Category:general -- posted at: 6:00am PDT

Fri, 16 October 2020

Patrick Rosenstiel joins us to discuss the The National Popular Vote.

Direct download: national-popular-vote-interstate-compact.mp3
Category:general -- posted at: 8:24am PDT

Mon, 12 October 2020

Defending the p-value

Yudi Pawitan joins us to discuss his paper Defending the P-value.

Direct download: defending-the-p-value.mp3
Category:general -- posted at: 6:00am PDT

Mon, 5 October 2020

Retraction Watch

Ivan Oransky joins us to discuss his work documenting the scientific peer-review process at retractionwatch.com.

Direct download: retraction-watch.mp3
Category:general -- posted at: 8:00am PDT

Mon, 21 September 2020

Crowdsourced Expertise

Derek Lim joins us to discuss the paper Expertise and Dynamics within Crowdsourced Musical Knowledge Curation: A Case Study of the Genius Platform.

Direct download: crowdsourced-expertise.mp3
Category:general -- posted at: 7:00am PDT

Mon, 14 September 2020

The Spread of Misinformation Online

Neil Johnson joins us to discuss the paper The online competition between pro- and anti-vaccination views.

Direct download: the-spread-of-misinformation-online.mp3
Category:general -- posted at: 7:00am PDT

Mon, 7 September 2020

Consensus Voting

Mashbat Suzuki joins us to discuss the paper How Many Freemasons Are There? The Consensus Voting Mechanism in Metric Spaces.

Check out Mashbat’s and many other great talks at the 13th Symposium on Algorithmic Game Theory (SAGT 2020)

Direct download: consensus-voting.mp3
Category:general -- posted at: 7:00am PDT

Mon, 31 August 2020

Voting Mechanisms

Steven Heilman joins us to discuss his paper Designing Stable Elections.

For a general interest article, see: https://theconversation.com/the-electoral-college-is-surprisingly-vulnerable-to-popular-vote-changes-141104

Steven Heilman receives funding from the National Science Foundation. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author and do not necessarily reflect the views of the National Science Foundation.

Direct download: voting-mechanisms.mp3
Category:general -- posted at: 7:00am PDT

Mon, 24 August 2020

False Consensus

Sami Yousif joins us to discuss the paper The Illusion of Consensus: A Failure to Distinguish Between True and False Consensus. This work empirically explores how individuals evaluate consensus under different experimental conditions reviewing online news articles.

More from Sami at samiyousif.org

Link to survey mentioned by Daniel Kerrigan: https://forms.gle/TCdGem3WTUYEP31B8

Direct download: false-concensus.mp3
Category:general -- posted at: 3:16pm PDT

Tue, 18 August 2020

Fraud Detection in Real Time

In this solo episode, Kyle overviews the field of fraud detection with eCommerce as a use case. He discusses some of the techniques and system architectures used by companies to fight fraud with a focus on why these things need to be approached from a real-time perspective.

Direct download: fraud-detection-in-real-time.mp3
Category:general -- posted at: 12:12am PDT

Tue, 11 August 2020

Listener Survey Review

In this episode, Kyle and Linhda review the results of our recent survey. Hear all about the demographic details and how we interpret these results.

Direct download: listener-survey-review.mp3
Category:general -- posted at: 10:01am PDT

Mon, 27 July 2020

Human Computer Interaction and Online Privacy

Moses Namara from the HATLab joins us to discuss his research into the interaction between privacy and human-computer interaction.

Direct download: human-computer-interaction-and-online-privacy.mp3
Category:general -- posted at: 2:43pm PDT

Mon, 20 July 2020

Authorship Attribution of Lennon McCartney Songs

Mark Glickman joins us to discuss the paper Data in the Life: Authorship Attribution in Lennon-McCartney Songs.

Direct download: authorship-attribution-of-lennon-mccartney-songs.mp3
Category:general -- posted at: 8:00am PDT

Fri, 10 July 2020

GANs Can Be Interpretable

Erik Härkönen joins us to discuss the paper GANSpace: Discovering Interpretable GAN Controls. During the interview, Kyle makes reference to this amazing interpretable GAN controls video and it’s accompanying codebase found here. Erik mentions the GANspace collab notebook which is a rapid way to try these ideas out for yourself.

Direct download: gans-can-be-interpretable.mp3
Category:general -- posted at: 7:42pm PDT

Mon, 6 July 2020

Sentiment Preserving Fake Reviews

David Ifeoluwa Adelani joins us to discuss Generating Sentiment-Preserving Fake Online Reviews Using Neural Language Models and Their Human- and Machine-based Detection.

Direct download: sentiment-preserving-fake-reviews.mp3
Category:general -- posted at: 3:48pm PDT

Fri, 26 June 2020

Interpretability Practitioners

Sungsoo Ray Hong joins us to discuss the paper Human Factors in Model Interpretability: Industry Practices, Challenges, and Needs.

Direct download: interpretability-practitioners.mp3
Category:general -- posted at: 9:43am PDT

Fri, 19 June 2020

Facial Recognition Auditing

Deb Raji joins us to discuss her recent publication Saving Face: Investigating the Ethical Concerns of Facial Recognition Auditing.

Direct download: facial-recognition-auditing.mp3
Category:general -- posted at: 11:34am PDT

Fri, 12 June 2020

Robust Fit to Nature

Uri Hasson joins us this week to discuss the paper Robust-fit to Nature: An Evolutionary Perspective on Biological (and Artificial) Neural Networks.

Direct download: robust-fit-to-nature.mp3
Category:general -- posted at: 8:56am PDT

Fri, 5 June 2020

Black Boxes Are Not Required

Deep neural networks are undeniably effective. They rely on such a high number of parameters, that they are appropriately described as “black boxes”.

While black boxes lack desirably properties like interpretability and explainability, in some cases, their accuracy makes them incredibly useful.

But does achiving “usefulness” require a black box? Can we be sure an equally valid but simpler solution does not exist?

Cynthia Rudin helps us answer that question. We discuss her recent paper with co-author Joanna Radin titled (spoiler warning)…

Why Are We Using Black Box Models in AI When We Don’t Need To? A Lesson From An Explainable AI Competition

Direct download: black-boxes-are-not-required.mp3
Category:general -- posted at: 12:59pm PDT

Sat, 30 May 2020

Robustness to Unforeseen Adversarial Attacks

Daniel Kang joins us to discuss the paper Testing Robustness Against Unforeseen Adversaries.

Direct download: robustness-to-unforeseen-adversarial-attacks.mp3
Category:general -- posted at: 8:29am PDT

Fri, 22 May 2020

Estimating the Size of Language Acquisition

Frank Mollica joins us to discuss the paper Humans store about 1.5 megabytes of information during language acquisition

Direct download: estimating-the-size-of-language-acquisition.mp3
Category:general -- posted at: 2:36pm PDT

Fri, 15 May 2020

Interpretable AI in Healthcare

Jayaraman Thiagarajan joins us to discuss the recent paper Calibrating Healthcare AI: Towards Reliable and Interpretable Deep Predictive Models.

Direct download: interpretable-ai-in-healthcare.mp3
Category:general -- posted at: 8:49am PDT

Fri, 8 May 2020

Understanding Neural Networks

What does it mean to understand a neural network? That’s the question posted on this arXiv paper. Kyle speaks with Tim Lillicrap about this and several other big questions.

Direct download: understanding-neural-networks.mp3
Category:general -- posted at: 10:07am PDT

Fri, 1 May 2020

Self-Explaining AI

Dan Elton joins us to discuss self-explaining AI. What could be better than an interpretable model? How about a model wich explains itself in a conversational way, engaging in a back and forth with the user.

We discuss the paper Self-explaining AI as an alternative to interpretable AI which presents a framework for self-explainging AI.

Direct download: self-explaining-ai.mp3
Category:general -- posted at: 10:23pm PDT

Fri, 24 April 2020

Plastic Bag Bans

Becca Taylor joins us to discuss her work studying the impact of plastic bag bans as published in Bag Leakage: The Effect of Disposable Carryout Bag Regulations on Unregulated Bags from the Journal of Environmental Economics and Management. How does one measure the impact of these bans? Are they achieving their intended goals? Join us and find out!

Direct download: plastic-bag-bans.mp3
Category:general -- posted at: 8:45am PDT

Sat, 18 April 2020

Self Driving Cars and Pedestrians

We are joined by Arash Kalatian to discuss Decoding pedestrian and automated vehicle interactions using immersive virtual reality and interpretable deep learning.

Direct download: self-driving-cars-and-pedestrians.mp3
Category:general -- posted at: 10:58am PDT

Fri, 10 April 2020

Computer Vision is Not Perfect

Computer Vision is not Perfect

Julia Evans joins us help answer the question why do neural networks think a panda is a vulture. Kyle talks to Julia about her hands-on work fooling neural networks.

Julia runs Wizard Zines which publishes works such as Your Linux Toolbox. You can find her on Twitter @b0rk

Direct download: computer-vision-is-not-perfect.mp3
Category:general -- posted at: 10:53am PDT

Sat, 4 April 2020

Uncertainty Representations

Jessica Hullman joins us to share her expertise on data visualization and communication of data in the media. We discuss Jessica’s work on visualizing uncertainty, interviewing visualization designers on why they don't visualize uncertainty, and modeling interactions with visualizations as Bayesian updates.

Homepage: http://users.eecs.northwestern.edu/~jhullman/

Lab: MU Collective

Direct download: uncertainty-representations.mp3
Category:general -- posted at: 8:18am PDT

Fri, 27 March 2020

AlphaGo, COVID-19 Contact Tracing and New Data Set

Announcing Journal Club

I am pleased to announce Data Skeptic is launching a new spin-off show called "Journal Club" with similar themes but a very different format to the Data Skeptic everyone is used to.

In Journal Club, we will have a regular panel and occasional guest panelists to discuss interesting news items and one featured journal article every week in a roundtable discussion. Each week, I'll be joined by Lan Guo and George Kemp for a discussion of interesting data science related news articles and a featured journal or pre-print article.

We hope that this podcast will give listeners an introduction to the works we cover and how people discuss these works. Our topics will often coincide with the original Data Skeptic podcast's current Interpretability theme, but we have few rules right now or what we pick. We enjoy discussing these items with each other and we hope you will do.

In the coming weeks, we will start opening up the guest chair more often to bring new voices to our discussion. After that we'll be looking for ways we can engage with our audience.

Keep reading and thanks for listening!

Kyle

Direct download: AlphaGo_COVID-19_Contact_Tracing_and_New_Data_Set.mp3
Category:general -- posted at: 11:00pm PDT

Fri, 20 March 2020

Visualizing Uncertainty

Direct download: visualizing-uncertainty.mp3
Category:general -- posted at: 8:00am PDT

Fri, 13 March 2020

Interpretability Tooling

Pramit Choudhary joins us to talk about the methodologies and tools used to assist with model interpretability.

Direct download: interpretability-tooling.mp3
Category:general -- posted at: 8:00am PDT

Fri, 6 March 2020

Shapley Values

Kyle and Linhda discuss how Shapley Values might be a good tool for determining what makes the cut for a home renovation.

Direct download: shapley-values.mp3
Category:general -- posted at: 12:29pm PDT

Fri, 28 February 2020

Anchors as Explanations

We welcome back Marco Tulio Ribeiro to discuss research he has done since our original discussion on LIME.

In particular, we ask the question Are Red Roses Red? and discuss how Anchors provide high precision model-agnostic explanations.

Please take our listener survey.

Direct download: anchors-as-explanations.mp3
Category:general -- posted at: 6:46am PDT

Fri, 21 February 2020

Mathematical Models of Ecological Systems

Direct download: mathematical-models-of-ecological-systems.mp3
Category:general -- posted at: 4:10pm PDT

Fri, 14 February 2020

Adversarial Explanations

Walt Woods joins us to discuss his paper Adversarial Explanations for Understanding Image Classification Decisions and Improved Neural Network Robustness with co-authors Jack Chen and Christof Teuscher.

Direct download: adversarial-explanations.mp3
Category:general -- posted at: 3:10pm PDT

Fri, 7 February 2020

ObjectNet

Andrei Barbu joins us to discuss ObjectNet - a new kind of vision dataset.

In contrast to ImageNet, ObjectNet seeks to provide images that are more representative of the types of images an autonomous machine is likely to encounter in the real world. Collecting a dataset in this way required careful use of Mechanical Turk to get Turkers to provide a corpus of images that removes some of the bias found in ImageNet.

http://0xab.com/

Direct download: objectnet.mp3
Category:general -- posted at: 8:00am PDT

Fri, 31 January 2020

Visualization and Interpretability

Enrico Bertini joins us to discuss how data visualization can be used to help make machine learning more interpretable and explainable.

Find out more about Enrico at http://enrico.bertini.io/.

More from Enrico with co-host Moritz Stefaner on the Data Stories podcast!

Direct download: visualization-and-interpretability.mp3
Category:general -- posted at: 8:00am PDT

Sat, 25 January 2020

Interpretable One Shot Learning

We welcome Su Wang back to Data Skeptic to discuss the paper Distributional modeling on a diet: One-shot word learning from text only.

Direct download: interpretable-one-shot-learning.mp3
Category:general -- posted at: 9:00pm PDT

Wed, 22 January 2020

Fooling Computer Vision

Wiebe van Ranst joins us to talk about a project in which specially designed printed images can fool a computer vision system, preventing it from identifying a person. Their attack targets the popular YOLO2 pre-trained image recognition model, and thus, is likely to be widely applicable.

Direct download: fooling-computer-vision.mp3
Category:general -- posted at: 10:38am PDT

Mon, 13 January 2020

Algorithmic Fairness

This episode includes an interview with Aaron Roth author of The Ethical Algorithm.

Direct download: algorithmic-fairness.mp3
Category:general -- posted at: 6:31pm PDT

Tue, 7 January 2020

Interpretability

Machine learning has shown a rapid expansion into every sector and industry. With increasing reliance on models and increasing stakes for the decisions of models, questions of how models actually work are becoming increasingly important to ask.

Welcome to Data Skeptic Interpretability.

In this episode, Kyle interviews Christoph Molnar about his book Interpretable Machine Learning.

Thanks to our sponsor, the Gartner Data & Analytics Summit going on in Grapevine, TX on March 23 – 26, 2020. Use discount code: dataskeptic.

Music

Our new theme song is #5 by Big D and the Kids Table.

Incidental music by Tanuki Suit Riot.

Direct download: interpretability.mp3
Category:general -- posted at: 12:33am PDT