Data Skeptic

April 2024
S	M	T	W	T	F	S

	1	2	3	4	5	6
7	8	9	10	11	12	13
14	15	16	17	18	19	20
21	22	23	24	25	26	27
28	29	30

Thu, 26 May 2022

Today, we are joined by Alexander Thor, a Product Manager at Vizlib, makers of Astrato. Astrato is a data analytics and business intelligence tool built on the cloud and for the cloud. Alexander discusses the features and capabilities of Astrato for data professionals.

Visit our website for additional show notes!

Direct download: modern-data-stacks.mp3
Category:general -- posted at: 7:00am PDT

Mon, 23 May 2022

Emoji as a Predictor

Emojis are arguably one of the most effective ways to express emotions when texting. In today’s episode, Xuan Lu shares her research on the use of emojis by developers. She explains how the study of emojis can track the emotions of remote workers and predict future behavior. Listen to find out more!

Direct download: emoji-as-a-predictor.mp3
Category:general -- posted at: 7:25am PDT

Mon, 16 May 2022

Polarizing Trends in the Gig Economy

On the show today, Fabian Braesemann, a research fellow at the University of Oxford, joins us to discuss his study analyzing the gig economy. He revealed the trends he discovered since remote work became mainstream, the factors causing spatial polarization and some downsides of the gig economy. Listen to learn what he found.

Direct download: polarizing-trends-in-the-gig-economy.mp3
Category:general -- posted at: 6:39am PDT

Thu, 12 May 2022

Remote Learning in Applied Engineering

On the show today, we interview Mouhamed Abdulla, a professor of Electrical Engineering at Sheridan Institute of Technology. Mouhamed joins us to discuss his study on remote teaching and learning in applied engineering. He discusses how he embraced the new approach after the pandemic, the challenges he faced and how he tackled them. Listen to find out more.

Click here for additional show notes on our website!

Thanks to our sponsor!
https://neptune.ai/

Log, store, query, display, organize, and compare all your model metadata in a single place

Direct download: remote-learning-in-applied-engineering.mp3
Category:general -- posted at: 5:29am PDT

Mon, 9 May 2022

Remote Productivity

It is difficult to estimate the effect on remote working across the board. Darja Šmite, who speaks with us today, is a professor of Software Engineering at the Blekinge Institute of Technology. In her recently published paper, she analyzed data on several companies' activities before and after remote working became prevalent. She discussed the results found, why they were and some subtle drawbacks of remote working. Check it out!

Click here for additional show notes on our website!

Direct download: remote-productivity.mp3
Category:general -- posted at: 5:45am PDT

Sun, 1 May 2022

Does Remote Learning Work?

We explore this complex question in two interviews today. First, Kasey Wagoner describes 3 approaches to remote lab sessions and an analysis of which was the most instrumental to students. Second, Tahiya Chowdhury shares insights about the specific features of video-conferencing platforms that are lacking in comparison to in-person learning.

Click here for additional show notes on our website!

Thanks to our sponsor!
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

Direct download: does-remote-learning-work.mp3
Category:general -- posted at: 6:00am PDT

Mon, 25 April 2022

Covid-19 Impact on Bicycle Usage

In this episode, we speak with Abdullah Kurkcu, a Lead Traffic Modeler. Abdullah joins us to discuss his recent study on the effect of COVID-19 on bicycle usage in the US. He walks us through the data gathering process, data preprocessing, feature engineering, and model building. Abdullah also disclosed his results and key takeaways from the study. Listen to find out more.

Click here for additional show notes on our website.

Thanks to our sponsor!
Astrato is a modern BI and analytics platform built for the Snowflake Data Cloud. A next-generation live query data visualization and analytics solution, empowering everyone to make live data decisions.

Direct download: covid-19-impact-on-bicycle-usage.mp3
Category:general -- posted at: 5:46am PDT

Fri, 22 April 2022

Learning Digital Fabrication Remotely

Today, we are joined by Jennifer Jacobs and Nadya Peek, who discuss their experience in teaching remote classes for a course that is largely hands-on. The discussion was focused on digital fabrication, why it is important, the prospect for the future, the challenges with remote lectures, and everything in between.

Click here for additional show notes on our website!

Thanks to our sponsor!
https://neptune.ai/

Log, store, query, display, organize, and compare all your model metadata in a single place

Direct download: learning-digital-fabrication-remotely.mp3
Category:general -- posted at: 4:55am PDT

Mon, 18 April 2022

Remote Software Development

Today, we are joined by Denae Ford, a Senior Researcher at Microsoft Research and an Affiliate Assistant Professor at the University of Washington. Denae discusses her work around remote work and its culminating impact on workers. She narrowed down her research to how COVID-19 has affected the working system of software engineers and the emerging challenges it brings.

Click here to access additional show notes on our website!

Thanks to our sponsor!

Weights & Biases : The developer-first MLOps platform. Build better models faster with experiment tracking, dataset versioning, and model management.

Direct download: remote-software-development.mp3
Category:general -- posted at: 9:49am PDT

Mon, 11 April 2022

Quantum K-Means

In this episode, we interview Jonas Landman, a Postdoc candidate at the University of Edinburg. Jonas discusses his study around quantum learning where he attempted to recreate the conventional k-means clustering algorithm and spectral clustering algorithm using quantum computing.

Click here to access additional show notes on our website!

Direct download: quantum-k-means.mp3
Category:general -- posted at: 6:00am PDT

Mon, 4 April 2022

K-Means in Practice

K-means is widely used in real-life business problems. In this episode, Mujtaba Anwer, a researcher and Data Scientist walks us through some use cases of k-means. He also spoke extensively on how to prepare your data for clustering, find the best number of clusters to use, and turn the ‘abstract’ result into real business value. Listen to learn. Click here to access additional show notes on our website! Thanks to our sponsor!
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

Direct download: k-means-in-practice.mp3
Category:general -- posted at: 6:00am PDT

Mon, 28 March 2022

Fair Hierarchical Clustering

Building a fair machine learning model has become a critical consideration in today’s world. In this episode, we speak with Anshuman Chabra, a Ph.D. candidate in Computer Networks. Chhabra joins us to discuss his research on building fair machine learning models and why it is important. Find out how he modeled the problem and the result found.

Click here to access additional show notes on our webiste!

Thanks to our sponsor!
https://astrato.io

Astrato is a modern BI and analytics platform built for the Snowflake Data Cloud. A next-generation live query data visualization and analytics solution, empowering everyone to make live data decisions.

Direct download: fair-hierarchical-clustering.mp3
Category:general -- posted at: 6:23am PDT

Mon, 21 March 2022

Matrix Factorization For k-Means

Many people know K-means clustering as a powerful clustering technique but not all listeners will be as familiar with spectral clustering. In today’s episode, Sibylle Hess from the Data Mining group at TU Eindhoven joins us to discuss her work around spectral clustering and how its result could potentially cause a massive shift from the conventional neural networks. Listen to learn about her findings.

Visit our website for additional show notes

Thanks to our sponsor, Weights & Biases

Direct download: matrix-factorization-for-k-means.mp3
Category:general -- posted at: 6:00am PDT

Mon, 14 March 2022

Breathing K-Means

In this episode, we speak with Bernd Fritzke, a proficient financial expert and a Data Science researcher on his recent research - the breathing K-means algorithm. Bernd discussed the perks of the algorithms and what makes it stand out from other K-means variations. He extensively discussed the working principle of the algorithm and the subtle but impactful features that enables it produce top-notch results with low computational resources. Listen to learn about this algorithm.

Direct download: breathing-k-means.mp3
Category:general -- posted at: 6:00am PDT

Mon, 7 March 2022

Power K-Means

In today’s episode, Jason, an Assistant Professor of Statistical Science at Duke University talks about his research on K power means. K power means is a newly-developed algorithm by Jason and his team, that aims to solve the problem of local minima in classical K-means, without demanding heavy computational resources. Listen to find out the outcome of Jason's study.

Click here to access additional show notes on our website!

Thanks to our Sponsors:
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale. https://clear.ml

Springboard
Springboard offers end-to-end online data career programs that encompass data science, data analytics, data engineering, and machine learning engineering.

Direct download: power-k-means.mp3
Category:general -- posted at: 6:00am PDT

Thu, 3 March 2022

Explainable K-Means

In this episode, Kyle interviews Lucas Murtinho about the paper "Shallow decision treees for explainable k-means clustering" about the use of decision trees to help explain the clustering partitions.

Check out our website for extended show notes!

Thanks to our Sponsors:
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

Direct download: explainable-k-means.mp3
Category:general -- posted at: 6:17am PDT

Mon, 28 February 2022

Customer Clustering

Have you ever wondered how you can use clustering to extract meaningful insight from a time-series single-feature data? In today’s episode, Ehsan speaks about his recent research on actionable feature extraction using clustering techniques. Want to find out more? Listen to discover the methodologies he used for his research and the commensurate results.

Visit our website for extended show notes!

https://clear.ml/

ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

Direct download: customer-clustering.mp3
Category:general -- posted at: 6:00am PDT

Mon, 21 February 2022

k-means Image Segmentation

Linh Da joins us to explore how image segmentation can be done using k-means clustering. Image segmentation involves dividing an image into a distinct set of segments. One such approach is to do this purely on color, in which case, k-means clustering is a good option.

Check out our website for extended show notes and images!

Thanks to our Sponsors:

Visit Weights and Biases mention Data Skeptic when you request a demo!

&
Nomad Data

In the image below, you can see the k-means clustering segmentation results for the same image with the values of 2, 4, 6, and 8 for k.

Lilac Crowned Amazon

Direct download: k-means-image-segmentation.mp3
Category:general -- posted at: 4:00pm PDT

Fri, 18 February 2022

Tracking Elephant Clusters

In today’s episode, Gregory Glatzer explained his machine learning project that involved the prediction of elephant movement and settlement, in a bid to limit the activities of poachers. He used two machine learning algorithms, DBSCAN and K-Means clustering at different stages of the project. Listen to learn about why these two techniques were useful and what conclusions could be drawn.

Click here to see additional show notes on our website!

Thanks to our sponsor, Astrato

Direct download: tracking-elephant-clusters.mp3
Category:general -- posted at: 2:43pm PDT

Mon, 14 February 2022

k-means clustering

Welcome to our new season, Data Skeptic: k-means clustering. Each week will feature an interview or discussion related to this classic algorithm, it's use cases, and analysis.

This episode is an overview of the topic presented in several segments.

Direct download: k-means-clustering.mp3
Category:general -- posted at: 8:44am PDT

Mon, 7 February 2022

Snowflake Essentials

Frank Bell, Snowflake Data Superhero, and SnowPro, joins us today to talk about his book “Snowflake Essentials: Getting Started with Big Data in the Cloud.”

Snowflake Essentials: Getting Started with Big Data in the Cloud by Frank Bell, Raj Chirumamilla, Bhaskar B. Joshi, Bjorn Lindstrom, Ruchi Soni, Sameer Videkar
Snowflake Solutions
Snoptimizer - Snowflake Cost, Security, and Performance Optimization - Coming Soon!

Thanks to our Sponsors:

Find Better Data Faster with Nomad Data. Visit nomad-data.com
Visit Springboard and use promo code DATASKEPTIC to receive a $750 discount

Direct download: snowflake-essentials.mp3
Category:general -- posted at: 6:00am PDT

Mon, 31 January 2022

Explainable Climate Science

Zack Labe, a Post-Doctoral Researcher at Colorado State University, joins us today to discuss his work “Detecting Climate Signals using Explainable AI with Single Forcing Large Ensembles.”
Works Mentioned
“Detecting Climate Signals using Explainable AI with Single Forcing Large Ensembles”
by Zachary M. Labe, Elizabeth A. Barnes

Sponsored by:
Astrato
and
BBEdit by Bare Bones Software

Direct download: explainable-climate-science.mp3
Category:general -- posted at: 8:24am PDT

Mon, 24 January 2022

Energy Forecasting Pipelines

Erin Boyle, the Head of Data Science at Myst AI, joins us today to talk about her work with Myst AI, a time series forecasting platform and service with the objective for positively impacting sustainability.

https://docs.myst.ai/docs Visit Weights and Biases at wandb.me/dataskeptic Find Better Data Faster with Nomad Data. Visit nomad-data.com

Direct download: energy-forecasting-pipelines.mp3
Category:general -- posted at: 6:00am PDT

Mon, 17 January 2022

Matrix Profiles in Stumpy

Sean Law, Principle Data Scientist, R&D at a Fortune 500 Company, comes on to talk about his creation of the STUMPY Python Library.

Sponsored by Hello Fresh and mParticle:

Go to Hellofresh.com/dataskeptic16 for up to 16 free meals AND 3 free gifts!

Visit mparticle.com to learn how teams at Postmates, NBCUniversal, Spotify, and Airbnb use mParticle’s customer data infrastructure to accelerate their customer data strategies.

Direct download: matrix-profiles-in-stumpy.mp3
Category:general -- posted at: 6:00am PDT

Thu, 13 January 2022

The Great Australian Prediction Project

Data scientists and psychics have at least one major thing in common. Both professions attempt to predict the future. In the case of a data scientist, this is done using algorithms, data, and often comes with some measure of quality such as a confidence interval or estimated accuracy. In contrast, psychics rely on their intuition or an appeal to the supernatural as the source for their predictions. Still, in the interest of empirical evidence, the quality of predictions made by psychics can be put to the test.

The Great Australian Psychic Prediction Project seeks to do exactly that. It's the longest known project tracking annual predictions made by psychics, and the accuracy of those predictions in hindsight. Richard Saunders, host of The Skeptic Zone Podcast, joins us to share the results of this decadal study.

Read the full report: https://www.skeptics.com.au/2021/12/09/psychic-project-full-results-released/

And follow the Skeptics Zone: https://www.skepticzone.tv/

Direct download: the-great-australian-prediction-project.mp3
Category:general -- posted at: 6:30pm PDT

Mon, 10 January 2022

Water Demand Forecasting

Georgia Papacharalampous, Researcher at the National Technical University of Athens, joins us today to talk about her work “Probabilistic water demand forecasting using quantile regression algorithms.”

Visit Springboard and use promo code DATASKEPTIC to receive a $750 discount

Direct download: water-demand-forecasting.mp3
Category:general -- posted at: 9:30am PDT

Mon, 3 January 2022

Open Telemetry

John Watson, Principal Software Engineer at Splunk, joins us today to talk about Splunk and OpenTelemetry.

Direct download: open-telemetry.mp3
Category:general -- posted at: 6:00am PDT

Mon, 27 December 2021

Fashion Predictions

Yusan Lin, a Research Scientist at Visa Research, comes on today to talk about her work "Predicting Next-Season Designs on High Fashion Runway."

Direct download: fashion-predictions.mp3
Category:general -- posted at: 6:00am PDT

Sat, 25 December 2021

Time Series Mini Episodes

Time series topics on Data Skeptic predate our current season. This holiday special collects three popular mini-episodes from the archive that discuss time series topics with a few new comments from Kyle.

Direct download: time-series-mini-episodes.mp3
Category:general -- posted at: 12:02am PDT

Mon, 20 December 2021

Forecasting Motor Vehicle Collision

Dr. Darren Shannon, a Lecturer in Quantitative Finance in the Department of Accounting and Finance, University of Limerick, joins us today to talk about his work "Extending the Heston Model to Forecast Motor Vehicle Collision Rates."

Direct download: forecasting-motor-vehicle-collision-rates.mp3
Category:general -- posted at: 6:00am PDT

Mon, 13 December 2021

Deep Learning for Road Traffic Forecasting

Eric Manibardo, PhD Student at the University of the Basque Country in Spain, comes on today to share his work, "Deep Learning for Road Traffic Forecasting: Does it Make a Difference?"

Direct download: deep-learning-for-road-traffic-forecasting.mp3
Category:general -- posted at: 6:00am PDT

Mon, 6 December 2021

Bike Share Demand Forecasting

Daniele Gammelli, PhD Student in Machine Learning at Technical University of Denmark and visiting PhD Student at Stanford University, joins us today to talk about his work "Predictive and Prescriptive Performance of Bike-Sharing Demand Forecasts for Inventory Management."

Direct download: bike-share-demand-forecasting.mp3
Category:general -- posted at: 6:00am PDT

Mon, 29 November 2021

Forecasting in Supply Chain

Mahdi Abolghasemi, Lecturer at Monash University, joins us today to talk about his work "Demand forecasting in supply chain: The impact of demand volatility in the presence of promotion."

Direct download: forecasting-in-supply-chain.mp3
Category:general -- posted at: 6:00am PDT

Fri, 26 November 2021

Black Friday

The retail holiday “black Friday” occurs the day after Thanksgiving in the United States. It’s dubbed this because many retail companies spend the first 10 months of the year running at a loss (in the red) before finally earning as much as 80% of their revenue in the last two months of the year.

This episode features four interviews with guests bringing unique data-driven perspectives on the topic of analyzing this seeming outlier in a time series dataset.

Direct download: black-friday.mp3
Category:general -- posted at: 7:24am PDT

Mon, 22 November 2021

Aligning Time Series on Incomparable Spaces

Alex Terenin, Postdoctoral Research Associate at the University of Cambridge, joins us today to talk about his work "Aligning Time Series on Incomparable Spaces."

Direct download: aligning-time-series-on-incomparable-spaces.mp3
Category:general -- posted at: 6:00am PDT

Mon, 15 November 2021

Comparing Time Series with HCTSA

Today we are joined again by Ben Fulcher, leader of the Dynamics and Neural Systems Group at the University of Sydney in Australia, to talk about hctsa, a software package for running highly comparative time-series analysis.

Direct download: comparing-time-series-with-hctsa.mp3
Category:general -- posted at: 6:01am PDT

Mon, 8 November 2021

Change Point Detection Algorithms

Gerrit van den Burg, Postdoctoral Researcher at The Alan Turing Institute, joins us today to discuss his work "An Evaluation of Change Point Detection Algorithms."

Direct download: change-point-detection-algorithms.mp3
Category:general -- posted at: 6:14am PDT

Mon, 1 November 2021

Time Series for Good

Bahman Rostami-Tabar, Senior Lecturer in Management Science at Cardiff University, joins us today to talk about his work "Forecasting and its Beneficiaries."

Direct download: time-series-for-good.mp3
Category:general -- posted at: 6:00am PDT

Mon, 25 October 2021

Long Term Time Series Forecasting

Alex Mallen, Computer Science student at the University of Washington, and Henning Lange, a Postdoctoral Scholar in Applied Math at the University of Washington, join us today to share their work "Deep Probabilistic Koopman: Long-term Time-Series Forecasting Under Periodic Uncertainties."

Direct download: long-term-time-series-forecasting.mp3
Category:general -- posted at: 6:00am PDT

Sun, 17 October 2021

Fast and Frugal Time Series Forecasting

Fotios Petropoulos, Professor of Management Science at the University of Bath in The U.K., joins us today to talk about his work "Fast and Frugal Time Series Forecasting."

Direct download: fast-and-frugal-time-series-forecasting.mp3
Category:general -- posted at: 1:13pm PDT

Mon, 11 October 2021

Causal Inference in Educational Systems

Manie Tadayon, a PhD graduate from the ECE department at University of California, Los Angeles, joins us today to talk about his work “Comparative Analysis of the Hidden Markov Model and LSTM: A Simulative Approach.”

Direct download: causal-inference-in-educational-systems.mp3
Category:general -- posted at: 6:00am PDT

Mon, 4 October 2021

Boosted Embeddings for Time Series

Sankeerth Rao Karingula, ML Researcher at Palo Alto Networks, joins us today to talk about his work “Boosted Embeddings for Time Series Forecasting.”

Works Mentioned
Boosted Embeddings for Time Series Forecasting
by Sankeerth Rao Karingula, Nandini Ramanan, Rasool Tahmasbi, Mehrnaz Amjadi, Deokwoo Jung, Ricky Si, Charanraj Thimmisetty, Luisa Polania Cabrera, Marjorie Sayer, Claudionor Nunes Coelho Jr

https://www.linkedin.com/in/sankeerthrao/

https://twitter.com/sankeerthrao3

https://lod2021.icas.cc/

Direct download: boosted-embeddings-for-time-series.mp3
Category:general -- posted at: 6:00am PDT

Mon, 27 September 2021

Change Point Detection in Continuous Integration Systems

David Daly, Performance Engineer at MongoDB, joins us today to discuss "The Use of Change Point Detection to Identify Software Performance Regressions in a Continuous Integration System".

Works Mentioned
The Use of Change Point Detection to Identify Software Performance Regressions in a Continuous Integration System
by David Daly, William Brown, Henrik Ingo, Jim O’Leary, David BradfordSocial Media

David's Website
David's Twitter
Mongodb

Direct download: change-point-detection-in-continuous-integration-systems.mp3
Category:general -- posted at: 6:00am PDT

Mon, 20 September 2021

Applying k-Nearest Neighbors to Time Series

Samya Tajmouati, a PhD student in Data Science at the University of Science of Kenitra, Morocco, joins us today to discuss her work Applying K-Nearest Neighbors to Time Series Forecasting: Two New Approaches.

Direct download: applying-k-nearest-neighbors-to-time-series.mp3
Category:general -- posted at: 6:00am PDT

Mon, 13 September 2021

Ultra Long Time Series

Dr. Feng Li, (@f3ngli) is an Associate Professor of Statistics in the School of Statistics and Mathematics at Central University of Finance and Economics in Beijing, China. He joins us today to discuss his work Distributed ARIMA Models for Ultra-long Time Series.

Direct download: ultra-long-time-series.mp3
Category:general -- posted at: 6:00am PDT

Mon, 6 September 2021

MiniRocket

Angus Dempster, PhD Student at Monash University in Australia, comes on today to talk about MINIROCKET: A Very Fast (Almost) Deterministic Transform for Time Series Classification, a fast deterministic transform for time series classification. MINIROCKET reformulates ROCKET, gaining a 75x improvement on larger datasets with essentially the same performance. In this episode, we talk about the insights that realized this speedup as well as use cases.

Direct download: minirocket.mp3
Category:general -- posted at: 6:00am PDT

Mon, 30 August 2021

ARiMA is not Sufficient

Chongshou Li, Associate Professor at Southwest Jiaotong University in China, joins us today to talk about his work Why are the ARIMA and SARIMA not Sufficient.

Direct download: arima-is-not-sufficient.mp3
Category:general -- posted at: 6:00am PDT

Mon, 23 August 2021

Comp Engine

Ben Fulcher, Senior Lecturer at the School of Physics at the University of Sydney in Australia, comes on today to talk about his project Comp Engine.

Follow Ben on Twitter: @bendfulcher
For posts about time series analysis : @comptimeseries
comp-engine.org

Direct download: comp-engine.mp3
Category:general -- posted at: 6:00am PDT

Mon, 16 August 2021

Detecting Ransomware

Nitin Pundir, PhD candidate at University Florida and works at the Florida Institute for Cybersecurity Research, comes on today to talk about his work “RanStop: A Hardware-assisted Runtime Crypto-Ransomware Detection Technique.”

FICS Research Lab - https://fics.institute.ufl.edu/

LinkedIn - https://www.linkedin.com/in/nitin-pundir470/

Direct download: detecting-ransomware.mp3
Category:general -- posted at: 6:00am PDT

Mon, 9 August 2021

GANs in Finance

Florian Eckerli, a recent graduate of Zurich University of Applied Sciences, comes on the show today to discuss his work Generative Adversarial Networks in Finance: An Overview.

Direct download: gans-in-finance.mp3
Category:general -- posted at: 6:00am PDT

Mon, 2 August 2021

Predicting Urban Land Use

Today on the show we have Daniel Omeiza, a doctoral student in the computer science department of the University of Oxford, who joins us to talk about his work Efficient Machine Learning for Large-Scale Urban Land-Use Forecasting in Sub-Saharan Africa.

Direct download: predicting-urban-land-use.mp3
Category:general -- posted at: 6:00am PDT

Mon, 26 July 2021

Opportunities for Skillful Weather Prediction

Today on the show we have Elizabeth Barnes, Associate Professor in the department of Atmospheric Science at Colorado State University, who joins us to talk about her work Identifying Opportunities for Skillful Weather Prediction with Interpretable Neural Networks. Find more from the Barnes Research Group on their site.

Weather is notoriously difficult to predict. Complex systems are demanding of computational power. Further, the chaotic nature of, well, nature, makes accurate forecasting especially difficult the longer into the future one wants to look. Yet all is not lost!

In this interview, we explore the use of machine learning to help identify certain conditions under which the weather system has entered an unusually predictable position in it’s normally chaotic state space.

Direct download: opportunities-for-skillful-weather-prediction.mp3
Category:general -- posted at: 6:00am PDT

Mon, 19 July 2021

Predicting Stock Prices

Today on the show we have Andrea Fronzetti Colladon (@iandreafc), currently working at the University of Perugia and inventor of the Semantic Brand Score, joins us to talk about his work studying human communication and social interaction.

We discuss the paper Look inside. Predicting Stock Prices by Analyzing an Enterprise Intranet Social Network and Using Word Co-Occurrence Networks.

Direct download: predicting-stock-prices.mp3
Category:general -- posted at: 6:00am PDT

Mon, 12 July 2021

N-Beats

Today on the show we have Boris Oreshkin @boreshkin, a Senior Research Scientist at Unity Technologies, who joins us today to talk about his work N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting.

Works Mentioned:
N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting
By Boris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua Bengio
https://arxiv.org/abs/1905.10437

Social Media
Linkedin

Twitter

Direct download: nbeats.mp3
Category:general -- posted at: 8:04am PDT

Mon, 5 July 2021

Translation Automation

Today we are back with another episode discussing AI in the work field. AI has, is, and will continue to facilitate the automation of work done by humans. Sometimes this may be an entire role. Other times it may automate a particular part of their role, scaling their effectiveness.

Carl Stimson, a Freelance Japanese to English translator, comes on the show to talk about his work in translation and his perspective about how AI will change translation in the future.

Direct download: translation-automation.mp3
Category:general -- posted at: 6:48pm PDT

Mon, 28 June 2021

Time Series at the Beach

Shane Ross, Professor of Aerospace and Ocean Engineering at Virginia Tech University, comes on today to talk about his work “Beach-level 24-hour forecasts of Florida red tide-induced respiratory irritation.”

Direct download: time-series-at-the-beach.mp3
Category:general -- posted at: 6:00am PDT

Mon, 21 June 2021

Automatic Identification of Outlier Galaxy Images

Lior Shamir, Associate Professor of Computer Science at Kansas University, joins us today to talk about the recent paper Automatic Identification of Outliers in Hubble Space Telescope Galaxy Images.

Follow Lio on Twitter @shamir_lior

Direct download: automatic-identification-of-outlier-galaxy-images.mp3
Category:general -- posted at: 12:11pm PDT

Wed, 16 June 2021

Do We Need Deep Learning in Time Series

Shereen Elsayed and Daniela Thyssens, both are PhD Student at Hildesheim University in Germany, come on today to talk about the work “Do We Really Need Deep Learning Models for Time Series Forecasting?”

Direct download: do-we-need-deep-learning-in-time-series.mp3
Category:general -- posted at: 9:10am PDT

Thu, 10 June 2021

Detecting Drift

Sam Ackerman, Research Data Scientist at IBM Research Labs in Haifa, Israel, joins us today to talk about his work Detection of Data Drift and Outliers Affecting Machine Learning Model Performance Over Time.

Check out Sam's IBM statistics/ML blog at: http://www.research.ibm.com/haifa/dept/vst/ML-QA.shtml

Direct download: detecting-drift.mp3
Category:general -- posted at: 5:05pm PDT

Mon, 31 May 2021

Darts Library for Time Series

Julien Herzen, PhD graduate from EPFL in Switzerland, comes on today to talk about his work with Unit 8 and the development of the Python Library: Darts.

Direct download: darts-library-for-time-series.mp3
Category:general -- posted at: 7:52am PDT

Mon, 24 May 2021

Forecasting Principles and Practice

Welcome to Timeseries! Today’s episode is an interview with Rob Hyndman, Professor of Statistics at Monash University in Australia, and author of Forecasting: Principles and Practices.

Direct download: forecasting-principles-and-practice.mp3
Category:general -- posted at: 7:57am PDT

Fri, 21 May 2021

Prequisites for Time Series

Today's experimental episode uses sound to describe some basic ideas from time series.

This episode includes lag, seasonality, trend, noise, heteroskedasticity, decomposition, smoothing, feature engineering, and deep learning.

Direct download: prequisites-for-time-series.mp3
Category:general -- posted at: 11:36am PDT

Fri, 7 May 2021

Orders of Magnitude

Today’s show in two parts. First, Linhda joins us to review the episodes from Data Skeptic: Pilot Season and give her feedback on each of the topics.

Second, we introduce our new segment “Orders of Magnitude”. It’s a statistical game show in which participants must identify the true statistic hidden in a list of statistics which are off by at least an order of magnitude. Claudia and Vanessa join as our first contestants. Below are the sources of our questions.

Heights

Bird Statistics

Amounts of Data

Our statistics come from this post

Direct download: oom.mp3
Category:general -- posted at: 11:55am PDT

Mon, 3 May 2021

They're Coming for Our Jobs

AI has, is, and will continue to facilitate the automation of work done by humans. Sometimes this may be an entire role. Other times it may automate a particular part of their role, scaling their effectiveness. Unless progress in AI inexplicably halts, the tasks done by humans vs. machines will continue to evolve. Today’s episode is a speculative conversation about what the future may hold.

Co-Host of Squaring the Strange Podcast, Caricature Artist, and an Academic Editor, Celestia Ward joins us today! Kyle and Celestia discuss whether or not her jobs as a caricature artist or as an academic editor are under threat from AI automation.

Mentions

https://squaringthestrange.wordpress.com/
https://twitter.com/celestiaward
The legendary Dr. Jorge Pérez and his work studying unicorns
Supernormal stimulus
International Society of Caricature Artists
Two Heads Studios

Direct download: theyre-coming-for-our-jobs.mp3
Category:general -- posted at: 9:00am PDT

Mon, 26 April 2021

Pandemic Machine Learning Pitfalls

Today on the show Derek Driggs, a PhD Student at the University of Cambridge. He comes on to discuss the work Common Pitfalls and Recommendations for Using Machine Learning to Detect and Prognosticate for COVID-19 Using Chest Radiographs and CT Scans.

Help us vote for the next theme of Data Skeptic!

Vote here: https://dataskeptic.com/vote

Direct download: pandemic-machine-learning-pitfalls.mp3
Category:general -- posted at: 12:00am PDT

Mon, 19 April 2021

Flesch Kincaid Readability Tests

Given a document in English, how can you estimate the ease with which someone will find they can read it? Does it require a college-level of reading comprehension or is it something a much younger student could read and understand?

While these questions are useful to ask, they don't admit a simple answer. One option is to use one of the (essentially identical) two Flesch Kincaid Readability Tests. These are simple calculations which provide you with a rough estimate of the reading ease.

In this episode, Kyle shares his thoughts on this tool and when it could be appropriate to use as part of your feature engineering pipeline towards a machine learning objective.

For empirical validation of these metrics, the plot below compares English language Wikipedia pages with "Simple English" Wikipedia pages. The analysis Kyle describes in this episode yields the intuitively pleasing histogram below. It summarizes the distribution of Flesch reading ease scores for 1000 pages examined from both Wikipedias.

Direct download: flesch-kincaid-readability-tests.mp3
Category:general -- posted at: 12:50am PDT

Fri, 9 April 2021

Fairness Aware Outlier Detection

Today on the show we have Shubhranshu Shekar, a Ph. D Student at Carnegie Mellon University, who joins us to talk about his work, FAIROD: Fairness-aware Outlier Detection.

Direct download: fairness-aware-outlier-detection.mp3
Category:general -- posted at: 8:30am PDT

Mon, 5 April 2021

Life May be Rare

Today on the show Dr. Anders Sandburg, Senior Research Fellow at the Future of Humanity Institute at Oxford University, comes on to share his work “The Timing of Evolutionary Transitions Suggest Intelligent Life is Rare.”

Works Mentioned:

Paper:
“The Timing of Evolutionary Transitions Suggest Intelligent Life is Rare.”by Andrew E Snyder-Beattie, Anders Sandberg, K Eric Drexler, Michael B Bonsall

Twitter:
@anderssandburg

Direct download: life-may-be-rare.mp3
Category:general -- posted at: 7:24am PDT

Mon, 29 March 2021

Social Networks

Mayank Kejriwal, Research Professor at the University of Southern California and Researcher at the Information Sciences Institute, joins us today to discuss his work and his new book Knowledge, Graphs, Fundamentals, Techniques and Applications by Mayank Kejriwal, Craig A. Knoblock, and Pedro Szekley.

Works Mentioned
“Knowledge, Graphs, Fundamentals, Techniques and Applications”by Mayank Kejriwal, Craig A. Knoblock, and Pedro Szekley

Direct download: social-networks.mp3
Category:general -- posted at: 7:21am PDT

Mon, 22 March 2021

The QAnon Conspiracy

QAnon is a conspiracy theory born in the underbelly of the internet. While easy to disprove, these cryptic ideas captured the minds of many people and (in part) paved the way to the 2021 storming of the US Capital.

This is a contemporary conspiracy which came into existence and grew in a very digital way. This makes it possible for researchers to study this phenomenon in a way not accessible in previous conspiracy theories of similar popularity.

This episode is not so much a debunking of this debunked theory, but rather an exploration of the metadata and origins of this conspiracy.

This episode is also the first in our 2021 Pilot Season in which we are going to test out a few formats for Data Skeptic to see what our next season should be. This is the first installment. In a few weeks, we're going to ask everyone to vote for their favorite theme for our next season.

Direct download: the-qanon-conspiracy.mp3
Category:general -- posted at: 7:39am PDT

Mon, 15 March 2021

Benchmarking Vision on Edge vs Cloud

Karthick Shankar, Masters Student at Carnegie Mellon University, and Somali Chaterji, Assistant Professor at Purdue University, join us today to discuss the paper "JANUS: Benchmarking Commercial and Open-Source Cloud and Edge Platforms for Object and Anomaly Detection Workloads"

Works Mentioned:

https://ieeexplore.ieee.org/abstract/document/9284314
“JANUS: Benchmarking Commercial and Open-Source Cloud and Edge Platforms for Object and Anomaly Detection Workloads.”

by: Karthick Shankar, Pengcheng Wang, Ran Xu, Ashraf Mahgoub, Somali ChaterjiSocial Media

Karthick Shankar
https://twitter.com/karthick_sh

Somali Chaterji
https://twitter.com/somalichaterji?lang=en
https://schaterji.io/

Direct download: benchmarking-vision-on-edge-vs-cloud.mp3
Category:general -- posted at: 5:00am PDT

Fri, 5 March 2021

Goodhart's Law in Reinforcement Learning

Hal Ashton, a PhD student from the University College of London, joins us today to discuss a recent work Causal Campbell-Goodhart’s law and Reinforcement Learning.

"Only buy honey from a local producer." - Hal Ashton

Works Mentioned:

“Causal Campbell-Goodhart’s law and Reinforcement Learning”by Hal AshtonBook

“The Book of Why”by Judea PearlPaper

Thanks to our sponsor!

When your business is ready to make that next hire, find the right person with LinkedIn Jobs. Just visit LinkedIn.com/DATASKEPTIC to post a job for

free! Terms and conditions apply

Direct download: goodharts-law-in-reinforcement-learning.mp3
Category:general -- posted at: 5:00am PDT

Mon, 1 March 2021

Video Anomaly Detection

Yuqi Ouyang, in his second year of PhD study at the University of Warwick in England, joins us today to discuss his work “Video Anomaly Detection by Estimating Likelihood of Representations.”Works Mentioned:

Video Anomaly Detection by Estimating Likelihood of Representations
https://arxiv.org/abs/2012.01468
by: Yuqi Ouyang, Victor Sanchez

Direct download: video-anomaly-detection.mp3
Category:general -- posted at: 6:00am PDT

Mon, 22 February 2021

Fault Tolerant Distributed Gradient Descent

Nirupam Gupta, a Computer Science Post Doctoral Researcher at EDFL University in Switzerland, joins us today to discuss his work “Byzantine Fault-Tolerance in Peer-to-Peer Distributed Gradient-Descent.”

Works Mentioned:
https://arxiv.org/abs/2101.12316

Byzantine Fault-Tolerance in Peer-to-Peer Distributed Gradient-Descent
by Nirupam Gupta and Nitin H. Vaidya

Conference Details:

https://georgetown.zoom.us/meeting/register/tJ0sc-2grDwjEtfnLI0zPnN-GwkDvJdaOxXF

Direct download: fault-tolerant-distributed-gradient-descent.mp3
Category:general -- posted at: 6:30am PDT

Mon, 15 February 2021

Decentralized Information Gathering

Mikko Lauri, Post Doctoral researcher at the University of Hamburg, Germany, comes on the show today to discuss the work Information Gathering in Decentralized POMDPs by Policy Graph Improvements.

Follow Mikko: @mikko_lauri

Github https://laurimi.github.io/

Direct download: decentralized-information-gathering.mp3
Category:general -- posted at: 5:30am PDT

Fri, 5 February 2021

Leaderless Consensus

Balaji Arun, a PhD Student in the Systems of Software Research Group at Virginia Tech, joins us today to discuss his research of distributed systems through the paper “Taming the Contention in Consensus-based Distributed Systems.”

Works Mentioned
“Taming the Contention in Consensus-based Distributed Systems”
by Balaji Arun, Sebastiano Peluso, Roberto Palmieri, Giuliano Losa, and Binoy Ravindran
https://www.ssrg.ece.vt.edu/papers/tdsc20-author-version.pdf

“Fast Paxos”
by Leslie Lamport
https://link.springer.com/article/10.1007/s00446-006-0005-x

Direct download: leaderless-consensus.mp3
Category:general -- posted at: 9:47am PDT

Fri, 29 January 2021

Automatic Summarization

Maartje ter Hoeve, PhD Student at the University of Amsterdam, joins us today to discuss her research in automated summarization through the paper “What Makes a Good Summary? Reconsidering the Focus of Automatic Summarization.”

Works Mentioned
“What Makes a Good Summary? Reconsidering the Focus of Automatic Summarization.”
by Maartje der Hoeve, Juilia Kiseleva, and Maarten de Rijke

Contact
Email:
m.a.terhoeve@uva.nl

Twitter:
https://twitter.com/maartjeterhoeve

Website:
https://maartjeth.github.io/#get-in-touch

Direct download: automatic-summarization.mp3
Category:general -- posted at: 8:00am PDT

Fri, 22 January 2021

Gerrymandering

Brian Brubach, Assistant Professor in the Computer Science Department at Wellesley College, joins us today to discuss his work “Meddling Metrics: the Effects of Measuring and Constraining Partisan Gerrymandering on Voter Incentives".

WORKS MENTIONED:
Meddling Metrics: the Effects of Measuring and Constraining Partisan Gerrymandering on Voter Incentives
by Brian Brubach, Aravind Srinivasan, and Shawn Zhao

Direct download: gerrymandering.mp3
Category:general -- posted at: 8:00am PDT

Fri, 15 January 2021

Even Cooperative Chess is Hard

Aside from victory questions like “can black force a checkmate on white in 5 moves?” many novel questions can be asked about a game of chess. Some questions are trivial (e.g. “How many pieces does white have?") while more computationally challenging questions can contribute interesting results in computational complexity theory.

In this episode, Josh Brunner, Master's student in Theoretical Computer Science at MIT, joins us to discuss his recent paper Complexity of Retrograde and Helpmate Chess Problems: Even Cooperative Chess is Hard.

Works Mentioned
Complexity of Retrograde and Helpmate Chess Problems: Even Cooperative Chess is Hard
by Josh Brunner, Erik D. Demaine, Dylan Hendrickson, and Juilian Wellman

1x1 Rush Hour With Fixed Blocks is PSPACE Complete
by Josh Brunner, Lily Chung, Erik D. Demaine, Dylan Hendrickson, Adam Hesterberg, Adam Suhl, Avi Zeff

Direct download: even-cooperative-chess-is-hard.mp3
Category:general -- posted at: 10:02am PDT

Mon, 11 January 2021

Consecutive Votes in Paxos

Eil Goldweber, a graduate student at the University of Michigan, comes on today to share his work in applying formal verification to systems and a modification to the Paxos protocol discussed in the paper Significance on Consecutive Ballots in Paxos.

Works Mentioned :
Previous Episode on Paxos
https://dataskeptic.com/blog/episodes/2020/distributed-consensus

Paper:
On the Significance on Consecutive Ballots in Paxos by: Eli Goldweber, Nuda Zhang, and Manos Kapritsos

Thanks to our sponsor:
Nord VPN : 68% off a 2-year plan and one month free! With NordVPN, all the data you send and receive online travels through an encrypted tunnel. This way, no one can get their hands on your private information. Nord VPN is quick and easy to use to protect the privacy and security of your data. Check them out at nordvpn.com/dataskeptic

Direct download: consecutive-votes-in-paxos.mp3
Category:general -- posted at: 6:00am PDT

Fri, 1 January 2021

Visual Illusions Deceiving Neural Networks

Today on the show we have Adrian Martin, a Post-doctoral researcher from the University of Pompeu Fabra in Barcelona, Spain. He comes on the show today to discuss his research from the paper “Convolutional Neural Networks can be Deceived by Visual Illusions.”

Works Mentioned in Paper:
“Convolutional Neural Networks can be Decieved by Visual Illusions.” by Alexander Gomez-Villa, Adrian Martin, Javier Vazquez-Corral, and Marcelo Bertalmio

Examples:

Snake Illusions
https://www.illusionsindex.org/i/rotating-snakes

Twitter:
Alex: @alviur

Adrian: @adriMartin13

Thanks to our sponsor!

Keep your home internet connection safe with Nord VPN! Get 68% off plus a free month at nordvpn.com/dataskeptic (30-day money-back guarantee!)

Direct download: visual-illusions-deceiving-neural-networks.mp3
Category:general -- posted at: 6:00am PDT

Fri, 25 December 2020

Earthquake Detection with Crowd-sourced Data

Have you ever wanted to hear what an earthquake sounds like? Today on the show we have Omkar Ranadive, Computer Science Masters student at NorthWestern University, who collaborates with Suzan van der Lee, an Earth and Planetary Sciences professor at Northwestern University, on the crowd-sourcing project Earthquake Detective.

Email Links:
Suzan: suzan@earth.northwestern.edu
Omkar: omkar.ranadive@u.northwestern.edu

Works Mentioned:

Paper: Applying Machine Learning to Crowd-sourced Data from Earthquake Detective
https://arxiv.org/abs/2011.04740
by Omkar Ranadive, Suzan van der Lee, Vivan Tang, and Kevin Chao
Github: https://github.com/Omkar-Ranadive/Earthquake-Detective
Earthquake Detective: https://www.zooniverse.org/projects/vivitang/earthquake-detective

Thanks to our sponsors!

Brilliant.org Is an awesome platform with interesting courses, like Quantum Computing! There is something for you and surely something for the whole family! Get 20% off Brilliant Premium at http://brilliant.com/dataskeptic

Direct download: earthquake-detection-with-crowd-sourced-data.mp3
Category:general -- posted at: 8:21am PDT

Tue, 22 December 2020

Byzantine Fault Tolerant Consensus

Byzantine fault tolerance (BFT) is a desirable property in a distributed computing environment. BFT means the system can survive the loss of nodes and nodes becoming unreliable. There are many different protocols for achieving BFT, though not all options can scale to large network sizes.

Ted Yin joins us to explain BFT, survey the wide variety of protocols, and share details about HotStuff.

Direct download: byzantine-fault-tolerant-consensus.mp3
Category:general -- posted at: 5:00am PDT

Fri, 11 December 2020

Alpha Fold

Kyle shared some initial reactions to the announcement about Alpha Fold 2's celebrated performance in the CASP14 prediction. By many accounts, this exciting result means protein folding is now a solved problem.

Thanks to our sponsors!

Brilliant is a great last-minute gift idea! Give access to 60 + interactive courses including Quantum Computing and Group Theory. There's something for everyone at Brilliant. They have award-winning courses, taught by teachers, researchers and professionals from MIT, Caltech, Duke, Microsoft, Google and many more. Check them out at brilliant.org/dataskeptic to take advantage of 20% off a Premium memebership.
Betterhelp is an online professional counseling platform. Start communicating with a licensed professional in under 24 hours! It's safe, private and convenient. From online messages to phone and video calls, there is something for everyone. Get 10% off your first month at betterhelp.com/dataskeptic

Direct download: alpha-fold.mp3
Category:general -- posted at: 9:45am PDT

Fri, 4 December 2020

Arrow's Impossibility Theorem

Above all, everyone wants voting to be fair. What does fair mean and how can we measure it? Kenneth Arrow posited a simple set of conditions that one would certainly desire in a voting system. For example, unanimity - if everyone picks candidate A, then A should win!

Yet surprisingly, under a few basic assumptions, this theorem demonstrates that no voting system exists which can satisfy all the criteria.

This episode is a discussion about the structure of the proof and some of its implications.

Works Mentioned

A Difficulty in the Concept of Social Welfare by Kenneth J. Arrow

Three Brief Proofs of Arrows Impossibility Theorem by John Geanakoplos

Thank you to our sponsors!

Better Help is much more affordable than traditional offline counseling, and financial aid is available! Get started in less than 24 hours. Data Skeptic listeners get 10% off your first month when you visit: betterhelp.com/dataskeptic

Let Springboard School of Data jumpstart your data career! With 100% online and remote schooling, supported by a vast network of professional mentors with a tuition-back guarantee, you can't go wrong. Up to twenty $500 scholarships will be awarded to Data Skeptic listeners. Check them out at springboard.com/dataskeptic and enroll using code: DATASK

Direct download: arrows-impossibility-theorem.mp3
Category:general -- posted at: 8:39am PDT

Fri, 27 November 2020

Face Mask Sentiment Analysis

As the COVID-19 pandemic continues, the public (or at least those with Twitter accounts) are sharing their personal opinions about mask-wearing via Twitter. What does this data tell us about public opinion? How does it vary by demographic? What, if anything, can make people change their minds?

Today we speak to, Neil Yeung and Jonathan Lai, Undergraduate students in the Department of Computer Science at the University of Rochester, and Professor of Computer Science, Jiebo-Luoto to discuss their recent paper. Face Off: Polarized Public Opinions on Personal Face Mask Usage during the COVID-19 Pandemic.

Works Mentioned
https://arxiv.org/abs/2011.00336

Emails:
Neil Yeung
nyeung@u.rochester.edu

Jonathan Lia
jlai11@u.rochester.edu

Jiebo Luo
jluo@cs.rochester.edu

Thanks to our sponsors!

Springboard School of Data offers a comprehensive career program encompassing data science, analytics, engineering, and Machine Learning. All courses are online and tailored to fit the lifestyle of working professionals. Up to 20 Data Skeptic listeners will receive $500 scholarships. Apply today at springboard.com/datasketpic
Check out Brilliant's group theory course to learn about object-oriented design! Brilliant is great for learning something new or to get an easy-to-look-at review of something you already know. Check them out a Brilliant.org/dataskeptic to get 20% off of a year of Brilliant Premium!

Direct download: face-mask-sentiment-analysis.mp3
Category:general -- posted at: 10:56am PDT

Fri, 20 November 2020

Counting Briberies in Elections

Niclas Boehmer, second year PhD student at Berlin Institute of Technology, comes on today to discuss the computational complexity of bribery in elections through the paper “On the Robustness of Winners: Counting Briberies in Elections.”

Links Mentioned:
https://www.akt.tu-berlin.de/menue/team/boehmer_niclas/

Works Mentioned:
“On the Robustness of Winners: Counting Briberies in Elections.” by Niclas Boehmer, Robert Bredereck, Piotr Faliszewski. Rolf Niedermier

Thanks to our sponsors:

Springboard School of Data: Springboard is a comprehensive end-to-end online data career program. Create a portfolio of projects to spring your career into action. Learn more about how you can be one of twenty $500 scholarship recipients at springboard.com/dataskeptic. This opportunity is exclusive to Data Skeptic listeners. (Enroll with code: DATASK)

Nord VPN: Protect your home internet connection with unlimited bandwidth. Data Skeptic Listeners-- take advantage of their Black Friday offer: purchase a 2-year plan, get 4 additional months free. nordvpn.com/dataskeptic (Use coupon code DATASKEPTIC)

Direct download: counting-briberies-in-elections.mp3
Category:general -- posted at: 8:26am PDT

Fri, 13 November 2020

Sybil Attacks on Federated Learning

Clement Fung, a Societal Computing PhD student at Carnegie Mellon University, discusses his research in security of machine learning systems and a defense against targeted sybil-based poisoning called FoolsGold.

Works Mentioned:
The Limitations of Federated Learning in Sybil Settings

Twitter:

@clemfung

Website:
https://clementfung.github.io/

Thanks to our sponsors:

Brilliant - Online learning platform. Check out Geometry Fundamentals! Visit Brilliant.org/dataskeptic for 20% off Brilliant Premium!

BetterHelp - Convenient, professional, and affordable online counseling. Take 10% off your first month at betterhelp.com/dataskeptic

Direct download: sybil-attacks-on-federated-learning.mp3
Category:general -- posted at: 10:25am PDT

Fri, 6 November 2020

Differential Privacy at the US Census

Simson Garfinkel, Senior Computer Scientist for Confidentiality and Data Access at the US Census Bureau, discusses his work modernizing the Census Bureau disclosure avoidance system from private to public disclosure avoidance techniques using differential privacy. Some of the discussion revolves around the topics in the paper Randomness Concerns When Deploying Differential Privacy.

WORKS MENTIONED:

“Calibrating Noise to Sensitivity in Private Data Analysis” by Cynthia Dwork, Frank McSherry, Kobbi Nissim, Adam Smith
"Issues Encountered Deploying Differential Privacy" by Simson L Garfinkel, John M Abowd, and Sarah Powazek
"Randomness Concerns When Deploying Differential Privacy" by Simson L. Garfinkel and Philip Leclerc

Check out: https://simson.net/page/Differential_privacy

Thank you to our sponsor, BetterHelp. Professional and confidential in-app counseling for everyone. Save 10% on your first month of services with www.betterhelp.com/dataskeptic

Direct download: differential-privacy-at-the-us-census.mp3
Category:general -- posted at: 8:13am PDT

Thu, 29 October 2020

Distributed Consensus

Computer Science research fellow of Cambridge University, Heidi Howard discusses Paxos, Raft, and distributed consensus in distributed systems alongside with her work “Paxos vs. Raft: Have we reached consensus on distributed consensus?”

She goes into detail about the leaders in Paxos and Raft and how The Raft Consensus Algorithm actually inspired her to pursue her PhD.

Paxos vs Raft paper: https://arxiv.org/abs/2004.05074

Leslie Lamport paper “part-time Parliament”
https://lamport.azurewebsites.net/pubs/lamport-paxos.pdf

Leslie Lamport paper "Paxos Made Simple"
https://lamport.azurewebsites.net/pubs/paxos-simple.pdf

Twitter : @heidiann360

Thank you to our sponsor Monday.com! Their apps challenge is still accepting submissions! find more information at monday.com/dataskeptic

Direct download: distributed-consensus.mp3
Category:general -- posted at: 10:36pm PDT

Fri, 23 October 2020

ACID Compliance

Linhda joins Kyle today to talk through A.C.I.D. Compliance (atomicity, consistency, isolation, and durability). The presence of these four components can ensure that a database’s transaction is completed in a timely manner. Kyle uses examples such as google sheets, bank transactions, and even the game rummy cube.

Thanks to this week's sponsors:

Monday.com - Their Apps Challenge is underway and available at monday.com/dataskeptic
Brilliant - Check out their Quantum Computing Course, I highly recommend it! Other interesting topics I’ve seen are Neural Networks and Logic. Check them out at Brilliant.org/dataskeptic

Direct download: acid-compliance.mp3
Category:general -- posted at: 6:00am PDT

Fri, 16 October 2020

Patrick Rosenstiel joins us to discuss the The National Popular Vote.

Direct download: national-popular-vote-interstate-compact.mp3
Category:general -- posted at: 8:24am PDT

Mon, 12 October 2020

Defending the p-value

Yudi Pawitan joins us to discuss his paper Defending the P-value.

Direct download: defending-the-p-value.mp3
Category:general -- posted at: 6:00am PDT

Mon, 5 October 2020

Retraction Watch

Ivan Oransky joins us to discuss his work documenting the scientific peer-review process at retractionwatch.com.

Direct download: retraction-watch.mp3
Category:general -- posted at: 8:00am PDT

Mon, 21 September 2020

Crowdsourced Expertise

Derek Lim joins us to discuss the paper Expertise and Dynamics within Crowdsourced Musical Knowledge Curation: A Case Study of the Genius Platform.

Direct download: crowdsourced-expertise.mp3
Category:general -- posted at: 7:00am PDT

Mon, 14 September 2020

The Spread of Misinformation Online

Neil Johnson joins us to discuss the paper The online competition between pro- and anti-vaccination views.

Direct download: the-spread-of-misinformation-online.mp3
Category:general -- posted at: 7:00am PDT

Mon, 7 September 2020

Consensus Voting

Mashbat Suzuki joins us to discuss the paper How Many Freemasons Are There? The Consensus Voting Mechanism in Metric Spaces.

Check out Mashbat’s and many other great talks at the 13th Symposium on Algorithmic Game Theory (SAGT 2020)

Direct download: consensus-voting.mp3
Category:general -- posted at: 7:00am PDT

Mon, 31 August 2020

Voting Mechanisms

Steven Heilman joins us to discuss his paper Designing Stable Elections.

For a general interest article, see: https://theconversation.com/the-electoral-college-is-surprisingly-vulnerable-to-popular-vote-changes-141104

Steven Heilman receives funding from the National Science Foundation. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author and do not necessarily reflect the views of the National Science Foundation.

Direct download: voting-mechanisms.mp3
Category:general -- posted at: 7:00am PDT

Mon, 24 August 2020

False Consensus

Sami Yousif joins us to discuss the paper The Illusion of Consensus: A Failure to Distinguish Between True and False Consensus. This work empirically explores how individuals evaluate consensus under different experimental conditions reviewing online news articles.