Data Skeptic

April 2024
S	M	T	W	T	F	S

	1	2	3	4	5	6
7	8	9	10	11	12	13
14	15	16	17	18	19	20
21	22	23	24	25	26	27
28	29	30

Fri, 31 October 2014

This episode explores the basis of why we can trust encryption. Suprisingly, a discussion of looking up a word in the dictionary (binary search) and efficiently going wine tasting (the travelling salesman problem) help introduce computational complexity as well as the P ?= NP question, which is paramount to the trustworthiness RSA encryption.

With a high level foundation of computational theory, we talk about NP problems, and why prime factorization is a difficult problem, thus making it a great basis for the RSA encryption algorithm, which most of the internet uses to encrypt data. Unlike the encryption scheme Ray Romano used in "Everybody Loves Raymond", RSA has nice theoretical foundations.

It should be noted that although this episode gives good reason to trust that properly encrypted data, based on well choosen public/private keys where the private key is not compromised, is safe. However, having safe encryption doesn't necessarily mean that the Internet is secure. Topics like Man in the Middle attacks as well as the Snowden revelations are a topic for another day, not for this record length "mini" episode.

Direct download: MINI_Is_the_Internet_Secure.mp3
Category:miniepisode -- posted at: 12:41am PDT

Thu, 23 October 2014

Practicing and Communicating Data Science with Jeff Stanton

Jeff Stanton joins me in this episode to discuss his book An Introduction to Data Science, and some of the unique challenges and issues faced by someone doing applied data science. A challenge to any data scientist is making sure they have a good input data set and apply any necessary data munging steps before their analysis. We cover some good advise for how to approach such problems.

Direct download: Practicing_and_Communicating_Data_Science.mp3
Category:data science -- posted at: 10:18pm PDT

Thu, 16 October 2014

[MINI] The T-Test

The t-test is this week's mini-episode topic. The t-test is a statistical testing procedure used to determine if the mean of two datasets differs by a statistically significant amount. We discuss how a wine manufacturer might apply a t-test to determine if the sweetness, acidity, or some other property of two separate grape vines might differ in a statistically meaningful way.

Check out more details and examiles found in the show notes linked below.

https://dataskeptic.com/blog/episodes/2014/t-test

Direct download: MINI_The_T-Test.mp3
Category:miniepisode -- posted at: 7:49pm PDT

Thu, 9 October 2014

Data Myths with Karl Mamer

This week I'm joined by Karl Mamer to discuss the data behind three well known urban legends. Did a large blackout in New York and surrounding areas result in a baby boom nine months later? Do subliminal messages affect our behavior? Is placing beer alongside diapers a recipe for generating more revenue than these products in separate locations? Listen as Karl and I explore these claims.

Direct download: Data_Myths_with_Karl_Mamer.mp3
Category:skepticism -- posted at: 7:35pm PDT

Tue, 7 October 2014

Contest Announcement

The Data Skeptic Podcast is launching a contest- not one of chance, but one of skill. Listeners are encouraged to put their data science skills to good use, or if all else fails, guess!

The contest works as follows. Below is some data about the cumulative number of downloads the podcast has achieved on a few given dates. Your job is to predict the date and time at which the podcast will recieve download number 27,182. Why this arbitrary number? It's as good as any other arbitrary number!

Use whatever means you want to formulate a prediction. Once you have it, wait until that time and then post a review of the Data Skeptic Podcast on iTunes. You don't even have to leave a good review! The review which is posted closest to the actual time at which this download occurs will win a free copy of Matthew Russell's "Mining the Social Web" courtesy of the Data Skeptic Podcast. "Price is Right" rules are in play - the winner is the person that posts their review closest to the actual time without going over.

More information at dataskeptic.com

Direct download: contest.mp3
Category:statistics -- posted at: 9:49pm PDT

Fri, 3 October 2014

[MINI] Selection Bias

A discussion about conducting US presidential election polls helps frame a converation about selection bias.

Direct download: MINI_Selection_Bias.mp3
Category:miniepisode -- posted at: 1:00am PDT

Thu, 25 September 2014

[MINI] Confidence Intervals

Commute times and BBQ invites help frame a discussion about the statistical concept of confidence intervals.

Direct download: MINI_Confidence_Intervals.mp3
Category:miniepisode -- posted at: 10:47pm PDT

Fri, 19 September 2014

[MINI] Value of Information

A discussion about getting ready in the morning, negotiating a used car purchase, and selecting the best AirBnB place to stay at help frame a conversation about the decision theoretic principal known as the Value of Information equation.

Direct download: MINI_Value_of_Information.mp3
Category:miniepisode -- posted at: 12:29am PDT

Wed, 17 September 2014

Game Science Dice with Louis Zocchi

In this bonus episode, guest Louis Zocchi discusses his background in the gaming industry, specifically, how he became a manufacturer of dice designed to produce statistically uniform outcomes.

During the show Louis mentioned a two part video listeners might enjoy: part 1 and part 2 can both be found on youtube.

Kyle mentioned a robot capable of unnoticably cheating at Rock Paper Scissors / Ro Sham Bo. More details can be found here.

Louis mentioned dice collector Kevin Cook whose website is DiceCollector.com

While we're on the subject of table top role playing games, Kyle recommends these two related podcasts listeners might enjoy:

The Conspiracy Skeptic podcast (on which host Kyle was recently a guest) had a great episode "Dungeons and Dragons - The Devil's Game?" which explores claims of D&Ds alleged ties to skepticism.

Also, Kyle swears there's a great Monster Talk episode discussing claims of a satanic connection to Dungeons and Dragons, but despite mild efforts to locate it, he came up empty. Regardless, listeners of the Data Skeptic Podcast are encouraged to explore the back catalog to try and find the aforementioned episode of this great podcast.

Last but not least, as mentioned in the outro, awesomedice.com did some great independent empirical testing that confirms Game Science dice are much closer to the desired uniform distribution over possible outcomes when compared to one leading manufacturer.

Direct download: Game_Science_Dice_with_Louis_Zocchi.mp3
Category:gaming -- posted at: 12:27am PDT

Fri, 12 September 2014

Data Science at ZestFinance with Marick Sinay

Marick Sinay from ZestFianance is our guest this weel. This episode explores how data science techniques are applied in the financial world, specifically in assessing credit worthiness.

Direct download: Zest_Finance.mp3
Category:financial -- posted at: 2:30am PDT

Fri, 5 September 2014

[MINI] Decision Tree Learning

Linhda and Kyle talk about Decision Tree Learning in this miniepisode. Decision Tree Learning is the algorithmic process of trying to generate an optimal decision tree to properly classify or forecast some future unlabeled element based by following each step in the tree.

Direct download: MINI_Decision_Tree_Learning.mp3
Category:miniepisode -- posted at: 12:49am PDT

Fri, 29 August 2014

Jackson Pollock Authentication Analysis with Kate Jones-Smith

Our guest this week is Hamilton physics professor Kate Jones-Smith who joins us to discuss the evidence for the claim that drip paintings of Jackson Pollock contain fractal patterns. This hypothesis originates in a paper by Taylor, Micolich, and Jonas titled Fractal analysis of Pollock's drip paintings which appeared in Nature.

Kate and co-author Harsh Mathur wrote a paper titled Revisiting Pollock's Drip Paintings which also appeared in Nature. A full text PDF can be found here, but lacks the helpful figures which can be found here, although two images are blurred behind a paywall.

Their paper was covered in the New York Times as well as in USA Today (albeit with with a much more delightful headline: Never mind the Pollock's [sic]).

While discussing the intersection of science and art, the conversation also touched briefly on a few other intersting topics. For example, Penrose Tiles appearing in islamic art (pre-dating Roger Penrose's investigation of the interesting properties of these tiling processes), Quasicrystal designs in art, Automated brushstroke analysis of the works of Vincent van Gogh, and attempts to authenticate a possible work of Leonardo Da Vinci of uncertain provenance. Last but not least, the conversation touches on the particularly compellingHockney-Falco Thesis which is also covered in David Hockney's book Secret Knowledge.

For those interested in reading some of Kate's other publications, many Katherine Jones-Smith articles can be found at the given link, all of which have downloadable PDFs.

Direct download: Jackson_Pollock_Authentication_Analysis_with_Kate_Jones-Smith.mp3
Category:art -- posted at: 6:00am PDT

Fri, 22 August 2014

[MINI] Noise!!

Our topic for this week is "noise" as in signal vs. noise. This is not a signal processing discussions, but rather a brief introduction to how the work noise is used to describe how much information in a dataset is useless (as opposed to useful).

Also, Kyle announces having recently had the pleasure of appearing as a guest on The Conspiracy Skeptic Podcast to discussion The Bible Code. Please check out this other fine program for this and it's many other great episodes.

Direct download: MINI_Noise.mp3
Category:miniepisode -- posted at: 6:00am PDT

Fri, 15 August 2014

Guerilla Skepticism on Wikipedia with Susan Gerbic

Our guest this week is Susan Gerbic. Susan is a skeptical activist involved in many activities, the one we focus on most in this episode is Guerrilla Skepticism on Wikipedia, an organization working to improve the content and citations of Wikipedia.

During the episode, Kyle recommended Susan's talk a The Amazing Meeting 9 which can be found here.

Some noteworthy topics mentioned during the podcast were Neil deGrasse Tyson's endorsement of the Penny for NASA project. As well as the Web of Trust and Rebutr browser plug ins, as well as how following the Skeptic Action project on Twitter provides recommendations of sites to visit and rate as you see fit via these tools.

For her benevolent reference, Susan suggested The Odds Must Be Crazy , a fun website that explores the statistical likelihoods of seemingly unlikely situations. For all else, Susan and her various activities can be found via SusanGerbic.com.

Direct download: Guerilla_Skepticism_on_Wikipedia_with_Susan_Gerbic.mp3
Category:wikipedia -- posted at: 6:00am PDT

Fri, 8 August 2014

[MINI] Ant Colony Optimization

In this week's mini episode, Linhda and Kyle discuss Ant Colony Optimization - a numerical / stochastic optimization technique which models its search after the process ants employ in using random walks to find a goal (food) and then leaving a pheremone trail in their walk back to the nest. We even find some way of relating the city of San Francisco and running a restaurant into the discussion.

Direct download: MINI_Ant_Colony_Optimization.mp3
Category:miniepisode -- posted at: 6:00am PDT

Fri, 1 August 2014

Data in Healthcare IT with Shahid Shah

Our guest this week is Shahid Shah. Shahid is CEO at Netspective, and writes three blogs: Health Care Guy, Shahid Shah, and HitSphere - the Healthcare IT Supersite.

During the program, Kyle recommended a talk from the 2014 MIT Sloan CIO Symposium entitled Transforming "Digital Silos" to "Digital Care Enterprise" which was hosted by our guest Shahid Shah.

In addition to his work in Healthcare IT, he also the chairperson for Open Source Electronic Health Record Alliance, an non-profit organization that, amongst other activities, is hosting an upcoming conference. The 3rd annual OSEHRA Open Source Summit: Global Collaboration in Healthcare IT , which will be taking place September 3-5, 2014 in Washington DC.

For our benevolent recommendation, Shahid suggested listeners may benefit from taking the time to read books on leadership for the insights they provide. For our self-serving recommendation, Shahid recommended listeners check out his company Netspective , if you are working with a company looking for help getting started building software utilizing next generation technologies.

Direct download: Data_in_Healthcare_IT_with_Shahid_Shah.mp3
Category:medicine -- posted at: 6:00am PDT

Fri, 25 July 2014

[MINI] Cross Validation

This miniepisode discusses the technique called Cross Validation - a process by which one randomly divides up a dataset into numerous small partitions. Next, (typically) one is held out, and the rest are used to train some model. The hold out set can then be used to validate how good the model does at describing/predicting new data.

Direct download: MINI_Cross_Validation.mp3
Category:general -- posted at: 7:51am PDT

Fri, 18 July 2014

Streetlight Outage and Crime Rate Analysis with Zach Seeskin

This episode features a discussion with statistics PhD student Zach Seeskin about a project he was involved in as part of the Eric and Wendy Schmidt Data Science for Social Good Summer Fellowship. The project involved exploring the relationship (if any) between streetlight outages and crime in the City of Chicago. We discuss how the data was accessed via the City of Chicago data portal, how the analysis was done, and what correlations were discovered in the data. Won't you listen and hear what was found?

Direct download: Streetlight_Outage_and_Crime_Rate_Analysis_with_Zach_Seeskin.mp3
Category:general -- posted at: 6:00am PDT

Fri, 11 July 2014

[MINI] Experimental Design

This episode loosely explores the topic of Experimental Design including hypothesis testing, the importance of statistical tests, and an everyday and business example.

Direct download: MINI_Experimental_Design.mp3
Category:miniepisode -- posted at: 6:00am PDT

Mon, 7 July 2014

The Right (big data) Tool for the Job with Jay Shankar

In this week's episode, we discuss applied solutions to big data problem with big data engineer Jay Shankar. The episode explores approaches and design philosophy to solving real world big data business problems, and the exploration of the wide array of tools available.

Direct download: Data_Skeptic_Podcast_-_Big_Data_Tools.mp3
Category:general -- posted at: 6:00am PDT

Fri, 27 June 2014

[MINI] Bayesian Updating

In this minisode, we discuss Bayesian Updating - the process by which one can calculate the most likely hypothesis might be true given one's older / prior belief and all new evidence.

Direct download: MINI_Bayesian_Updating.mp3
Category:miniepisode -- posted at: 6:00am PDT

Fri, 20 June 2014

Personalized Medicine with Niki Athanasiadou

In the second full length episode of the podcast, we discuss the current state of personalized medicine and the advancements in genetics that have made it possible.

Direct download: Data_Skeptic_Podcast_ep_004_-_Personalized_Medicine_with_Niki_Athanasiadou.mp3
Category:medicine -- posted at: 6:00am PDT

Fri, 13 June 2014

[MINI] p-values

In this mini, we discuss p-values and their use in hypothesis testing, in the context of an hypothetical experiment on plant flowering, and end with a reference to the Particle Fever documentary and how statistical significance played a role.

Direct download: MINI_p-values_.mp3
Category:miniepisode -- posted at: 6:00am PDT

Fri, 6 June 2014

Advertising Attribution with Nathan Janos

A conversation with Convertro's Nathan Janos about methodologies used to help advertisers understand the affect each of their marketing efforts (print, SEM, display, skywriting, etc.) contributes to their overall return.

Direct download: Data_Skeptic_Podcast_ep_002_-_Advertising_Attribution_with_Nathan_Janos.mp3
Category:advertising -- posted at: 6:00am PDT

Fri, 30 May 2014

[MINI] type i / type ii errors

In this first mini-episode of the Data Skeptic Podcast, we define and discuss type i and type ii errors (a.k.a. false positives and false negatives).

Direct download: type_i_type_ii.mp3
Category:miniepisode -- posted at: 6:00am PDT

Fri, 23 May 2014

Introduction

The Data Skeptic Podcast features conversations with topics related to data science, statistics, machine learning, artificial intelligence and the like, all from the perspective of applying critical thinking and the scientific method to evaluate the veracity of claims and efficacy of approaches.

This first episode is a short discussion about what this podcast is all about.

Direct download: Data_Skeptic_Podcast_ep000_-_Introduction.mp3
Category:metadata -- posted at: 3:00am PDT