Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Sunday, October 26, 2025

Causal Inference: The Mixtape

Causal Inference: The Mixtape by Scott Cunningham

This is primarily an econometrics textbook. There are dense sets of equations and programs (in Stata and R) for performing Causal inference. I skipped over a lot of that and focussed on the narrative. The book primarily deals with teasing out the cause from complex data. Natural Experiments often play a key role. Two things that are almost the same with a slight difference can often provide insight into a policy based on that difference. However, you must also be careful to control for other factors. A lot of the math in this involves doing that. Often graphic visualizations can help to tease out these bits of interest. Much of the results are "freakonomics" style. These results are often well debated. Often some other factors can be found that swing the results in another direction. (Though sometimes the study is just bad.) Having an openness to accept a different result when the data shows otherwise is important for the process. I would prefer a book that is heavier on the stories and less on the mechanics, but I could not expect more from a textbook. The author also pulls in song titles for the chapter headings (thus a "mixtape").

Monday, September 16, 2024

Factfulness: Ten Reasons We're Wrong About the World--and Why Things Are Better Than You Think

Factfulness: Ten Reasons We're Wrong About the World--and Why Things Are Better Than You Think by Hans Rosling, Anna Rosling Rönnlund, Ola Rosling

As humans, we often get stuck in our beliefs and confuse relative differences with absolute change. We are drawn to the urgent, and often fail to notice the slow, long term changes. Activists use these tendencies to entice us to donate to their cause via the urgency. This all combines us to view the world as a horrible place that is getting worse. 

The facts tell a different story.

Child mortality is a simple measure that is difficult to fudge. It has been improving throughout the world. Even the least developed countries in Africa tend to be better off than European countries were a few centuries ago. Though we often get caught comparing them to Europe today. 

One way to look at development is to divide people into 4 groups. Most of Europe and North America are in Group 4, though there are pockets in lower groups. Just about every country has some Group 4 citizens. Group 1 is the lowest group and what we think of as living in poverty and often on the brink of starvation. The other groups are in between. It is a significant step up from one group to the other. More and more people throughout the world have been moving to higher groups. This is typically accompanied by reduced birthrates, increased consumption and increased standards of living. Many of the countries that are "developing" are also increasing the carbon emissions - but still emit much lower per person than developed countries.

Today we see climate change advocates try to tie every extreme weather event to climate change. This may be a good way of advancing their cause today. Unfortunately, it damages long term credibility. The facts are not as dire and extreme as they would make them out. Humans are just not wired to pay attention to something that happens slowly. 

We can use statistics to help us to take a step back and make wise decisions. We may be missing opportunities due to preconceived notions. Slow change can gradually trick us into making changes that we would otherwise not make all at once. This can be used for both good and bad.

Thursday, April 11, 2024

Everydata: The Misinformation Hidden in the Little Data You Consume Every Day

Everydata: The Misinformation Hidden in the Little Data You Consume Every Day by John H. Johnson and Mike Gluck

We see different types of data throughout the day. From the few bytes of data that represent the time on the alarm clock, to the megabytes representing pictures and videos, we are constantly bombarded. Understanding what the data contains continues to be a challenge. Sometimes the data is intentionally cherry-picked to help encourage a certain response. Other times, the changes may be unintentional. We are are also likely to be more influenced by anecdotes than the true data. The book, like others provides hints on how we can identify the "truth" behind the data we see. 

Uncharted: Big Data as a Lens on Human Culture

Uncharted: Big Data as a Lens on Human Culture by Jean-Baptiste Michel and Erez Aiden

The Google book-scanning project gives us easy access to a great library of books and can help us understand the evolution of language. The authors were able to analyze how words changed and when they changed. Many irregular verbs have slowly gone out of style. The least used ones are the first to regularize, while the more common ones are slower to change. We can also see how words like babysitter gradually evolve.

The describe their findings as well as the work they have gone through to be able to do the research. There are plenty of great nuggets in the book. However, the writing, while personable comes up lacking.

Saturday, December 09, 2023

Random Acts of Medicine: The Hidden Forces That Sway Doctors, Impact Patients, and Shape Our Health

​Natural experiments can be used to perform experiments that would be difficult (or impossible) to perform in a fully randomized trial. These "experiments" are carried out by analyzing data to  focus on certain attributes of interest while controlling for other attributes. One example was looking at flu shots by birthday. Children with summer birthdays typically have their well child visits at a time when a flu vaccine is not available. They must return at a later time to get the shot. This extra friction leads to a lower flu vaccination rate and a higher prevalence of the flu. 

The book looks at many other random experiments. They typically used anonymized insurance or medicare claims. They also pull down datasets of doctors, events and other things to correlate. Somet results are counter-intuitive. Emergency heart operations tend to have a higher success rate when a cardiology convention is going on. Some lend themselves to explanation. Younger hospitalists tend to have better outcomes because they are more knowledgeable in the latest practices. However, older doctors that see the greatest number cases tend to also have good performance, likely because they have had a chance to learn modern practices. With surgeons, older surgeons do have better outcomes.

There were many other examples of experiments in the book. The authors have a very open approach, with a willingness to accept criticism of their methods. There may be confounding circumstances that impact the results that are seen. People often change their behavior when they know they are being measured. Trying to improve medical outcomes can be extremely challenging. If we find a good measure, participants will often find a way to maximize that measure without necessarily improving overall outcomes. How do we make medicine better? We need to look at the big picture of the human experience, but also realize that humans are providing the medical care. The experience of receiving the care can be as impactful as the actual care received. How do we maximize everything? We have a significant challenge.

Monday, January 30, 2023

The Art of Statistics: How to Learn from Data

Statistics can be powerful, but also mystical. This book attempts to take the mystery out by building up knowledge from the ground up. Statistical topics are developed using real world applications and then analyzed from a "brute force" method. Then the math is applied as an "easy solution" to the problems. The author also pull from his experience using statistics in various cases. One that is covered multiple times is the doctor that had killed is patients. Naive statistical analysis could have caught him. However, it would also have caught innocent doctors. It is important to do more thorough statistical analysis as well as look at the real world conditions. The book concludes with an analysis of statistics and the state of scientific research. P-values have been taken as gospel and many "tricks" have been used by researchers to get publishable results, even if these are not really meaningful.

Thursday, January 26, 2023

Naked Statistics: Stripping the Dread from the Data

Statistics can be a source of much confusion. I recent newspaper article stated "only [small number]% of coaches are women, yet [large number] girls play this sport." The goal was to saw there is a great shortage of female coaches. But using a percentage in one place and number in another is misleading. "Just 5% of coaches are female, while 5,000,000 girls play this sport" sounds really bad. But if there are also 500,000,000 boys that play the sport, it actually shows female coaches are over-represented. Numbers and statistics are a great way to tell these "lies".

The author initially shied away from "math heavy" classes. Abstract ideas did not appeal to him. However, once he was able to connect concrete interests it became much more appealing. In this book he tries to impart that feeling with a lighthearted look at the practical nature of statistics.

Statistics can be used to help infer important conclusions. They can also be used to come up with many false explanations. Knowing the difference can be challenging. The sampling method, the data, the methodology and the questions asked can all play a role in the quality. It is very easy to come up with "accurate" statistics that are totally wrong. We have a duty to find these and understand them. This book is a helpful, easy to understand primer.

Friday, January 06, 2023

The Biggest Bluff: How I Learned to Pay Attention, Master Myself, and Win

Poker is a complex game. Unlike other casino games, there is an important element of skill. However, unlike chess, there is a degree of luck involved. It is similar to life in general. We can exercise skill, but success is still dependent on some luck. 

The Biggest Bluff is the author's attempt to learn and compete in poker at the highest level. She seeks out an expert to coach her. There is the basic skill, but more important is understanding the psychology. You must play offense and defense. The more you learn, the more you realize you don't know. The more advanced players know they have much more they need to continue to learn. You don't want to be unconsciously incompetent.

Friday, January 28, 2022

Bad Science: Quacks, Hacks, and Big Pharma Flacks

There is plenty of "bad science" out in the world. Some is intentionally misleading. Other times they accidentally produce bad results. Bad Science explores different issues with science, with a strong focus on medicine.

Studies funded by pharmaceutical companies frequently validate their own product. Ones that don't are unlikely to be published. A competitor may be "handicapped" to help make the company's drug appear better. Researchers are also likely to fish for a result among the study. (Perhaps Asian men between 51 and 60 show a better result.)

Even outside pharma, there is plenty of bad science. Selective publication results in plenty of studies not being passed. Various means of p-hacking may be used to find interesting results in studies. Natural variation may be mistaken from something significant. The placebo effect can be a very real thing. (Improvements can even be seen when patients are given a pill that they know is a placebo.) The quest for a "magic pill" can lead to many cures of limited value being common. 

The communication of results can also impact the reception of studies. Bigger numbers get more press. However, without understanding the scope of the dataset, they may be irrelevant. Something may appear to cause a "doubling", but that may be only going from 1 in 1000 to 2 in 1000. It could have just as well been a freak account. There are also plenty of "smooth talking" experts that may have little training or even understanding of basic science. It is common for them (as well as the public in general) to put credit in research that backs up their view, while finding faults in studies that don't. Sometimes "scientific" explanations can be given that have little grounding in science (such as for homeopathy's water memory.) 

The book ends with an analysis of the link between vaccines and autism. There were a number of players involved. A doctor was working with patients that wanted to file lawsuits due to autism diagnoses. Media was happy to jump on the story (but only some time after initial publication.) There were plenty of poorly done studies, bad interpretation of data and selective publication that all added to a scare with little basis in reality.

Thursday, November 04, 2021

The Data Detective: Ten Easy Rules to Make Sense of Statistics

On the surface, statistics are concrete facts that describe the world. However, in actuality, they are highly susceptible to manipulation, both intentional and unintentional. Understanding the details is key to making accurate decisions. 

In the UK, one part of the country appeared to have a much greater infant mortality rate. What could be causing that? The mortality rate even differed among similar demographics in the different locations. The "difference" turned out to just be in the way the statistics were reported. The same event would be reported as a "late miscarriage" in one place and a "live birth followed by death" in another. The events were the same, but the different reporting, made one appear to reduce infant mortality. This difference also explains some of the differences in infant mortality across countries. It is not that outcomes are worse, just that they are reported differently. It is common for elaborate charts and analysis to be built on unreliable data. 

Scientific studies, can be very prone to publication bias. An unexpected novel result is likely to get a great deal of press coverage. The "boring" results are less likely to even be published. There is also a strong incentive to publish novel results, rather than try to reproduce previous results. There may have been multiple people trying similar experiments, but only the novel results made it to publication. Then there are issues with the study itself. The significance threshold is arbitrary. A researcher may choose to change the size of the dataset in order to get the results that they are desiring. Psychology and medicine are both very prone to only publishing the "good stuff", leaving us with little ability to know what is truly "good".

On an individual level, we are also susceptible to bias from our pre-existing beliefs. We are more likely to consume media that backs up our beliefs. People that think like us are more likely to be given the benefit of the doubt over those that contradict our viewpoints. The same things happens in research. Results that are differ from what is expected are more likely to be discarded.

Visualizations are a great way to mislead. Different types of numbers can be compared to produce misleading results. Changes in color and scale can change our interpretation of the data. Unrelated data can be shown together to imply relationships were none exists. Graphs and infographics are easily digestible, but not necessarily show the whole picture.

The use of data can impact the data itself. If people know that data will be used for a purpose, their behavior may be changed to account for the data collection. Wrestlers target a specific weight class before they weigh in. Looking at data, it will be rare for a wrestler to be at the bottom of a weight class. Teachers that are graded based on student performance have incentive to manipulate that performance. The student performance is no longer fully objective. Government officials that can control data may manipulate it for their objectives. (This it is important to have non-partisan data organizations.)

Data Detective does a great job of helping to guide us through the pitfalls and benefits of statistics that we encounter in our life. It builds on past work while presenting concrete cases to help us to better understand what is going on in the world around us.

Monday, October 25, 2021

Humble Pi: When Math Goes Wrong in the Real World

Humble Pi is a fun book about math gone bad. Or perhaps smore accurately, it is about humans' poor understanding of numbers. There are big disasters and little flubs described. Bridges collapsed because minor changes greatly increased loads. McDonald's went to court because they seemingly exaggerated the number of meal combinations. Other companies understated the number of combinations. People have been given lethal doses of medications due to misunderstandings of the unit of measurements. The Gimli Glider was a 767 airplane that ran out of fuel midflight. Luckily, the pilot was able to glide it down to a landing. A number of things together created the issue. However, a key issue was the use of wrong units for fuel calculations.

People can sometimes have very bad logic when confronted with big numbers. One meme mistakenly assumed dividing a large number in the millions by another in the millions would have an answer in the millions. There have been internet flame wars over whether a week has 7 or 8 days. (Having something 0-based or 1-based does make a big difference.) 

The intersection of computers and people also create all sorts of problems. In one hilarious case, the upgrade of a an operating system (and downgrade of mail software) resulted in emails only being able to reach recipients within 500 miles. (Thank you speed of light!) Strange results often appear due to the binary representation of numbers. Results that exceed the maximum capacity can be very unpredictable. (When the 8-bit level number in pac-man rolled over, things went bezerk.) However, a similar overrun has also lead to a rocket to crash. (Legacy code also contributed to this. The part that caused the self-destruct did not even need to be run at the time.)

People's inability to behave like "math" can also help finding frauds. A professor can distinguish between those that really did record a large number of coin flips and those that made things up. Those that really did it are more likely to have some longer "runs" than a faker would appear. Forensic accounting can also be used to identify fraud due to bad distributions of numbers, indicating more likely "fakes".

What do we learn from all this? People can make bad mistakes with numbers. However, there are also many cases of bad number usage that we are still living with. There are also plenty of others that have been identified by the perpetrators, but have been "swept under the rug" as trade secrets.

Saturday, July 10, 2021

A Field Guide to Lies: Critical Thinking in the Information Age

This title has gone through a few title changes in the last five years. From A Field Guide to Lies: Critical Thinking in the Information Age to Weaponized Lies: How to Think Critically in the Post-truth Era, before settling on A Field Guide to Lies: Critical Thinking With Statistics and the Scientific Method. The core is similar to other books that explain in detail how statistics, charts and numbers are used to manipulate. Sometimes the manipulation is done to knowingly mislead. Many times, both the reader and writer are complicit in the poor understanding of what they are presenting. Covid-19 and politics have brought even more statistical lying out in the world.

The first part of the book focuses on how statistics and graphs can be used to mislead us. Then it goes on to cover other ways that we can be misled in a narrative fashion. Understanding bayesian probability can be extremely helpful in making sense of things. Slight differences in expressing inclusion and exclusion can trigger vastly different responses in our perception. News coverage tends to focus on the more rare, colorful events. This leads people to become more worried about the least probably things they hear more about. Science news coverage can be even worse. A new exciting research study makes the news. However, this could just as easily have been random noise. Once it is repeated multiple times and analyzed over different populations can we start to draw accurate conclusions. Alas, these meta-analysis do not get the great coverage. 

In analyzing media coverage, Bayes theorem can help us. We can use the 4-way table to analyze the difference from the current understanding. It is often more likely that extreme results are just flukes. The conventional analysis is more likely to be true. However, with more evidence, we have a greater likelihood of the "new" solution supplanting the current.

This book is loaded with examples of many ways that people "lie" (either intentional or accidentally) with the data they are presenting. There are plenty of tools to help us overcome these lies. Information is readily available today. It is now up to us to do the analysis.

Friday, February 12, 2021

Calculated Risks: How to Know When Numbers Deceive You

The best way to get lie to somebody is to present accurate statistics that they do not understand. In Calculated Risks, Gerd Gigerenzer presents many examples in which people make bad decisions based on a poor understanding of the numbers. The first part of the book focusses on medicine. Doctors tend to be very smart, but not experts with statistics. People are often confused when risk is expressed in probabilities rather than frequencies. Relative risk reduction is especially misleading. If 1000 people participate in a screening and 3 die, while 4 of 1000 in a control group die, the screening would have a relative risk reduction of 25%, which sounds impressive. However, the absolute risk reduction is only 0.1%, which doesn't sound as good. If analyzed further it is found that the average increase in life expectancy is 12 days, which sounds even worse. And that doesn't even take into account the costs. 

Breast cancer is called out as something where bad statistical understanding has lead to negative outcomes. Studies have shown no increase in life lived with universal mammograms for people under 50. However, there is a significant risk of false positives. There is also an increased risk of cancer caused by the radiation exposure in the test. Alas, rather than being upset at needless intervention, people are often relieved when a doctor cuts into the breast only to find no cancer. The many statistics about breast cancer lead to confusion. There is a fairly high prevalence of the cancer in women. However, it is often not the cause of death. Similar to prostate cancer, many people that die of other causes are found to have the cancer. These cancers are often benign and do not negatively impact the quality of life. However, treating these with surgery and/or radiation does negatively impact the life. 

The book gives an example of expressing data differently. "If cancer probability is .8%. If somebody has breast cancer, there is a 90% chance of a positive mammogram. If they don't have cancer, there is a 7% chance of positive mammogram." From this example, it looks like a positive test result makes it extremely likely that one has cancer. However, expressing in frequencies makes it more clear. "8 of 1000 women have breast cancer. Of these 8, 7 will have a positive test. Of the 992 without cancer, 70 will have a positive test". This is the same data expressed differently. However, it makes it much more clear that somebody with a positive result most likely does not have cancer. 

Knowing the prevalence in a population can be extremely important in understanding results. If somebody has HIV, there is a 99.9% chance they get a positive test result. If they are not infected, there is a 99.99% chance they will get a negative test. This seems like the test is rock solid. However, if somebody is not in a high risk category, their chance of having HIV is about 0.01%. Thus in a group of 10,000 low-risk men, 1 will have the virus and 9999 will not. The one with the virus will likely get a positive result. Of the 9999 men without HIV, there will also be 1 likely test result. Thus, even with the very high specificity and sensitivity, a positive result from a low risk population only has a 50% chance of being accurate. In the case of AIDS, many people have committed suicide after getting a positive result, even though the accuracy of this result for their population is about the same as a coin flip.

Criminal Justice often misuses statistics on both sides. In the OJ Simpson trial, the defense portrayed his beating of his wife as irrelevant to the his wive's eventual murder. There are millions of men that beat their wives, yet only a small number that eventually murder them. However, expressing in different frequencies produces a different result. Of 100,000 battered women, 45 were murdered. This seems to back up the defense. However, of the 45 murdered women, 40 were murdered by their partners, while 5 were murdered by somebody else. Oops! Now it seem that the wife beating may be very relevant indicator of a murder. The inverse of this is the "prosecutors fallacy". Here, the prosecutor infers that the probability of observing a set of characteristics is the probability that a defendant is innocent.

The book ends with some hypothetical problems and examples of "deliberate misleading" with statistics. (In one case, risk was expressed using low absolute numbers, while benefit was shown using high relative numbers.) Statistical literacy is a key trait in society today. Alas, even most highly educated people do not have it and will fall prey to intentional (and unintentional) misleading representations.


Sunday, March 01, 2020

How Not to Be Wrong: The Power of Mathematical Thinking

How Not To Be Wrong is part an exploration of the joys of math, and part a manual on how to avoid getting duped (or duping oneself) with math. There is also a bit of namedropping of famous mathematicians for good measure. Politicians are great at coming up with good sounding, but meaningless numbers. Wisconsin may have claimed that its 5000 net new jobs accounted for 50% of all new jobs for a year. That could be technically correct. However, since some states lost jobs, a state with 12000 net new jobs would claim to account for 120% of net new jobs, showing the craziness of the data. Science is also rife with "statistically significant random results" from insufficiently large or small studies. There is also a "survivors bias", with many failed studies not being published. If 20 people study something, one is likely to randomly discover a significant result. If that person is the only one to publish, we don't realize the significant result was just random. The "5%" p-value acceptance threshold is just an arbitrary value. However, it does result in a serious amount of 'p-hacking' There are a lot more studies that just meet the threshold than would be suspected by a normal distribution. Similarly, people tend to favor numbers ending in 7 as "random numbers". Thus, an excessive preponderance of vote counts ending in 7, may indicate a rigged election.

Wednesday, August 02, 2017

The Theory That Would Not Die: How Bayes' Rule Cracked the Enigma Code, Hunted Down Russian Submarines, and Emerged Triumphant from Two Centuries of Controversy

Bayes' rule seems very simple. However, it has produced a great deal of controversy throughout its history. Even the name itself is controversial. Bayes does appear to have been one of the first people to produce a paper on it. However, the paper we have was substantially edited by Richard Price and presented at the Royal Society after Bayes' death. Laplace later independently discovered the algorithm and ran with it. Since he was the more renowned mathematical mind, it would make sense to name it after him. However, it was later referred to as Bayes rule by those who popularized it, and that is what we have today.
The Theory That Would Not Die does not spend much time covering the details of Bayes' rule. It is assumed the reader already understands it, or will be able to understand it well enough by following the story line. Instead, the focus is on the conflicts between the "bayeseians" and the "frequentists". Bayes can help determine probabilities given scant data or unknown occurrences and was derided as "subjective". Frequency analysis deals with known observations as was considered a more theoretically accurate. Bayesian analysis would come and go in spurts during its history. In world war ii, it helped lead to cracking the German code and significantly helping the Allied war efforts. Alas, it was deemed so important that it was classified, and thus not disclosed to the general public. The ability to adjust probabilities based on past outcomes made it especially useful for insurance actuaries. Today it has applications in multitudes of fields from medical research to spam filters. It is great at helping to tease out the signal from the noise and find high probability answers given scant data.