Showing posts with label big data. Show all posts
Showing posts with label big data. Show all posts

Wednesday, May 07, 2025

Making Sense of Chaos: A Better Economics for a Better World

Making Sense of Chaos: A Better Economics for a Better World by J. Doyne Farmer

The author uses data models to look at complex problems. This has been done at times to piece together opportunities in the market. (The author also points out that efficient markets imply that there are no easy arbitrage opportunities. However, markets are efficient because arbitrageurs will come in and eliminate inefficiencies.) Some of these models can help predict things with greater accuracy than other more conventional approaches. He is able to point out cases where his models have worked well. He then uses this as a way to predict future activity. This may get a little more questionable as things move to the unknown. (For example, the cost of renewable energy may be steadily going down making it cheaper to focus investments there. However, storage and roll out could have other impacts.)

Economics awards Nobel prizes for "good ideas" while physics requires "proof". Should economics theories be fully validated in the real world? It is an interesting concept. The big data models could be the future. How would this impact economics and the world in the future?

Thursday, April 11, 2024

Uncharted: Big Data as a Lens on Human Culture

Uncharted: Big Data as a Lens on Human Culture by Jean-Baptiste Michel and Erez Aiden

The Google book-scanning project gives us easy access to a great library of books and can help us understand the evolution of language. The authors were able to analyze how words changed and when they changed. Many irregular verbs have slowly gone out of style. The least used ones are the first to regularize, while the more common ones are slower to change. We can also see how words like babysitter gradually evolve.

The describe their findings as well as the work they have gone through to be able to do the research. There are plenty of great nuggets in the book. However, the writing, while personable comes up lacking.

Wednesday, July 05, 2023

Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy

Big data has the potential to erase past prejudices and provide an equal playing field for everybody. Rather than rely on biased feelings, computers can look just at the numbers to make objective decisions. This could allow previously redlined groups better access to loans and money. It could help with hiring and termination of employees. It could also better target resources to improve education and reduce crime. Big data could also help companies make a lot of money.

Alas, in the real world things have not quite turned out that way. The one area where big data has had the most success is in earning companies a lot of money. However, this has been fraught with significant negativity.  At its worst, big data may make totally erroneous predictions based on limited or erroneous data. It can be very difficult to appeal these decisions. At other times, the models serve to ingrain past prejudices in seemingly objective ways.

A good big data model needs a continuous feedback loop to improve predictions and improve the results. It relies on good data. It also must have data that is meaningful and well collected. A grading of teachers by improved student test scores may seem objective, but it is easily manipulated. If there was cheating the previous year, the improvement would be lower. The single data point could also be subject to isolated factors (such as sick students on testing day.)

Bucketing people into groups may make analysis more simple, but it could also produce misleading results for individuals. Using proxies can also have problems. Some data may be easier to collect than what is attempted to be measured, but does not provide the same result.

Data also has the challenge that it can lead to unintended changes in behavior. This is especially problematic when proxies are used. Instead of showing an improvement in the desired behavior, the model encourages changes in the proxy. Data may also be fudged or the model tweaked to get a desired result. The "objective" analysis is merely an insulating layer between subjectiveness.

Big data is a tool. It can help humans uncover hidden gems. However, it can also lead to bad subjectivity in the name of objectivity. It is important to understand the details before relying on data to make decisions.

Tuesday, June 13, 2017

The Signal and the Noise: Why So Many Predictions Fail--but Some Don't

We spend a hug amount of money and computational power on predicting weather, yet we still complain about inaccurate weather forecasts. In spite of this, weather forecasting is one of the "success stories". We have much more accurate forecasts than we did prior to the computational advances. However, beyond a week, weather forecasts fare no better than guesswork. Small changes in variables can significantly alter the long term prognosis. People can also cause kinks in the operations. Commercial weather forecasts have a tendency to overestimate precipitation. (They would rather have someone pleasantly surprised by a sunny day than have an activity ruined by rain.) There was also a case of flooding in North Dakota where an accurate river level prediction was made. However, only the average number was shared rather than the range. The river crested within the range, which happened to be just above the flooding level (and above the average level predicted.)
Baseball provides a rich source of data about many players. It also provides many opportunities for inaccurate predictions about players. Successful teams use a mixture of scouting and statistical analysis to find the best players for the money.
There are a number of biases in the data analysis and predictions we see. Bold predictions are most likely to get press coverage, but are least likely to be right. It is almost always possible to find a significant pattern in the "noise", but that doesn't do much good for predicting future signals. People also tend to be really bad at understanding what data means. Furthermore, news coverage tends to focus on the outliers, even though they tend to be the most inaccurate.
Coverage of global warming provides a cautionary example. There is scientific consensus on the negative impact of human activities on the earth's climate. However, consensus does not necessarily mean good science. Rather than being a balanced average of different opinions, a consensus tends to be dominated by the loudest or most forceful voice. For climate change, the initial view was simply that a greenhouse effect existed and that human activity contributed to increasing in gases. After this point, things got wonky. Discussion switched to global warming, with precise numbers given for warming predictions. When these tended to overstate the warming, the predictions were revised down and models were calibrated. However, the more extreme predictions were the ones that received more press coverage. This would distort the public's view of the situation and give greater credence to the opponents. The response to global warming involves politics, and politics is concerned with the short term impacts, not the long term results. Thus, the noise ends up being twisted towards short term purpose, while the signal is left in the scientific circles.
Predicting terrorism is a lot like predicting earthquakes. We know it is likely to happen, with the minor activities being more frequent than the high-body-count ones. However, we are not good at knowing the specifics. The September 11, 2001 attacks were "unknown unknowns". They were just not something we expected or thought to expect. This made prediction difficult. Terrorism does tend to follow a power-law distribution giving us an idea that a terror attack might be due, but no more than that. (Ironically, Israel seems to buck the power law trend. They permit small-scale attacks to happen, but focus efforts on limiting the more damaging large scale attacks.)
The real problem with predictions is people. The sensational tends to be more appealing than well-thought out. A single pronouncement is given more weight than one couched in uncertainty - even though the uncertain one is much more truthful. People also tend to value "loyalty", giving more credit to those that stick by their guns, even though a willingness to change predictions in face of data makes for more valuable predictions. What are we to do?

Tuesday, May 30, 2017

Everybody Lies: Big Data, New Data, and What the Internet Can Tell Us About Who We Really Are

Everybody Lies weaves together two interweaving threads. The first is that what people type in private in a search engine tells a lot more about their true feelings than what they tell to pollsters. (People will often search for "non socially acceptable" things such as racist or sexual keywords, but would not admit that to pollsters.) The second is that analyzing big data can provide us answers that we could not find using smaller data sets. An example of a finding was on crimes caused by violent movies. Using hourly crime data and violent movie box office, it was found there was a decrease in violent crimes when a popular violent movie was playing. This may be due to the violence-inclined watching the movie instead of getting drunk and violent that evening.

Most of the work delves in to the big data analysis of the non-socially acceptable topics. People would not tell a pollster something, but they would be willing to type it into a search engine. The difference between the public postings on Facebook and the private searches on Google would be an interesting avenue of exploration. (But, alas, the book doesn't go there.) The author acknowledges that there are weaknesses in using extremely large data sets. It is very well possible to find an answer that is merely coincidence. You can almost always find the answer that you are looking for, but need to be careful to make sure it is really true. (However, I wonder if he also falls victim to this also. If something is unacceptable in an area, the people publicly supporting it may be lower, while searches may be higher. However, would the reverse also be true, with people in an area where it is acceptable to say they support it, while not actually searching for it?)

Anti-muslim behavior provides an interesting case study. After attacks in San Bernadino, anti Muslim sentiment was on the rise. Obama tried to quell this by giving a speech stressing peace. The speech went over well with the media. However, a spike in Muslim-hate searchers occurred during the speech. The one time when it went down was when he talked about Muslim athletes. People were then more interested in searching out how Muslims were similar to them. This helped provide a base for future attempts at reducing violent behavior.

The analysis of "border cases" provides some interesting insights. There is a strict test score cut off for admittance to the most prestigious high school in New York. However, people that barely make the cut off seem to get into equally prestigious colleges as those who barely make the cutoff. This seems to show the high school has very little value. (However, it could also show that admissions officers favor students from a diversity of schools and may make it more difficult for those in the best school to get into their desired college.) Similar results were shown for people who got into Penn State and Harvard. Regardless of which school they chose, they seemed to have equally successful careers. (This does leave plenty of questions. Did people that went to Penn State work harder? Does Penn State have an honors program that provides a similar environment as an ivy? Is Harvard a mediocre experience for those without wealthy connections? Would people with similar academic profiles that did not apply to either show similar results?) It does seem to show that it is the person, not the circumstances that lead to success. Or perhaps that people that try and fail have a chip on their shoulder and are likely to work harder to succeed in the long run.

The author pays a debt of gratitude to Steven Levitt and Freakanomics in inspiring him to look for quirky answers to other problems. (Though he does claim that Levitt has fallen from grace to to political incorrectness and a coding issue - I guess I missed that one.) Now big data is the force that can finally put the "science" into social science. With large data sets, we can legitimately probe human behavior in a way that natural scientists can probe nature. However, there still are challenges. In some instances, "little data" can be better used. Often the best results can be found by combining multiple sources that include big data, enriched by more traditional "little data".

Sunday, December 30, 2012

Information Filtering

The corporate computer world is currently all over "Big Data". Companies are collecting all kinds of data. Now they just have to have means of analyzing it in order to help their corporate mission. You can already see how some things are done now. Safeway has a "just for you" programs that offers customized discounts based on an individual's shopping habits. Staples adjusts online prices based on somebody's location and proximity to Staples and competitors' stores. And those are only the cases that get big Wall Street Journal feature articles. There are many other cases that fall under the radar.

There are also many companies struggling to figure out what to do with the data.

But what about individual people?

A few decades ago, the radio pretty much told you what you would listen to. The radio stations played the Beatles. Everyone listened to the Beatles. That was that. If you happened to live close to an indie radio station or indie record store you may get something different. Or you could really scour mail-order catalogs or friends with demo tapes. It took significant effort to find anything unique.

Today, you can turn on Spotify, Amazon, Pandora, or even iTunes and find millions of songs. Finding many obscure artists is not the problem. Filtering through the mass of obscurity is now the challenge. You almost long for the DJs to tell you what to buy.

Specialization and customization have also made things worse. If you wanted to buy a home computer in the 80s, you chose between the Commodore 64, Apple II, or Atari 800. You might have to hunt around town to try to find the best prices, but you knew what each would do. Now you can choose among windows, macs, linux, ChromeOS, Android and iOS. And then there are near infinite variations of each model. Different retailers may have different model numbers (with retailers often trying to have custom version to prevent people from "showrooming" in their store.) The information is now abundantly available. Filtering it has become a problem. (Just try to say "give me the cheapest computer with a Blue Ray, 4 GB memory and 1TB hard drive - and that doesn't even go in to processor, OS, cores, or whatnot.)

Where does this lead us? Will businesses begin to customize their consumer offerings so well that there is an individual model number for each consumer? So much for shopping around.

Even outside of retail, the information glut has a way of overtaking us. Now you can flip on the internet and catch the outcome of every college football game as it happens. There is no need to worry about which games are televised, or what scores the TV announcer feels are interesting. You get everything. You also get the polls as soon as they are available. No more waiting for the Monday paper to see how your team is doing. But, with this media dominated football, you also start to loose the gameday experience. Saturday afternoon games are becoming endangered. Local teams? Why bother. You can just get everything on TV. But, you can now trash talk with fans all over the nation right as the game is happening. You just have to find the right message board - and there are hundreds. Now just try to filter through all these to find things you really are interesting in. Maybe comment #653 on board #203. Or maybe his Twitter feed is the best place to look. Before, it might be difficult to find somebody interested in an out of town game. Now, its difficult to filter through all the garbage to find the intelligent conversation.

How do you deal with the information overload? How do you focus on what data is really needed, and filter out the garbage? In its infancy Google did a great job of showing the most relevant results. Today, the spammers and SEO-gurus are catching up and sometimes winning. A simple query may return page upon page of "junk" results. Is it time for the "new" search engine to help us finally filter through the glut?