Showing posts with label information filtering. Show all posts
Showing posts with label information filtering. Show all posts

Tuesday, November 23, 2021

Too Much Information: Understanding What You Don’t Want to Know

Today we live in a time of abundant information. A single click or tap can get us anything we want to know. But does knowing make our lives better? There are many government initiatives out there that make that assumption. However, they often don't stand up to the results. 

Sometimes the act of reporting can trigger a reaction, even if unjustified. Labeling items as to whether or not the contain GMO ingredients makes it appear that GMO ingredients are harmful and should be avoided. However, most research shows them to be perfectly fine for humans. Other labeling has encouraged change that is helpful (such as trans-fat labeling.)

In some cases, the results can have very different behaviors. Putting calories to the left of food on menus instead of to the right causes different behaviors. The general result of calorie labeling also differs by groups. For some people (often the poor), the labeling encourages seeking out larger calorie counts to get more "food for the buck". Some people may have the discipline to use the calories to reduce consumption. Others may become more stressed out over it and eat more.

Providing information requires effort. The costs involved with producing the information need to be balanced with the benefit received. Government requests a great deal of information. How much of this is really needed? Are there alternatives to collecting it? Many social programs are means tested. In order to reduce misallocation of funds, huge amounts of information are required. This ends up excluding some of the people that are most in need. Would we be better off just accepting that some non-deserving people would receive benefits in order to ensure access to all those in need. (And even better yet, why not just allow everybody to take advantage?) Information requirements is both intentionally and unintentionally a huge barrier to entry.

Excessive disclosures can be useless. (Who pays attention to the multitude of privacy policies?) They can also encourage the opposite behavior. A doctor that discloses an interest in a certain treatment may thus unwittingly encourage the patient to select that treatment. The patient feels an obligation to support the doctor with that treatment. The doctor now feels relieved from conflict of interest concerns. Is this really what was intended by disclosure?

We as individuals and society need to focus on information that provides us benefit. We also need to realize that we are biased to what we have. Could are lives be better with less?

Sunday, December 01, 2013

Filter Bubble

Companies like Google, Facebook, Axiom and Blue Kai know a lot about you. They can use this to taylor websites and advertising to individual users. This can have serious impact on our society.

Huge amounts of information are produced every day. Determining what is relevant is a huge challenge. That is where algorithms come in. Google taylors search results to the individual users. Facebook's news feed is based on what it thinks users are interested in seeing. Advertisers and retailers also customize their online promotions based on what they think users will be interested in seeing. This can cause people to reside in a "bubble" where they only see things that they like. They may not see opposing viewpoints or anything that allows them to think. This is even more worrisome since most of this data filtering is done transparently without users knowing they are living in a bubble. (Perhaps this is a form of "mind control" where the big companies can gradually nudge people in a direction they would like.)

Personalization is treated as a "black box" by most companies. The companies themselves may not even know exactly what results are being returned. They can tweak algorithms based on feedback. However, they probably could not tell exactly what type of results would be returned. The algorithms can use various different data points to personalize. Thus, even if a user is not logged on, their location, computer or web browser could identify them and provide personalized results.

This book rambles on to sound the alarm against "filter bubbles" that allow individuals to live in their isolated worlds filtered to provide what they want. This will keep out information from opposing viewpoints. It will also tend to stock them up with the most "sensational" junk-food content rather than the "good for you content." This could result in a dumbing down of society as people don't work their brains to get around new thoughts. Also, it can make it more difficult for new thoughts and media to get out there. (Since discovery is often based on "likes", only things most like what already exist will tend to get more exposure. Thus, even though it is theoretically easier for new things to be discovered, it is actually much more difficult to find innovation.

The filter bubble is often transparent to users, and can gradually steer content to more extreme viewpoints. This can lead to highly cantankerous partisan discussions. (Since each side is not exposed to the opposing view, they not even understand how others can feel that way.) Even worse, people wont realize they are living in the filter bubble and assume that everybody else is receiving the same information. Discovering "new" ideas can actually be more difficult than it was in the days when everybody was force-fed the same broadcasts.

I was expecting this book to be a discussion of the difficulty we have today of "filtering" through all the information out there. Instead, it focussed on the danger of a personalized web. The two are tightly related. However, this book seemed to spend a lot of time rambling from bullet point to bullet point. It contained plenty of good ideas, but the connections where not very strong.

The amount of information out there is enormous. Discovering useful information is becoming more and more difficult. The quantity of "junk" out there seems to be growing at a faster rate than the amount of useful data. Fifteen years ago, it was easy to put a website out there and get visitors interested in the content. You might get a few random spammers or bots, but most traffic was legitimate. Similarly, if you wanted to search for something, you could use one of the numerous search engines and find relevant sites. The search results may contain a bunch of sites that you were not interested in, but this was more a result of bad algorithms than bad sites.

Today, however, there is so much junk out there. You may have to wade through numerous spam and junk sites to get to the site you want. People are much less likely to find a site that somebody just put up. I see more traffic on this blog than on my proto-blog from 1996. However, the quality of traffic seems to be much worse (at least judging from the ratio of real comments to spam comments.) It is harder to discover quality new content. And it is harder for quality new content to be discovered.

I find myself spending more time on "established" sites. They may occasionally guide me to independent sites. However, they are more likely to simply direct to other well-known commercial sites. With so much information out there, curation has to be done somewhere. I don't have the time to do it (or even to create an algorithm to do it.) I'm dependent on somebody to do it for me. Since this is a huge undertaking, these "somebodies" will likely be large corporations that need to earn money. Since I am cheep, putting up with advertising is my most likely "payment". This puts me in a position vulnerable to being influenced by whatever the corporation or the advertisers desire. (Ironically, at the same time the internet is giving away unlimited content for the price of advertising, broadcast media has become more reliant on "subscriber fees" as part of its business model.) Thus, we become subject to whatever whims the big algorithms have. Is this really much different from being beholden to the broadcaster's desires? At least with the broadcasters, we were likely to find something new we liked. With personalization we can find ourselves further ghettoized. (I'm often finding that problem with online radio. I can create a station that plays songs I like within a very narrowly defined range. However, I get sick of the same type of music and want more variety. However, it is difficult to get variety without a bunch of junk. I'd almost prefer to have a DJ picking the music for me.) I still haven't found a recommendation engine that does a really good job. With the glut of information out there, one of the big challenge is filtering the unique from the derivative. Perhaps now is the time to reinvent the web.

Sunday, December 30, 2012

Information Filtering

The corporate computer world is currently all over "Big Data". Companies are collecting all kinds of data. Now they just have to have means of analyzing it in order to help their corporate mission. You can already see how some things are done now. Safeway has a "just for you" programs that offers customized discounts based on an individual's shopping habits. Staples adjusts online prices based on somebody's location and proximity to Staples and competitors' stores. And those are only the cases that get big Wall Street Journal feature articles. There are many other cases that fall under the radar.

There are also many companies struggling to figure out what to do with the data.

But what about individual people?

A few decades ago, the radio pretty much told you what you would listen to. The radio stations played the Beatles. Everyone listened to the Beatles. That was that. If you happened to live close to an indie radio station or indie record store you may get something different. Or you could really scour mail-order catalogs or friends with demo tapes. It took significant effort to find anything unique.

Today, you can turn on Spotify, Amazon, Pandora, or even iTunes and find millions of songs. Finding many obscure artists is not the problem. Filtering through the mass of obscurity is now the challenge. You almost long for the DJs to tell you what to buy.

Specialization and customization have also made things worse. If you wanted to buy a home computer in the 80s, you chose between the Commodore 64, Apple II, or Atari 800. You might have to hunt around town to try to find the best prices, but you knew what each would do. Now you can choose among windows, macs, linux, ChromeOS, Android and iOS. And then there are near infinite variations of each model. Different retailers may have different model numbers (with retailers often trying to have custom version to prevent people from "showrooming" in their store.) The information is now abundantly available. Filtering it has become a problem. (Just try to say "give me the cheapest computer with a Blue Ray, 4 GB memory and 1TB hard drive - and that doesn't even go in to processor, OS, cores, or whatnot.)

Where does this lead us? Will businesses begin to customize their consumer offerings so well that there is an individual model number for each consumer? So much for shopping around.

Even outside of retail, the information glut has a way of overtaking us. Now you can flip on the internet and catch the outcome of every college football game as it happens. There is no need to worry about which games are televised, or what scores the TV announcer feels are interesting. You get everything. You also get the polls as soon as they are available. No more waiting for the Monday paper to see how your team is doing. But, with this media dominated football, you also start to loose the gameday experience. Saturday afternoon games are becoming endangered. Local teams? Why bother. You can just get everything on TV. But, you can now trash talk with fans all over the nation right as the game is happening. You just have to find the right message board - and there are hundreds. Now just try to filter through all these to find things you really are interesting in. Maybe comment #653 on board #203. Or maybe his Twitter feed is the best place to look. Before, it might be difficult to find somebody interested in an out of town game. Now, its difficult to filter through all the garbage to find the intelligent conversation.

How do you deal with the information overload? How do you focus on what data is really needed, and filter out the garbage? In its infancy Google did a great job of showing the most relevant results. Today, the spammers and SEO-gurus are catching up and sometimes winning. A simple query may return page upon page of "junk" results. Is it time for the "new" search engine to help us finally filter through the glut?