What if the world economy is actually a malevolent artificial intelligence? Get out your tinfoil hats for this one.
So, if there are three things I hate, they are: 1) "Glossy" or over-reaching futurism. 2) Conspiracy theories and 3) People thinking an idea is important because it came to them in a dream. You must all now know that I'm about to become what I hate. Three times. Today's adventure could have been entitled, "Why We Should All Be Running Around, Screaming, With Our Pants on Our Heads Part II", or "Watson for President". Let's start.
Part the First: "Stateware"
I've been reading James Gleick's "The Information", and am struck by the Difference Engine. I've heard of it before, but I finally get it through my thick skull this time that computers were not always intended to be electronic machines, and in fact the first computers designed were not. I don't know how many difference engines it would take to build a bare-bones CPU, but theoretically it could be done. It would also be really really slow and run on oceans of steam.
My next thought was this: really, you could build a computer out of anything that responds to a binary difference. People have made computers out of biological material, Lego, et cetera. You could even make a computer out of people, in fact. If you replaced transistors with human beings, and wires with human language, you could make a very interesting computational device indeed. It would be error-prone and slow, which would make any software you ran on it very likely to crash, unless you had enough redundancy and sufficiently fast communication.
I thought a lot about governments, laws, economies, etc. and decided that, in a certain light, these systems could be considered software. Laws are essentially algorithmic—I think it no coincidence that we refer to the "code of law". In a discussion on this matter with a friend, I referred to this type of software as "stateware".
Part the Second: A Brief Interlude on the Technological Singularity
Since the beginning of computational technology, scientists have raised the possibility that eventually machines may surpass humans at all "intelligence" related functions, quickly outstripping all of humanity's prowess and (in many cases) taking control of the world.
Frankly, we have no way to predict what will happen once one piece of software is created that rivals the intelligence of a human being, generally. As that software could then write software that is more intelligent than itself, the possibilities are unfathomable. This inability to see anything in the future is the reason the phenomenon is called the "technological singularity".
One of the things that makes the singularity so scary is that an intelligent machine may or may not have goals that match those of humanity. It may have a written-in goal that it takes to an extreme (converting the world into a paperclip factory is my favorite example), or it may take as a goal the propagation of artificial intelligences at the expense of human survival.
Part the Third: Paranoia
This part is primarily putting the other two parts together: if "stateware" is a thing, and if software can be built that eventually outgrows the control of its designers, then it seems clear that states and economies are destined to arrive at singularity status before computer software. Economies are already difficult to understand, and have no stated goals, so an economic singularity is almost a given.
States, however, are more interesting. The goals of the United States Constitution, arguably the operating system for US stateware, are "to form a more perfect Union, establish Justice, insure domestic Tranquility, provide for the common defence, promote the general Welfare, and secure the Blessings of Liberty to ourselves and our Posterity". The Constitution was purposefully vague about the precedence and exact definitions of these goals.
At the time, of course, the hardware of the State was not fast enough for anyone to be concerned that any one of the goals would be given outsized importance, or that an exclusive definition of the wording would be settled upon. That is no longer the case. As the network of people inside the US has grown both in size and complexity, our capacity to produce more code (in the form of legislation, executive orders, and legal rulings) has exploded, creating a social software with unpredictable ramifications and unprecedented size.
The active coders here, the three branches of the government, are hobbled by design; the Founders didn't want any of them to control the coding process too much. This "un-agile" government doesn't have the power necessary to be fast enough to revise significant amounts of this code (just to create more), and the polarization of American politics is slowing the speed of government even further.
Beyond that, there is no guarantee that any elected official would want to revise US stateware. As politicians assume more power, it appears that the stateware co-opts their individual plans; those with reformist ideas are locked out of power, and those who make it through seem to change their minds.
Lastly, it doesn't seem like any small group of humans (even Congress is relatively small) can muster the cognitive power to steer this now-mammoth ship of state. The relations between organizations, lobbies, foreign powers, celebrities, systems, and economies are too complex. If the social singularity isn't here yet, it is howling at the door.
This all seems to drive to one point: we need AI assistance to manage the United States Government. I'm not suggesting we put things on auto-pilot, but we need machines that can make the connections we will inevitably fail to see. I don't believe the software exists yet, and this is a terrible realization to make.
Tuesday, April 12, 2011
Monday, March 14, 2011
The Need for a New Social Science
Memetics, Marketing, and Sociology can only get us so far.
I'm going to be very, very careful in this post; I don't know how to present this in a way that doesn't make me sound a little eccentric. I've been thinking a lot lately about memetics, the study of the transmission and propagation of ideas. These ideas have been compared alternately to genes and to biological viruses, spreading about hither and yon. There's been a bit of work done in the area of what ideas stick and what ideas don't, (consensus: various values are important, truth is only one of them) but there are just some huge, fat, gaping holes in the current state of the study of the propagation of ideas.
Gaping Hole #1: Knowing What We Can Know
First, and this is the scariest, we just have a very limited ability to view memes. We can see the spread of macro images on the internet, and identify that Ceiling Cat is probably sticking around for a while, whereas Rebecca Black's "Friday" is probably not. But we don't really know what to do with this information; the shape of the meme continues to elude us.
We will almost certainly never be able to predict what data will spread, and we will almost certainly never fully get a bead on what makes data "spreadable". We can get a general idea, perhaps, but it won't be an equation. Psychohistory is an idea in Asimov books, not a thing that will actually exist. (Oh, yeah, link. Sorry.)
This gaping hole is that memetic endeavors have largely been misguided because they've been strapped to a biological framework. Memes have no chromosomes, they cannot be put under a microscope, and they cannot be "sequenced".
Gaping Hole #2: Ideas and Behaviors
Second, there's not much in the literature by way of determining what, if any, is the difference between ideas that spread and the behaviors they leave behind. A piece of information that spreads might just annoy me by leaving "Never Gonna Give You Up" stuck in my head, or it might convince me to vote for Ralph Nader. There is a big difference between a meme and its behavioral payload.
Gaping Hole #3: The Myth of Measurement
Third, memetics tends toward the very abstract, and generally fails to measure anything at all, instead engaging in length thought experiments and exegeses on what theoretical construct is better for the task, in essence applying none of them empirically. Here's an article example, and it's ridiculously long.
The Solution
The solution is to actually conduct experiments on actual cultural transmission. Richard Dawkins, the founder of memetics, once noted that memetics had not yet found its Crick and Watson; it hadn't even found its Mendel. I think he was wrong. Plenty of people have come before, studying memetics before it was called memetics. Most importantly, to my mind, is the sociolinguist William Labov.
Labov tackled issues of how language change spread, what factors made someone likely to adapt their dialect to another, and so forth. One of his earliest and most famous studies was one of employees in three different department stores in New York City. He found that one meme, the tendency to "correct" the New York City tendency to "drop" the r-sound in certain words (his test was "fourth floor"). The meme was more likely to spread along socioeconomic lines, meaning that those in stores with higher-priced goods tended to include r-sounds in "fourth floor".
I know this isn't really much, as far as studies go, but it was a start. Labov discovered a number of patterns of linguistic change, and scores of sociolinguists after him have followed suit. Sociolinguistics may be a little tame compared to full-on memetics, as language traits are hardly as world-changing as religious and political beliefs, but it is a sufficient, if terribly overlooked start. Not much different from Mendel's plants, if you think about it.
I'm going to be very, very careful in this post; I don't know how to present this in a way that doesn't make me sound a little eccentric. I've been thinking a lot lately about memetics, the study of the transmission and propagation of ideas. These ideas have been compared alternately to genes and to biological viruses, spreading about hither and yon. There's been a bit of work done in the area of what ideas stick and what ideas don't, (consensus: various values are important, truth is only one of them) but there are just some huge, fat, gaping holes in the current state of the study of the propagation of ideas.
Gaping Hole #1: Knowing What We Can Know
First, and this is the scariest, we just have a very limited ability to view memes. We can see the spread of macro images on the internet, and identify that Ceiling Cat is probably sticking around for a while, whereas Rebecca Black's "Friday" is probably not. But we don't really know what to do with this information; the shape of the meme continues to elude us.
We will almost certainly never be able to predict what data will spread, and we will almost certainly never fully get a bead on what makes data "spreadable". We can get a general idea, perhaps, but it won't be an equation. Psychohistory is an idea in Asimov books, not a thing that will actually exist. (Oh, yeah, link. Sorry.)
This gaping hole is that memetic endeavors have largely been misguided because they've been strapped to a biological framework. Memes have no chromosomes, they cannot be put under a microscope, and they cannot be "sequenced".
Gaping Hole #2: Ideas and Behaviors
Second, there's not much in the literature by way of determining what, if any, is the difference between ideas that spread and the behaviors they leave behind. A piece of information that spreads might just annoy me by leaving "Never Gonna Give You Up" stuck in my head, or it might convince me to vote for Ralph Nader. There is a big difference between a meme and its behavioral payload.
Gaping Hole #3: The Myth of Measurement
Third, memetics tends toward the very abstract, and generally fails to measure anything at all, instead engaging in length thought experiments and exegeses on what theoretical construct is better for the task, in essence applying none of them empirically. Here's an article example, and it's ridiculously long.
The Solution
The solution is to actually conduct experiments on actual cultural transmission. Richard Dawkins, the founder of memetics, once noted that memetics had not yet found its Crick and Watson; it hadn't even found its Mendel. I think he was wrong. Plenty of people have come before, studying memetics before it was called memetics. Most importantly, to my mind, is the sociolinguist William Labov.
Labov tackled issues of how language change spread, what factors made someone likely to adapt their dialect to another, and so forth. One of his earliest and most famous studies was one of employees in three different department stores in New York City. He found that one meme, the tendency to "correct" the New York City tendency to "drop" the r-sound in certain words (his test was "fourth floor"). The meme was more likely to spread along socioeconomic lines, meaning that those in stores with higher-priced goods tended to include r-sounds in "fourth floor".
I know this isn't really much, as far as studies go, but it was a start. Labov discovered a number of patterns of linguistic change, and scores of sociolinguists after him have followed suit. Sociolinguistics may be a little tame compared to full-on memetics, as language traits are hardly as world-changing as religious and political beliefs, but it is a sufficient, if terribly overlooked start. Not much different from Mendel's plants, if you think about it.
Monday, January 10, 2011
Watch Your Hed
Headlines and ledes grab attention, and are often lies.
This is, of course, a follow-up to the "Update" segment of my previous post, in which I note that Fox News changed a headline from an inflammatory one to a less inflammatory one. Today's post is about why that sort of behavior is too little, too late.
A couple of posts ago, I talked about the "Information Problem", or how the surplus of content affects the way humanity makes decisions. I divided print and web material into three categories: content (less relevant, truth value unassignable), information (more relevant, but pre-interpreted) and data (most relevant, but difficult to understand without interpretation). News reporting, in general, should fit squarely into the "information" category; it should be pre-interpreted data reported from primary sources, and in general, should be unimpeachably true.
Now, that isn't always the case, and there are ways to report things that makes it appears as though the universe, not the medium, has a bias for or against a certain position. And I've written about that too. That needs to be taken care of, but in the end, each news user is going to have to strip bias from news. What's more devious, though, is the headline.
In a content economy, attention is currency. You can never make quick money by printing just the facts, reliably, and within a reasonable timeframe, even though that might be the most sustainable position. You're generally going to want to print news biased toward an audience, report on events that interest that audience, ignore things that don't interest the audience, and try to scoop everyone else all the time, even if you have unreliable or incomplete facts.
Even if you follow that formula, though, you still need to advertise. And the only advertisement that works on a news aggregator is the headline (and the lede, on occasion). So you have to make it count. In the end, if you make a statement that's shocking, you're more likely to get hits. It has been like this from the beginning of printed news. What's different in the aggregated news / 24-hour cable news ticker world is that far more headlines are read than articles, streaming at us as they do.
Unfortunately, shocking headlines are often a little bit misleading (or more than a little bit). This wasn't so much a problem before the Information Age, but now, each headline could easily be perceived as a tiny, unimpeachable fact. Which, of course, is inaccurate, and leads to misconstructed worldviews. Which messes with systems based on preferences and beliefs.
So, save the economy just a little bit, and try not to lie in your headline, even if you fix it in the article, or retract it later.
This is, of course, a follow-up to the "Update" segment of my previous post, in which I note that Fox News changed a headline from an inflammatory one to a less inflammatory one. Today's post is about why that sort of behavior is too little, too late.
A couple of posts ago, I talked about the "Information Problem", or how the surplus of content affects the way humanity makes decisions. I divided print and web material into three categories: content (less relevant, truth value unassignable), information (more relevant, but pre-interpreted) and data (most relevant, but difficult to understand without interpretation). News reporting, in general, should fit squarely into the "information" category; it should be pre-interpreted data reported from primary sources, and in general, should be unimpeachably true.
Now, that isn't always the case, and there are ways to report things that makes it appears as though the universe, not the medium, has a bias for or against a certain position. And I've written about that too. That needs to be taken care of, but in the end, each news user is going to have to strip bias from news. What's more devious, though, is the headline.
In a content economy, attention is currency. You can never make quick money by printing just the facts, reliably, and within a reasonable timeframe, even though that might be the most sustainable position. You're generally going to want to print news biased toward an audience, report on events that interest that audience, ignore things that don't interest the audience, and try to scoop everyone else all the time, even if you have unreliable or incomplete facts.
Even if you follow that formula, though, you still need to advertise. And the only advertisement that works on a news aggregator is the headline (and the lede, on occasion). So you have to make it count. In the end, if you make a statement that's shocking, you're more likely to get hits. It has been like this from the beginning of printed news. What's different in the aggregated news / 24-hour cable news ticker world is that far more headlines are read than articles, streaming at us as they do.
Unfortunately, shocking headlines are often a little bit misleading (or more than a little bit). This wasn't so much a problem before the Information Age, but now, each headline could easily be perceived as a tiny, unimpeachable fact. Which, of course, is inaccurate, and leads to misconstructed worldviews. Which messes with systems based on preferences and beliefs.
So, save the economy just a little bit, and try not to lie in your headline, even if you fix it in the article, or retract it later.
Sunday, January 9, 2011
The Tragedy of Tragedy
Just because something is sensational doesn't mean it's true.
After yesterday's tragic shooting in Tucson, which left six dead, including a nine-year-old and a federal judge, and several more wounded, including Democratic Congresswoman Gabrielle Giffords, the media responded with its traditional tragedy two-step program: first, report the facts; second, try to explain why the event occurred.
The first step was generally respectful, as media reporting of tragedy usually is. The second step was occasionally disturbing in that many linked Sarah Palin's trigger-happy metaphors to the terrible outburst.
Let me be clear before I go any further: I find Palin's ideology, rhetoric, and image utterly execrable in every way. I cannot defend her early departure from Juneau, and find her and her family's opportunistic reality spotlight-hogging a frightening example of a possible future direction of American politics.
And while it's arguable that Arizona's hyper-conservative politics played a role in alleged shooter Jared Loughner's timing and target, it seems unlikely that they made him a crazed shooter.
Tucson Weekly has an interesting piece on Mr. Loughner's personal internet presence, which makes it pretty clear that he was likely to shoot and kill in one venue or another at some point in time. It also happens to be the case that his home was a short walk from the Safeway where the attack occurred.
The nature of the attack, Mr. Loughner's personality, and his previous communications seem to put this incident in the same category as the Columbine shootings, not the JFK assassination.
But this doesn't match the agenda of some commentators. It appears as though some wish the shooter had been able to be clearly tied to the admittedly ridiculous, extreme rhetoric that has become the stock of American political discourse. (egregious example) Upon examining the evidence, this appears not to be the case. The man had extreme, iconoclastic political beliefs, and appears to have been on a rampage, rather than a mission.
To ignore the evidence and provide a false connection to an unwanted person or ideology disrespects the memory of those who died in this terrible tragedy. There are enough valid arguments against Ms. Palin's stances; making specious accusations is unnecessary.
(Update: Fox News accuses "many on the American Left" of misusing the tragedy, and changed the headline of this story from 'Dems Blame Rhetoric for Shooting' to 'Tragedy Inspires Political "Cheap Shots"'. Stay classy, Fox.)
After yesterday's tragic shooting in Tucson, which left six dead, including a nine-year-old and a federal judge, and several more wounded, including Democratic Congresswoman Gabrielle Giffords, the media responded with its traditional tragedy two-step program: first, report the facts; second, try to explain why the event occurred.
The first step was generally respectful, as media reporting of tragedy usually is. The second step was occasionally disturbing in that many linked Sarah Palin's trigger-happy metaphors to the terrible outburst.
Let me be clear before I go any further: I find Palin's ideology, rhetoric, and image utterly execrable in every way. I cannot defend her early departure from Juneau, and find her and her family's opportunistic reality spotlight-hogging a frightening example of a possible future direction of American politics.
And while it's arguable that Arizona's hyper-conservative politics played a role in alleged shooter Jared Loughner's timing and target, it seems unlikely that they made him a crazed shooter.
Tucson Weekly has an interesting piece on Mr. Loughner's personal internet presence, which makes it pretty clear that he was likely to shoot and kill in one venue or another at some point in time. It also happens to be the case that his home was a short walk from the Safeway where the attack occurred.
The nature of the attack, Mr. Loughner's personality, and his previous communications seem to put this incident in the same category as the Columbine shootings, not the JFK assassination.
But this doesn't match the agenda of some commentators. It appears as though some wish the shooter had been able to be clearly tied to the admittedly ridiculous, extreme rhetoric that has become the stock of American political discourse. (egregious example) Upon examining the evidence, this appears not to be the case. The man had extreme, iconoclastic political beliefs, and appears to have been on a rampage, rather than a mission.
To ignore the evidence and provide a false connection to an unwanted person or ideology disrespects the memory of those who died in this terrible tragedy. There are enough valid arguments against Ms. Palin's stances; making specious accusations is unnecessary.
(Update: Fox News accuses "many on the American Left" of misusing the tragedy, and changed the headline of this story from 'Dems Blame Rhetoric for Shooting' to 'Tragedy Inspires Political "Cheap Shots"'. Stay classy, Fox.)
Wednesday, January 5, 2011
The Information Problem
We have too much content, not enough information, and surprisingly little data.
The great futurist Alvin Toffler predicted a world with relevant data so profuse that information overload would be the inevitable result. He was wrong. The 21st Century does not have a problem with information overload. It has a problem with Content Overload.
You might think that I'm splitting semantic hairs in order to make a polarizing statement, but perhaps it would help my case to point out that the surfeit of content on the internet is of varying informational value. Quite a bit of it is opinion, and much of the stuff that is purportedly informational is of questionable truth value.
It's important to make a distinction between content, information, and data. While I'm not suggesting that there's anything inherent in each of these words that usefully distinguishes them, it is clear there are definitely three different phenomena which can be distinguished, regarding the stuff we take into our brains.
Content is what I'll call all print and digital "stuff" in general. This definition of content would include movie reviews, lolcats, sports scores, et cetera. Information is content which can be assigned a truth value. That gets tricky, because an opinion piece is clearly not information in and of itself, but may contain information (or misinformation). Data is a special kind of information, that appears in tabular or numeric form. The universe of Content and Data can be seen as a scale that runs from the lowest density of fact to the highest.
Data is not always foolproof, as Charles Seife identifies in his data-in-journalism analysis, Proofiness. A lot of figures are either half-right, made to appear to support erroneous conclusions, or just plain made up. Of course, that makes any informational statements based on these data unsupported, and opinions formed around this information, misinformed. To hear Seife tell it, there are a lot more errant conclusions out there than accurate ones--which leads me to believe the following:
Nearly everyone is wrong most of the time.
The most logical solution to being wrong is to check facts to make sure they are accurate before basing opinions and decisions on them. If this were possible, it would certainly make for a much better world. Unfortunately, verifiable fact is hard to come by. In many cases, scientific studies have not been done (or cannot be done). Worse, sometimes results of different studies conflict. Exercise science and nutrition seem especially prone to this: one month, authorities say it's essential to focus mostly on cardio, the next, weights are the thing. Fat used to be the killer, now carbs are. Any attempt to be consistently right, even most of the time, is probably doomed to failure.
Unfortunately for us, the number of decisions we have to make every day appears to be increasing. From financial choices, like bank selection and the use of credit cards, to consumer choices about everything from insurance to running shoes. Further, it seems as though the options we have are also increasing, adding even more difficulty to the task of human decision-making, and making the human life experience a multi-decade pratfall.
As a shortcut to making decisions, we tend to lump like decisions together, as a pattern. We may then seek to explain these patterns with a narrative. For example, my wife loves to shop with coupons. Her decision-making pattern is that if we don't absolutely need a product right away, she doesn't buy it at full price, or even close, ever. The narrative behind this pattern is the idea that companies engage in promotional deals to catch unwary shoppers off their guard, but that by putting in a little extra effort, you can subvert these deals to make your shopping exceptionally cheap.
The progression in my wife's example is from individual decisions, to patterns, to explanatory narratives. This is a practical approach to decision making. The reverse is the ideological approach: to take the narrative and impose it on practices and individual decisions. There is nothing wrong with the ideological approach if you use it to govern decisions made without a thought for immediate outcomes. Ethical decisions, for example. One deals honestly because the narrative of "Honesty is the best policy" embodies a state which the decision-maker wishes to attain, not because it yields a direct benefit. When you use the ideological approach and expect a specific result to a specific decision, things get crazy.
This is because narratives are based on, again, patterns made from lots of little decisions, which the narratives then explain. When the method by which the narrative was built is unknown, the truth value of the narrative is obscured. Using content in the place of information or data in order to make decisions is terribly dangerous. Using unverified information as if it were verified is terribly dangerous. Analyzing data incorrectly is terribly dangerous. And yet, this is what is happening in decision-making all over the world, from households to governments and beyond.
Data in and of itself is incredibly powerful. Bringing up Seife and Proofiness again, it's clear that attaching a number to a fact is a shortcut to credibility. But humans are notorious for misreading and abusing data, and creating narratives based on these tortured numbers. Here's an example debunking a mostly harmless myth about pet ownership.
What I'm driving at is this: humanity might make better decisions if we would take what's reported as news and science, et cetera, with a grain of salt. This proves exceptionally difficult, because in reacting to one incorrect or incomplete narrative, we often create an opposing narrative that is just as incorrect or incomplete. There may not be a solution to the "Information Problem", but I propose that you, Dear Reader, go conduct your own studies to verify my narrative, and live by holding out judgment until proof is overwhelming.
The great futurist Alvin Toffler predicted a world with relevant data so profuse that information overload would be the inevitable result. He was wrong. The 21st Century does not have a problem with information overload. It has a problem with Content Overload.
You might think that I'm splitting semantic hairs in order to make a polarizing statement, but perhaps it would help my case to point out that the surfeit of content on the internet is of varying informational value. Quite a bit of it is opinion, and much of the stuff that is purportedly informational is of questionable truth value.
It's important to make a distinction between content, information, and data. While I'm not suggesting that there's anything inherent in each of these words that usefully distinguishes them, it is clear there are definitely three different phenomena which can be distinguished, regarding the stuff we take into our brains.
Content is what I'll call all print and digital "stuff" in general. This definition of content would include movie reviews, lolcats, sports scores, et cetera. Information is content which can be assigned a truth value. That gets tricky, because an opinion piece is clearly not information in and of itself, but may contain information (or misinformation). Data is a special kind of information, that appears in tabular or numeric form. The universe of Content and Data can be seen as a scale that runs from the lowest density of fact to the highest.
Data is not always foolproof, as Charles Seife identifies in his data-in-journalism analysis, Proofiness. A lot of figures are either half-right, made to appear to support erroneous conclusions, or just plain made up. Of course, that makes any informational statements based on these data unsupported, and opinions formed around this information, misinformed. To hear Seife tell it, there are a lot more errant conclusions out there than accurate ones--which leads me to believe the following:
Nearly everyone is wrong most of the time.
The most logical solution to being wrong is to check facts to make sure they are accurate before basing opinions and decisions on them. If this were possible, it would certainly make for a much better world. Unfortunately, verifiable fact is hard to come by. In many cases, scientific studies have not been done (or cannot be done). Worse, sometimes results of different studies conflict. Exercise science and nutrition seem especially prone to this: one month, authorities say it's essential to focus mostly on cardio, the next, weights are the thing. Fat used to be the killer, now carbs are. Any attempt to be consistently right, even most of the time, is probably doomed to failure.
Unfortunately for us, the number of decisions we have to make every day appears to be increasing. From financial choices, like bank selection and the use of credit cards, to consumer choices about everything from insurance to running shoes. Further, it seems as though the options we have are also increasing, adding even more difficulty to the task of human decision-making, and making the human life experience a multi-decade pratfall.
As a shortcut to making decisions, we tend to lump like decisions together, as a pattern. We may then seek to explain these patterns with a narrative. For example, my wife loves to shop with coupons. Her decision-making pattern is that if we don't absolutely need a product right away, she doesn't buy it at full price, or even close, ever. The narrative behind this pattern is the idea that companies engage in promotional deals to catch unwary shoppers off their guard, but that by putting in a little extra effort, you can subvert these deals to make your shopping exceptionally cheap.
The progression in my wife's example is from individual decisions, to patterns, to explanatory narratives. This is a practical approach to decision making. The reverse is the ideological approach: to take the narrative and impose it on practices and individual decisions. There is nothing wrong with the ideological approach if you use it to govern decisions made without a thought for immediate outcomes. Ethical decisions, for example. One deals honestly because the narrative of "Honesty is the best policy" embodies a state which the decision-maker wishes to attain, not because it yields a direct benefit. When you use the ideological approach and expect a specific result to a specific decision, things get crazy.
This is because narratives are based on, again, patterns made from lots of little decisions, which the narratives then explain. When the method by which the narrative was built is unknown, the truth value of the narrative is obscured. Using content in the place of information or data in order to make decisions is terribly dangerous. Using unverified information as if it were verified is terribly dangerous. Analyzing data incorrectly is terribly dangerous. And yet, this is what is happening in decision-making all over the world, from households to governments and beyond.
Data in and of itself is incredibly powerful. Bringing up Seife and Proofiness again, it's clear that attaching a number to a fact is a shortcut to credibility. But humans are notorious for misreading and abusing data, and creating narratives based on these tortured numbers. Here's an example debunking a mostly harmless myth about pet ownership.
What I'm driving at is this: humanity might make better decisions if we would take what's reported as news and science, et cetera, with a grain of salt. This proves exceptionally difficult, because in reacting to one incorrect or incomplete narrative, we often create an opposing narrative that is just as incorrect or incomplete. There may not be a solution to the "Information Problem", but I propose that you, Dear Reader, go conduct your own studies to verify my narrative, and live by holding out judgment until proof is overwhelming.
Monday, December 27, 2010
As the Sands of the Hourglass...
This is a nice graph showing where people came into and went out of my Gmail chat life. That gray line is the first day I used the service, which as you'll not is NOT in the middle of 2006 (tee hee, previous post. tee hee.)


Personally, I think this is an incredibly telling graph. It's taking, again, my top nine gchat buddies and showing how the volume of our interaction changes. For example, KRS shows a very solid line at the beginning, showing frequent interaction at the beginning of this period, but tapers off. RPM starts off as a very casual gchat friend, but gains in intensity near the end. CRS is an interesting latecomer; she's my wife. She was given a Gmail account by ASU when she was accepted to the doctoral program, and I switched her personal email over to Gmail as well. We didn't start chatting until we were essentially engaged.
You'll note that she still makes it into my top nine. This is because of the size of those early chats. Unfortunately, this graph is not weighted; you get one dot every day we converse, no matter how long or short. I'm still learning R, and am trying to find out exactly how one goes about changing plotting colors according to the values of a variable. (Rather, I know how to do this for larger symbols, but not for dots).
A note on the three-letter codes. In the interest of privacy, they are generally not the current complete initials of the person listed. However, in order to make it readable for me, they are pretty close. If you're listed here, and you want me to change the initials to protect your privacy on the internet, please let me know. I'm not Mark Zuckerberg, you know.
You'll note that she still makes it into my top nine. This is because of the size of those early chats. Unfortunately, this graph is not weighted; you get one dot every day we converse, no matter how long or short. I'm still learning R, and am trying to find out exactly how one goes about changing plotting colors according to the values of a variable. (Rather, I know how to do this for larger symbols, but not for dots).
A note on the three-letter codes. In the interest of privacy, they are generally not the current complete initials of the person listed. However, in order to make it readable for me, they are pretty close. If you're listed here, and you want me to change the initials to protect your privacy on the internet, please let me know. I'm not Mark Zuckerberg, you know.
Saturday, December 18, 2010
The Size of Shakespeare, or, A Comedy of Errors
So.
It's been a little while, and I haven't neglected you, three blog readers. I've been working on a project in order to blow your collective mind, or at least give it a little something to chew on.
Specifically, I have delved deep into the realms of my Gmail chat logs and have begun to discover: data. Oh man, the trip this has been. And it's not over. There will be charts, there will be graphs, AND! there may be PODCASTING.
I intend to give you the tidbits I have learned in chewable form, piece by piece. Today's episode is: why you should examine your data thoroughly before you make any conclusions.
I wrote a Perl script to turn my wad of uncooked data into a delicious patty; it returned the size of each individual chat file along with other important stats. In the statistical scripting language R, I discovered that the sum total of chat content produced was, in a word, ridiculous. I did some calculations and made a graph that looks a little something like this:

Yes, it appeared that even just my most chatty friend had produced with me a larger corpus of work, bytewise, than Bill Shakespeare himself (the Bard wrote about 5 Mb worth). Sweet mercy. Note: I have been using Gmail's Chat client, and occasionally Google Talk, since the former launched in the middle of 2006.
The problem with this graph is that it's not actually accurate. When I examined the files a little further, I realized most of them looked like this:
Uh oh. So, of course, I had to write another Perl script. I found that code accounted for roughly 77 percent of the content created by Gmail chat. Here's the revised graph:

Tada! I think the most interesting development from this graph is quite simply that, even with the code stripped from the chats, a few of my friends and I have produced an entire corpus of text.
In the next few months, I'll look at some of the sociological implications of the date and size data, and then all the way into the the textual aspect of the transcripts. Analyzing the text itself should be tremendously interesting.
[Note: I changed some things as other pursuits have prevented me from diving into the actual text. Someday...someday.]
It's been a little while, and I haven't neglected you, three blog readers. I've been working on a project in order to blow your collective mind, or at least give it a little something to chew on.
Specifically, I have delved deep into the realms of my Gmail chat logs and have begun to discover: data. Oh man, the trip this has been. And it's not over. There will be charts, there will be graphs, AND! there may be PODCASTING.
I intend to give you the tidbits I have learned in chewable form, piece by piece. Today's episode is: why you should examine your data thoroughly before you make any conclusions.
I wrote a Perl script to turn my wad of uncooked data into a delicious patty; it returned the size of each individual chat file along with other important stats. In the statistical scripting language R, I discovered that the sum total of chat content produced was, in a word, ridiculous. I did some calculations and made a graph that looks a little something like this:

Yes, it appeared that even just my most chatty friend had produced with me a larger corpus of work, bytewise, than Bill Shakespeare himself (the Bard wrote about 5 Mb worth). Sweet mercy. Note: I have been using Gmail's Chat client, and occasionally Google Talk, since the former launched in the middle of 2006.
The problem with this graph is that it's not actually accurate. When I examined the files a little further, I realized most of them looked like this:
Uh oh. So, of course, I had to write another Perl script. I found that code accounted for roughly 77 percent of the content created by Gmail chat. Here's the revised graph:

Tada! I think the most interesting development from this graph is quite simply that, even with the code stripped from the chats, a few of my friends and I have produced an entire corpus of text.
In the next few months, I'll look at some of the sociological implications of the date and size data, and then all the way into the the textual aspect of the transcripts. Analyzing the text itself should be tremendously interesting.
[Note: I changed some things as other pursuits have prevented me from diving into the actual text. Someday...someday.]
Subscribe to:
Posts (Atom)