Tuesday, July 29, 2008

Wikipedia's Kittens

Just found a great blog
http://wikip.blogspot.com/search/label/sociology

Quoting:

"....a generalized unit of contributor motivation called a kitten.

1 kitten = the amount of motivation needed to get 1 person to spend 1 minute trying to improve an article

We can say, quite literally, that Wikipedia runs on kittens. In fact, entrepreneurs discover this every day when they try to start a "crowdsourcing" site and nobody shows up. So, what generates kittens? Foremost, it's the possibility of someone else learning from what you wrote -- not just immediately, but at any time in the future"

Kittens are born when there is a perception that the words one write will survive for some time.... something like this, to go by

Number of views on a given day = (Number of views per day).(Chance of surviving one day) ^ (Number of Days that have passed)
With a 1-in-ten-thousand chance of being destroyed each day, the article will rack up exactly seven million views over its lifetime.


Like the author says - thats a LOT of kittens!

Friday, July 25, 2008

Why Wikipedia Succeeded


Larry Sanger (Wikipedia's cofounder)'s take on why Wikipedia succeeded.
Although rather old (2005), the feature has some great insights.

http://features.slashdot.org/article.pl?sid=05/04/19/1746205&tid=95


In short, these are the factors
  1. Open content license. We promised contributors that their work would always remain free for others to read. This, as is well known, motivates people to work for the good of the world--and for the many people who would like to teach the whole world, that's a pretty strong motivation.
  2. Focus on the encyclopedia. We said that we were creating an encyclopedia, not a dictionary, etc., and we encouraged people to stick to creating the encyclopedia and not use the project as a debate forum.
  3. Openness. Anyone could contribute. Everyone was specifically made to feel welcome. (E.g., we encouraged the habit of writing on new contributors' user pages, "Welcome to Wikipedia!" etc.) There was no sense that someone would be turned away for not being bright enough, or not being a good enough writer, or whatever.
  4. Ease of editing. Wikis are pretty easy for most people to figure out. In other collaborative systems (like Nupedia), you have to learn all about the system first. Wikipedia had an almost flat learning curve.
  5. Collaborate radically; don't sign articles. Radical collaboration, in which (in principle) anyone can edit any part of anyone else's work, is one of the great innovations of the open source software movement. On Wikipedia, radical collaboration made it possible for work to move forward on all fronts at the same time, to avoid the big bottleneck that is the individual author, and to burnish articles on popular topics to a fine luster.
  6. Offer unedited, unapproved content for further development. This is required if one wishes to collaborate radically. We encouraged putting up their unfinished drafts--as long as they were at least roughly correct--with the idea that they can only improve if there are others collaborating. This is a classic principle of open source software. It helped get Wikipedia started and helped keep it moving. This is why so many original drafts of Wikipedia articles were basically garbage (no offense to anyone--some of my own drafts were sometimes garbage), and also why it is surprising to the uninitiated that many articles have turned out very well indeed.
  7. Neutrality. A firm neutrality policy made it possible for people of widely divergent opinions to work together, without constantly fighting. It's a way to keep the peace.
  8. Start with a core of good people. I think it was essential that we began the project with a core group of intelligent good writers who understood what an encyclopedia should look like, and who were basically decent human beings.
  9. Enjoy the Google effect. We had little to do with this, but had Google not sent us an increasing amount of traffic each time they spidered the growing website, we would not have grown nearly as fast as we did. (See below.)

Thursday, July 24, 2008

Defining Knowledge

Came across this quite elucidating (yes, I've used the word "elucidating"!) definition, explanation about Knowledge...

http://www.systems-thinking.org/kmgmt/kmgmt.htm

"
    • A collection of data is not information.
    • A collection of information is not knowledge.
    • A collection of knowledge is not wisdom.
    • A collection of wisdom is not truth.

The idea is that information, knowledge, and wisdom are more than simply collections. Rather, the whole represents more than the sum of its parts and has a synergy of its own.

We begin with data, which is just a meaningless point in space and time, without reference to either space or time. It is like an event out of context, a letter out of context, a word out of context. The key concept here being "out of context." And, since it is out of context, it is without a meaningful relation to anything else. When we encounter a piece of data, if it gets our attention at all, our first action is usually to attempt to find a way to attribute meaning to it. We do this by associating it with other things. If I see the number 5, I can immediately associate it with cardinal numbers and relate it to being greater than 4 and less than 6, whether this was implied by this particular instance or not. If I see a single word, such as "time," there is a tendency to immediately form associations with previous contexts within which I have found "time" to be meaningful. This might be, "being on time," "a stitch in time saves nine," "time never stops," etc. The implication here is that when there is no context, there is little or no meaning. So, we create context but, more often than not, that context is somewhat akin to conjecture, yet it fabricates meaning."

Monday, July 21, 2008

Model for Viral Growth

One of the things in my mind is to create a sufficiently accurate model to predict viral growth - While digging, I came across this very interesting blog post
http://lsvp.wordpress.com/2008/03/10/an-excellent-excel-model-of-viral-growth/
with a link here:
http://andrewchen.typepad.com/andrew_chens_blog/2008/03/facebook-viral.html?cid=106420002#comment-106420002

Some others
http://www.insidefacebook.com/2007/07/17/predicting-growth-with-appaholic/
http://www.skelliewag.org/the-butterfly-growth-model-224.htm



Friday, May 16, 2008

Watson and Crick


On Feb. 28, 1953, Francis Crick walked into the Eagle pub in Cambridge, England, and, as James Watson later recalled, announced that "we had found the secret of life."....

Slightly off-topic, but worth it for the 1959 picture alone

Full TIME article here.


Tracking Memes in the Infosphere

Infosphere is a term used since the 1990s to speculate about the common evolution of the Internet, society and culture. It is a neologism composed of information and sphere. More about its origins here.
The difficulties with memetics are many - and it has been bogged in controversy since a long time. One of the main problems is how to isolate a "meme" ; What IS a meme, anyway?
Start here, but don't expect to find a definite answer - it's not there yet!
The whole discipline (if it can be called that!) is in a similar state to that of genetics in the 1950s. What was a gene? It took Watson and Crick to come up with the molecular structure - the double helix - of DNA, before genetics really took off.
The meme sounds very vague when defined as "a unit of cultural information". An abstract but precise mathematical notion is required - perhaps it can be found in information theory?
So, not even knowing precisely what a meme is, how are we supposed to track them, and build theories around them, and maybe even try and predict stuff with them?

The meme-tracking problem.....
Some possible routes:
Web publication volume and search trends

Hitwise tracks search data of all major search engines, including Google.
Google Trends also tells us the history of search volume on keywords, i.e. how many searches were executed on these keywords over time. This sounds like a good indicator of what's "hot". I am not entirely sure it is a truly accurate indicator of "meme" though. For example, Breaking News of any kind will cause a peak in News Coverage - But News is not Meme!

Published Reports by Professional Market Research Firms:
E.g The Harris Interactive Annual RQ™ study, conducted yearly since 1999, assesses the reputation of the 60 most visible companies in the United States, as perceived by the general public. Changes in reputation are what we want to learn. Perceptions of 'brand' by consumer are one part of it, of course - but again, this is not quite enough or good enough data.

WOM data
: 'Positive word of mouth' data can be sourced from places like Keller and Fay's Talk-Track, a research service that tracks consumer conversations via a weekly survey
sample of 700 consumers aged 13+. Online brand mentions data can be sourced from a service like Nielsen Buzzmetrics. It searches the net for mentions of specific words or phrases on discussion boards, blogs or other places where consumers communicate online.

Advertising Spending: Weekly advertising spending data for television and national magazines can be had from Nielsen's Monitor + database. Online advertising spending can be obtained from AdRelevance, owned by Nielsen Netratings.

Agencies like ComScore MediaMetrix track website visits through a representative panel of 2 million users - another valuable storehouse of data but not "ready made" for meme-tracking by any means.

One problem for any non-US study is the possible difficulty in getting location-specific non-US data.

Research design incorporated open-ended, discoveryoriented in-depth interviews are another option, with the obvious limitation of being impossibly difficult to scale up, or even trust.

The meme-tracking space is supposed to be HOT round about now...
http://www.techcrunch.com/2006/02/04/a-look-at-the-memeorandum-killers/
It’s not easy to define this space.....but as Alex Barnett says, these are not meme-trackers - there seems to be no real meme-tracker around. I agree.
"Memeorandum, Megit and Chuquet are not 'meme' trackers. They are news trackers. Or tittle-tattle trackers. Or gossip trackers. Again, generally speaking, there are no 'memes' being tracked at these sites". I especialy like his comment that "The idea that these are 'memetrackers' is actually quite a good example of a meme."
Which brings us back to the question: How does one define a meme, at least in a way for a bot can measure it? (I think if you can define something that a machine can understand then you have done a good job at the definition!)

Some interesting papers, using innovative means to find and interpret data can be found in the Journal of Advertising Research, December 2007. I am going to fish around for those again.

Hint : Technology Memes

While learning quite a lot about Project Management from DavidT, I quite accidentally ran across his post on adoption .

"There are a couple of largely accepted theories that model or predict technology lifecycle and adoption patterns:- The Diffusion of Innovations theory offers a model for how a given technology gets accepted and spreads through markets. Its central point is that technologies spread by gradually addressing the needs of 4 types of users: innovators, early adopters, the early majority, and the late majority (a fifth category, the laggards, might just never get it)- The Technology Acceptance Model (TAM) offers some prediction to End User adoption. The key concept here is that individual users adopt a given technology based on its perceived usefulness and its perceived ease of use.To my knowledge, there isn't an established theory or framework that models evolution trends of Technologies.When looking at the history and evolution of web services, we seem to be in front of species that are spreading, adapting, and diverging much like finches in the Galapagos.The immediate thought that then comes to mind is whether Darwin's Theory of Evolution has some or any relevance to Technology.The theory of evolution defines three basic mechanisms of evolutionary change:. Natural Selection is a process by which traits that are more useful in a given environment become more common over time (because they give better chances of survival), while traits that are harmful become rarer. Gene Flow is the exchange of genes within and between populations, which translates in traits being transferred between populations and species.. Genetic Drift is a purely random shift of the frequency of traits within a population - traits become more or less common in a population because of the long-term statistical effect of the random distribution of genes in each generationHow could these mechanisms apply to technology?- Natural Selection is probably the mechanism most relevant to technology trending.The fitter a technology is to the needs of its market, the more likely it is to stick around, and potentially supersede other technologiesThis is why PCs are more likely to be found today than mainframes, why java is more often used than Fortran, and why soap-based web services have replaced xml-rpc.- Gene Flow is also common in the tech field (although we'd probably want to call it something else).Features and concepts are constantly exchanged between complementary or competing technologies.That's how C# got a memory garbage collection mechanism similar to the one in java, and how row-level locking made it in MS SQL Server after years of Oracle claiming it as a key differentiator.Gene flow is also at the root of hybridization, where traits of different species end-up being combined. This is what might be truly going on right now with REST - which is applying concepts of simpler web protocols, most notably HTTP and RSS, onto Web Services.- Genetic Drift seems at first least relevant to the tech field, but might in fact be the most interesting bit.The core concept in genetic drift is that the random distribution of genes in each generation can have a long-term effect on the frequency of traits in a population (because of the statistical law of large numbers, genetic drift is less likely to occur in large population than in smaller ones).What, if anything, could have a similar impact in the evolution of technologies? What type of mechanisms, if any, can have an effect on the evolution and adoption of a technology, without being connected to its intrinsic fit or value?Obviously there are a lot more forces that dictate the success or demise of technologies than just their core virtues.A strategic alliance with IBM propelled MS-DOS into market dominance; technology companies like Oracle spend millions trying to influence the market; and there is a whole ecosystem of media, analysts, and venture capitalists who strive on generating buzz (PointCast or Twitter come to mind).Who knows - if LISP had been able to be more hip, we might all be using more parenthesis today."

Again, my "meme theory" interest (obsession?) means I cannot but help notice 'one more case that fits'. Technology memes playing out their game of survival in the world....
The question is: How can I model, simulate, and more importantly - validate, prove....and Predict the future?