Showing posts with label stumble. Show all posts
Showing posts with label stumble. Show all posts

Wednesday, December 17, 2008

Image Evolution

Ten days ago... No, wait. That's something else.

Eight days ago, I stumbled onto Genetic Programming: Evolution of Mona Lisa. I thought it was pretty awesome. Today, I stumbled upon this simulation of the same idea you can run in your browser. Doesn't work in IE, but I really hope anyone reading this isn't using IE (any version) as your default browser anyway. Right now, mine is at 1,077/17,000. That means out of 17,000 “mutations” 1,077 have been an improvement over the previous best fit. I just hit 90% fit. Which looks like 005874.jpg on the example shots from the original post. That leads me to believe the offline version is a bit more efficient than the online version, since it's taken about 3 times the number of mutations to get there. But that's to be expected.

I think the ability to watch it develop over time in your own browser hammers home the idea a bit better than the static example provided in the first post from Roger Alsing. But I think there's still room for improvement. For one thing, Roger's program makes it look like it would take almost 1,000,000 generations to get a decent replication of the low res detail of the Mona Lisa he used as his example. That's a necessary abstraction to get this kind of simulation to run on a computer. In real evolution, each generation produces several nodes, each with their own mutations. The node tree would fork out to very huge, very quickly, due to exponential growth.

But applying survival of the fittest to the tree structure would get us to the best fit much more quickly. The version that runs in a browser addresses this somewhat with the x/y display. I'm currently at 91.10% with 1238/23000. On average so far, it's taken about 18.5 generations to find a better fit. If we could fork that into a node tree, even a simple node tree of 2 child nodes per parent node, we'd hit the first improvement in about 5 generations. Actually, we'd almost hit the first 2 improvements in 5 generations. After 5 more generations, we'd be hitting quantum leaps where over 2,000 mutations are tried. And that number doubles every generation.

Ok, I'm proving to not be a strong enough math geek to really get these thoughts out of my head in a way that makes sense. But the way these programs are running is a bit like selection sort whereas the node tree approach is more like heap sort. In another browser tab, I just hit ~94% fit after ~ 50,000 mutations (but only 1900 improvements). Selection sort has an efficiency rating of n2. So if it's taken ~50,000 steps to reach ~94%, that means n is about 224. Heap sort runs at n(log n), so to reach ~94% using that method would only take about 527 steps rather than ~50,000. That's a closer fit to the number of true, natural generations it would take to get the same sort of results through evolution.

A better approach to visualizing this sort of thing would allow for multiple source images and create a node tree instead of using the linear approach. The multiple source images would allow for a simulation of speciesization. The node tree would be less of an abstraction from the natural processes (although still astronomically simplified in comparison) and give a more natural indication of true “generations” required to reach a goal. Ideally, the number of descendant nodes would depend on past history of improvements. So “blood lines” that have produced a high number of improvements in the past would be more fruitful. Those with poor histories would get fewer chances, and eventually die out. There would also need to be a threshold after which a given line is only compared to the source image it seems to be naturally drifting towards. This would simulate the migration into more specialized environments over time. It would also cut down on the processing power required to make it all run, but such a system would experience exponential growth and would quickly overrun any processor or pile of RAM currently available. As abstract as such a program would still be, it would require hardware we won't see until quantum computing becomes a reality.

Sunday, December 14, 2008

The Social Side of Social Bookmarking

There's a decent amount of discovery power available through Ma.gnolia as well. I still think Stumble Upon does it better. But it would be silly to not explore the 2nd best tool available for the job. Ironically, but doing so, I was quickly reminded of the problems with default tagging in SU. I pulled up a couple of recent bookmarks from Jeffry Zeldman.

The first is a blog post from Simon Clayson on feeding IE6 a basic style sheet using the sort of techniques that were once common for targeting Netscape Navigator 4 with a set of specific, dumbed down styles while simultaneously protecting NN4 users from that browsers botched implementation of the majority of CSS which was safe to show to less craptastic browsers. Now NN4 is little but a ghost to haunt the nightmares of us old school CSS scribes and IE6 is now the crappiest browser still in common usage. I'll probably spend the rest of December debugging the redesign in IE6. Had I found this idea a year ago, I probably would have served IE6 a very simple style sheet and skipped the debugging. In all honesty, even at this stage it may be less work to implement these ideas rather than try to “fix” IE6.

So anyway, I thought this was a potentially useful technique, so I thumbed it up. This didn't pull up the form for submitting new content to SU, so I knew someone else had already submitted this particular link. This gave me a great opportunity to see what the default tag would be. jackosborne says this page is primarily about “graphic-design”. It deals exclusively with serving specific CSS code targeted at a specific web browser. I can think of at least half a dozen tags more useful for this content than “graphic-design”. But the current SU system gives too much power to the person submitting the content. Jack's actually got more stumbles tagged “web-design” (54) than “graphic-design” (45), but apparently that's due to other people's default tags on the pages he is thumbing up. Looking at his discoveries, he's also submitted this article on Five CSS Design Browser Differences I Can Live With by Andy Clarke and Using jQuery for Background Image Animations from Snook.ca as “graphic-design”. Maybe that tagging scheme server Jack well. It makes SU virtually worthless for me when it comes to organizing and retrieving the resources I discover through it.

The other page I discovered via Zeldman is Western Civ's guide to CSS browser support. Again, this page deals exclusively with CSS and web browsers, so for my purposes it would be pretty easy to tag. It was submitted by SU user kancerman uh...wow, 3 years ago. If I'm reading this right, I'm only the 10th person to thumb this up in those 3 years. That could be because it was submitted into the category “internet-tools”. For me, that category is better suited for things like online mortgage calculators or WriteBoard. But due to the way kancerman submitted this page, “internet-tools” is the default tag. Now I can look at his entry for this page and see that his 2nd tag is in fact “CSS”, but since that's the 2nd tag on his entry, it holds no bearing for how it is tagged by default when I thumb up the page.

Maybe this problem in the design of SU is worse than I thought. Not only does the default tagging scheme make it harder for me to go back and look up stuff I have previously thumbed up without bothering to write a review and/or manually tag myself. But it also seems to have a negative impact on the effectiveness of SU to function as a discovery engine. How many times have I found a page via means other than SU, thumbed it up, didn't see the new content submission form pop up, assumed whoever beat me to the punch on submitting the content at least submitted it properly, and went on my way? How often does the average SU user do that? One thing I've noticed since I started paying attention to the default tagging scheme in SU is how often I see content that is submitted into the wrong category. If I have found this content via SU, then I can use the “report last stumble” feature. But that only works if I stumble into something in one of my defined interests that really should be tagged as some other of my defined interest. If someone submits a CSS gallery as “photography” for example. But if I get to that page without being referred there by SU, there's no way for me to bring the miscategorization to the attention of whoever addresses such things. That is most likely to happen if someone submits content that should fall within one of my defined interests as pertaining to a topic of interest that isn't on my list.

Oh look, this jQuery plug-in has been submitted under “alternative-medicine”

There's no way I can do anything about that. All I can do is tag it properly within my own account. But since very few web designers are going to be stumbling through the alternative medicine category (then again, maybe I assume too much), and very few people looking for alternative medicine information will give a rat's ass about a jQuery plug-in, very few people who care about that content will ever stumble into that content. I can't even resubmit it. Once a page is submitted, all I can do is tag and review it myself. In effect, such content is quarantined, cut off from it's true target audience. I've got to think there are ways for SU to address this. If Mac OS X can have a pretty effective summarize tool built in, can't a similar algorithm be run against the content of new submissions to SU in an attempt to verify the categorization of that content? Couldn't meta tags, key words, or the sort of tricks search engines use to categorize content be applied? I know these things aren't cheap, but they are possible, and SU has a larger user base than delicious (which may actually be a big part of the problem).

Wednesday, December 10, 2008

My personal knowledge management problem

Full Disclosure:

This post and probably a couple of future posts will serve to fulfill a requirement in my graduate course on knowledge management. But I'm trying hard to approach this in such a way that such requirements are totally transparent asside from this note. Maybe I'll pull it off and this will bear some interest for folks other than my proff. Or maybe I'll totally drop the ball and not engage anyone with this content and totally screw up the assignment to boot. If so, maybe I'll at least fail spectacularly enough to get some good schadenfreude going.


I tried going through the exercises Kirby put together. The basic goal is to figure out which tasks I perform as a knowledge worker bring the most value to my organization, then figure out how much of my time I spend on those tasks vs. less valuable tasks, then try to maximize the time I can devote to the valuable stuff and minimize the time wasted (although that's a slightly harsher term than it needs to be in this context) on less valuable tasks.

I'll be honest, I don't think those exercises work for me right now. There's a couple of reasons for this.

  1. I'm 18 months in to this job and the task that has dominated my time thus far, redesigning the website (we launched the beta by the way, I don't think I took the time to announce that officially here, although I did on the Vol State blog), is not typical of the work someone in this position would be doing otherwise. Once the redesign launches, the way I work will shift, rather radically. It's hard if not impossible for me to look at the last 18 months and make conjectures for the next 18 months.
  2. The biggest drain on my productivity falls outside of my realm of influence; IE it's a trend I am powerless to address. I won't go into detail here but Kirby if you want specifics just email me and I'll fill you in.

So I've been looking at the way Kirby breaks down his model of personal knowledge management and one area I see a lot of room for improvement in the way I currently handle things is with information organization and retrieval. The sad part is I've been sitting on the tools to address this issue for years. I just need to be mindful of how I use them.

I signed up for a Ma.gnolia account back when they were still in beta. I used it for a while, then I found Stumble Upon (hereafter: SU). My thoughts at the time were that SU did all the social bookmarking stuff I had been using Ma.gnolia for with the added element of discovery of new content at the push of a button. That's true in theory. 2 years later, it's obvious that it falls apart in practice.

SU does the whole discovery thing very well. I don't think I ever would have found jQuery without SU. I had been underwhelmed by the JavaScript libraries I had seen during the first few months of that whole buzz and had pretty much written off the whole idea. I was just gonna stick to writing my own custom unobtrusive JavaScript using the Document Object Model. Now I literally use jQuery every day. My job would not be the same without it.

But I discovered jQuery at a time when I actually had the time to take on the learning curve, as gentle as it may be. The official documentation is complete enough that managing access to the information I needed to direct my own learning wasn't an issue either. The only exception I could find to that would be the plug ins, but truth be told if the plug in doesn't make intuitive sense and isn't well documented, I don't use it.

Compare that to some things I hope to learn more about in the near future, such as Drupal and Cake PHP or Perl. Or even compare it to some of the stuff I'm already using but need to reference source material rather than working from the top of my head, like regular expressions and PEAR or Active Directory. Now we're talking about steeper learning curves just as what little time to learning new skills is shaved away as I try to push the redesign through the beta testing phase and into launch. I keep stumbling onto sources for these topics, but lacking the time to fully digest them, I thumb them up and move on.

Ok, that last statement begs the question, if I don't have time to digest this stuff how do I have time to keep stumbling onto new content? First of all, SU is addictive. On top of that, it's so easy to just click the stumble button (with or without specifying a topic to stumble through, such as web design) that I can click through a fresh page or two while I'm checking in the files I just completed working on in Dreamweaver (hereafter: DW). Or while I wait for DW to generate the broken link reports I've been running lately. Actually, now that I bring that up, I really hope DW performs better on the redesigned site. The current site is such a mess of spaghetti code that DW is prone to take its sweet time or even crash when I ask it to perform a site wide action. The redesign is much leaner. Based on the work done so far, the code we shave off should be equivelant to about 68 copies of the complete works of Shakespeare. No, really. Project Gutenberg has the complete works of Shakespeare as a plain text file. I've done the math. :)

This is where the problem comes in. I find these great resources, or at least potentially great, but going back to find them later gets to be a real pain. Sometimes I don't take the time to write my own tags for a page. I just thumb it up and switch back to DW or click the stumble button again. But SU being socially driven, the tags default to the category chosen by the person submitting the site. I currently have 143 stumbles tagged with “graphic-design”. I'm not a graphic designer. I don't really even consider myself a web designer. If you want to split hairs, I consider myself more of a web developer. When I tag an article relating to design, I use “web-design”. I've got 418 of those. But it's possible I didn't personally tag all of them. If the person submitting them tagged it as “web-design” and I just thumbed it up and moved on without bothering to apply my own tags to it, then that's how it would default. It's obvious that graphic design is a pretty popular tag in the wild and the zeitgeist is polluting my tag cloud.

An Example

Over a year ago (November 1st of 2007 according to my SU history), I stumbled upon Scott Jehl's StyleMap script. At the time I thought, “This is how we need to do the org charts.”. Previously we had done the org charts in Microsoft Viso and then those files were exported to HTML. But that produces a tangled mess of frames and images. It's hard to navigate, hard to maintain, and doesn't even work on my Mac (thanks Microsoft). I thumbed it up and moved on.

In October of this year, I finally turned my attention to the org charts for the redesign. I remembered stumbling upon this script a long time ago that would be perfect. But I couldn't find it in my SU history. I tried Google searching every combination I could think of. I literally wasted an entire day trying to find this script.

The problem was the default tag the page was assigned had nothing to do with how I conceptualized the content of the page. I don't even remember what it was now and I have since gone back and edited the entry with my own tags. Google wasn't working because I had forgotten that it was written as a script to do site maps rather than org charts. To add insult to injury, when I finally dug up the article and tried to put the script into use, our org chart proved to be way too complex. But I could have discovered that in an hour had I not wasted an entire day (and part of the following morning) digging up the script.

The Problem

Partly due to flaws in the way I use it, and partly due to flaws in the way it's designed, SU is failing me as a means of efficient information organization and retrieval. In defense of the development team behind SU, it is designed more as a discovery engine than as an organization tool. And I can't sing enough praises as to how well it performs its core function.

The Solution?

So I turn my attention to my neglected Ma.gnolia account. If I start using both these tools to perform the tasks for wich they were designed, and approach my use of these tools in a mindful way, I think I can milk a lot more productivity out of my days. I'll map out that plan in a future entry. Stay tuned.