November 11, 2009

Search: Less Useful Due to Massive Info Growth, the Flow?

In a forward-looking presentation at the Defrag Conference this morning, Stowe Boyd pushed attendees to think about how the Web would look by the year 2019, with the aid of seeing the massive amounts of change that has taken place over the previous decade. One of Boyd's most-aggressive comments stated that the world of search is falling apart, as the problems it initially aimed to solve have been eroded thanks to the information explosion and the corresponding ease of access to social connections in a world of real time. Without saying that social networks would render the established search giants, irrelevant, he suggested, as he has on his blog frequently in the last few years, that the "flow" will replace the world of Web pages - and change the game on search entirely.

Boyd essentially argued that social tools are in the process of changing the culture. He said people were incentivized to discover breaking news from social friends through networks like Twitter and Facebook, which makes the new "real-time Web" interesting. He further suggested that how one interprets this news to define "meaning" is what will replace search.

One of the biggest reasons he thinks meaning will replace search is that the initial argument for search engines was trying to find the few documents on the Web that were relevant to your query, and now, practically any search can deliver millions of results.

"Search is starting to fail because scarcity has been replaced by infinity," Boyd said. "We are heading toward a world where all the critical information is available publicly, and breaking news is a few seconds away - at the most. We will switch to instead relying on finding things through our social connections - engines of meaning, and the source of what is important."

Assuming social elements are going to trump algorithms and crawlers that power today's engines, Boyd said he believes that the most important dimension is now time, not space - and that for the most part, this dimension is shared.

"We are not sharing space online, we are sharing time," he said. "Our time is increasingly not our own. A shared thread of time will be the norm, and how we will get work done."

This new shared thread of time, or "flow", as Stowe referred to it, is poised to become the replacement for today's static Web pages, a new element in today's social Web, which he pontificated could be "the most defining moment of our civilization."

Skepticism Over Current State of Social Web at Defrag

At the Defrag conference in Denver this morning, there was an acknowledgement that social elements are infiltrating practically every aspect of businesses and interpersonal engagement online, but unlike other events, which have seen a practical hugfest over the latest apps or services, the morning's speakers expressed a great deal of frustration over trying to find real benefits and utility to all the activity that is happening online. Speakers suggested today's tools have a stark lack of context, that businesses are too obsessed with having a complete data set and aren't focused enough on the actability on that data, and that many developers are focused on designing apps that simply don't drive benefits.

Eric Marcoullier, CEO of GNIP, was most direct in his comments, saying that "the business world doesn't give a (crap) about your lifestream app," saying that designing yet another application that sorts all your content online is essentially a list of lists - a list of "my stuff" or "my friends' stuff", which is cute, but not necessarily valuable in decision making.

GNIP is best known for offering managing data collection as a service. The company has seen some ups and downs over the last 18 months, culminating in a significant layoff in September that saw the company reduce staff - cutting seven heads from the dozen on their roster. But since the move, Marcoullier said the last few weeks have been "stellar" in terms of productivity, even as his clients aren't necessarily looking for the answers to data - just more data.

He asked, "Is there an opportunity to drive business decisions and revenue for your company?", saying "Data is useless without effort. When you get data, it is a lot of work to do something useful with it, yet market research companies are obsessed with completeness of data."

Similarly, T.A. McCann, CTO of Gist, said that leading social services, like LinkedIn, have curated millions of nodes, tracking millions of relationships. But for most, it hasn't yet been clear how these connections can be leveraged to drive real daily utility - beyond suggesting new connections and companies that should be known due to shared interests.

Much of these shared interests have been displayed in social streams including Twitter and Facebook, which despite their meteoric rise in visibility, are still struggling to provide more than a simple flow of updates and links.

Tim Young of SocialCast complained, "What I find on Twitter is link vomit, or link carpet bombing and swarming about events. During the day, I get all these links, and the issue is I click the link and there isn't a lot of context. Why did they share this and how did it get here?"

Tim called for a new solution to be built that would save traces and paths of content to help communicate new findings to derive value - something made ever more difficult when the most common real time search repository, Twitter search, is now hosting a database that can track as few as only two days.

And despite many people's claims that finding this data ever more quickly is going to make us more productive as a species, Stowe Boyd dumped on that, saying "the myth of increased productivity is a failed world view," adding, "people will trade personal productivity for connectedness, and they will accept an interrupt to help somebody in their social connections."

That's not to say all is dark. Eric of GNIP promised he was still a huge fan of social media, and Stowe pontificated that the rise of the social Web may already be "the most valuable artifact ever created". But from a raft of useless lifestreaming applications and a gap between link visibility and link utility, the speakers seem to agree that we have a long way to go from today's promises to tomorrow's solutions.

Twitter Plucks Data Management Guru from Yahoo!

That Twitter is dealing with massive amounts of data flowing through its servers these days would be an understatement, as the service sees strong growth and significant mindshare. With the company having passed what looks to have been its rockiest struggles over the last twelve months, Twitter is now getting to focus on rolling out some significant new features, from Lists to geolocation, trend definitions and retweets. But the microblogging giant looks like it is taking extra steps to harness the power of its rapidly-expanding data set.

If the company's own team list is to be believed, they just picked up Utkarsh Srivastava, a highly respected senior research scientist at Yahoo!, who is best known for his work on building large-scale distributed systems, specifically his efforts with Hadoop.

Hadoop, similar to the Google File System, is a framework that enables applications to work over distributed server nodes and significant data sets - potentially ranging in the petabytes. Yahoo!, Google's off and on competitor, has been the company most associated with Hadoop. While at Yahoo!, Srivastava was one of the original designers of "Pig", an Apache project for analyzing large data sets, which leveraged Hadoop. (See also the research paper: Pig Latin: A Not-So-Foreign Language for Data Processing)

Srivastava, a PhD graduate from Stanford University in Computer Science, has been working at Yahoo! Research since 2006. (See his home page and LinkedIn profile)

Not knowing what aspects of Twitter Srivastava may be working on, it's premature to assume whether his efforts will be primarily focused on new initiatives, or simply helping the company scale its growth. I can dream and hope that he can be the missing piece that brings Twitter's high potential search engine fully online, but that is no doubt a big project indeed.

Update: This hire has been confirmed by Srivastava and also covered by TechCrunch.

November 08, 2009

The Story of Google's Closure: Advanced JavaScript Tools

On Thursday, Google caught the eyes of Web developers around the world with the company's move to open source its Closure JavaScript compiler, library and template system to the Web community - the very same tools that power popular applications, including GMail, Google Docs, Google Maps, Google Reader, and no doubt many others. The Closure tools optimize Web code to be compact and high-performance, essentially reducing page load and redraw times while also enabling uncompromising capabilities. Around the Web, you could see the release elated geeks both inside and outside Google, many of whom previously worked with the tools while working for the Mountain View tech giant.

To better understand these tools, and get a real-world perspective on Closure, I reached out to Mihai Parparita, an engineer on the Google Reader team, to hear of his experience. He was gracious enough to extend a very thorough overview, explaining the tools' origin and use case, by e-mail, much of which is summarized below.

The Closure compiler dates back to GMail's launch in April of 2004. Paul Buchheit, now of Facebook, via FriendFeed and previously Google, largely credited for the founding of GMail, highlighted the announcement this week on his FriendFeed, calling it the "Gmail JavaScript compiler". The library and template system were initiated a few years following.

As Google Reader development started in early 2005, with Mihai, Jason Shellen, Chris Wetherell (the latter pair now are at Thing Labs working on Brizzly, which also uses Closure) and others working to make a top-notch Web-based RSS reader, the team leveraged Closure immediately after the initial prototypes. At the time, the team was less focused on download size than they are today, but the compiler's aggressive function checking improved error detection.

Mihai writes:
"Until the last month or so leading up to the Reader launch in October 2005, the size benefits of the compiler were less important, since we were less focused on download time (and performance in general) and more on getting basic functionality up and running. Instead, the extra checks that the compiler does (e.g. if a function is called with the wrong number of parameters, typos in variable names) made it easier to catch errors much earlier. We have set up our development mode for Reader so that when the browser is refreshed, the JavaScript is recompiled on the server and is used with the page when it is reloaded. This results in a tight development loop that makes it possible to catch JavaScript errors as early as possible."
As the library and template systems did not arrive until approximately 2006, Reader utilized homegrown code in their place that provided similar functionality, including handing different browser versions and quirks, Mihai said. But as soon as they were available, Reader used the new tools for new code, and later, to replace old shared libraries and homegrown code. Mihai said he performed an audit to detect usage of the old code, and find their Closure equivalents, so work could be distributed among the team during so-called "fixit" periods, when attention was given to code quality instead of new functionality.

With Closure implemented, benefits to Google Reader users are clear. Mihai estimates that without Closure, Reader's JavaScript code would be a massive 2 megabytes, which reduces to 513 kilobytes with Closure, and all the way down to 184 kilobytes using gzip, supported by nearly all browsers. Additional benefits include the near-elimination of concerns around browser differentiation, and an extremely manageable large JavaScript codebase "that doesn''t get out of control as it ages and accumulates features", he said. (Note download time was given as the main reason Robert Scoble has moved away from Reader and that the team recently made a push to even further optimize the code)

Closure's role at Reader, initially utilized in low level code, has "moved up the UI stack" to to the point where it is leveraged for UI widgets. Mihai says "this means that it's not a lot of work to do auto-complete widgets, menus, buttons, dialogs, drag-and-drop, etc. in Reader."

The excitement around Closure's release was palpable from developers through Silicon Valley and beyond as you could see from blog posts by Erik Arvidsson, a co-creator along with Dan Pupius, and a series of posts at bolinfest.com. Other excited Tweets came from Mike Knapp, the aforementioned Chris Wetherell and Kushal Dave.

As Mihai says, "You can tell that there's something special about this when you look at the ex-Googlers cheering about its release. If it had been some proprietary antiquated system that they had all been forced to use, they wouldn't have been so excited that it was out in the open now."

Like many other projects at Google, Closure's compiler, library and templates were derived solely as 20% projects and are largely still dependent on work done in so-called 20% time at Google. Mihai says that if one project needs a feature from the compiler or the library, they are encouraged to contribute to it as well.
"To give a specific example, Reader had some home-grown code for locating elements by class name and tag name (a much more rigid and simplified version of the flexible CSS selector-based queries that you can do with jQuery or with the Dojo-based goog.dom.query)," Mihai said. "As part of the process of "porting" to the Closure library, we realized that though there was an equivalent library function, goog.dom.getElementsByTagNameAndClass, it didn't use some of the more recent browser APIs that could it make it much faster (e.g.getElementsByClassName and the W3C Selector API). Therefore we not only switched Reader's code to use the Closure version, but we also incorporated those new API calls in it. This ended up making all other apps faster; it was very nice to get a message from Dan Pupius saying that the change had shaved off a noticeable amount of time in a common Gmail operation."
Now clearly I'm no developer beyond simple HTML and JavaScript, but I know good Web apps when I see them, and Google's Web apps (as well as Brizzly) are among the best in the world. They have managed to take what used to require massive software installs and make them relatively lightweight Web instances with similar functionality between services. With the release of Closure, sharp Web developers will be looking to leverage these JavaScript libraries and tools to make their own products best of breed - something that will benefit the Web as a whole. I appreciate Mihai's openness, and his willingness to share the story behind the story.