Showing posts with label Collecta. Show all posts
Showing posts with label Collecta. Show all posts

December 19, 2011

Time Shifting In a World of Realtime


Nearly three short years ago, the buzz word du jour in tech was “realtime”. Real time discovery. Real time search. Real time serendipity. The explosion of interest in social sharing tools like Twitter, Facebook and FriendFeed (remember this was early 2009) had people (myself included) saying that “Delayed news will no longer be acceptable for early adopters, who will gravitate to the quickest sources of news, wherever they may be.” In practice, while this has occasionally been true, I’ve found a completely divergent innovation to play as big a role in the way I (and others) consume news content and entertainment - that of time shifting, which has remained valuable at a time when most real-time search engines have pivoted or vanished.

Best exemplified by TiVo and other DVRs, preceded by the creaky VCR, the act of consuming media at a time much after its initial airing is so commonplace that live viewings are so uncommon that friends often tiptoe around current storylines for top shows. In some social circles, only the most breaking drama series get the “day it actually aired” treatment - like Breaking Bad, Dexter or Homeland, while everything else goes to TiVo, to be consumed later. (Obviously, I saw the season finales for Dexter and Homeland last night)

News, with some exceptions, can be similarly stored away for later viewing, be it through RSS readers or on your social network of choice. One must not be glued to the real time stream to make sure you don’t miss anything. Instead, the RSS reader traps your own hand-picked links, ready for viewing when you get the opportunity, not necessarily tied to their time of posting.

On the big screen, movies may bank on a massive opening weekend, but with consumers having so many options for entertainment sources, it’s common to see people mention they’ll “wait for Netflix”, which could be months or years away, content to save a few dollars while also getting the comfort of watching in their own home. And if you do find yourself suddenly interested in a show your friends have been seeing which has been out a few seasons, don’t fret, as you can, in almost all cases, catch up - tapping into many options, be they Netflix, Hulu, Xfinity, iTunes or Android Market.

This fall, I made it a personal mission to watch all of Mad Men, after hearing people go on and on about its quality. I powered through it with many late-night Netflix marathons. After finally ordering Showtime, I caught up on this season’s Dexter on Xfinity, and then did the same for Homeland. If my wife misses her favorite shows, she can do the same, tapping into the various video repositories on the web, including the big three networks, typically slower to adapt to the innovation of the web.

I watch my evening talk shows 3 to 5 in a row, from Jon Stewart to Conan, fast forwarding through commercials and skipping uninteresting guests - efficiently getting the best and skipping the rest. It’s almost the same approach I take to my RSS reader or activity on the social networks, skimming, reading, clicking and leaving no prisoners. Even if I’m not constantly connected, and I do a good job of getting close, I don’t feel this sense of missing something.

Realtime reactions to breaking news events, kicked off by an initial discovery, and then rattling around search engines and social media, can’t be duplicated by time shifted content, but for most buckets of content, be they text, audio or video, the drive to be first and in the mix of the story as it is interpreted and curated, is not essential. Advents in information and content sharing over the last few years have instead made “on demand” a reality, getting me what I want when I want it, not when someone else decides for me.

March 13, 2010

Users vs. Companies: Conflicts over the Real-Time Web?

If 2009 was the year of real-time Web, with practically every major service finding ways to bring content to its users instantly, 2010 is about optimizing the new real-time world, expanding interoperability between sites, finding more ways for users' content to be discovered, and taking the potential of real-time out of the status world and into the real world. Today, at the South by Southwest Interactive event in Austin, Texas, one panel asked if we were making serious progress in this vision, and if companies, feeling increased competitive pressures, are short-changing users in the process.

Marshall Kirkpatrick of ReadWriteWeb, who moderated the panel, featuring representatives from Collecta, Google, Gowalla and Microsoft, said "the real-time Web is a big, complex and multi-headed beast," adding, "almost as many people you talk to on the subject will give a different perspective."

For most, the real-time Web represents reducing latency from the time updates are published and when they are experienced practically to zero. This can be anything from updates from blogs to downstream aggregators and RSS feed readers, status updates from social networks to other points in the ecosystem, or instant alerts from the Web at large that a saved search you requested has found a positive match.

But one of the existing problems with the real-time Web that has occurred is that despite the focus by many services to solve the same problem, many have done so without delivering true data interoperability - and other services are trying to solve for real-time without having full access to users' public data.

"Back in the day, you couldn't send e-mail from AOL to Compuserve, and today, you can't send data from Google Buzz to Facebook," said Brett Slatkin of Google's App Engine team, and co-author of Pubsubhubbub. "Part of what we are trying to work on is breaking down these barriers that connect to different sites. If I am on Buzz and Marshall is on Identica and Jack is on Twitter, we should all be able to communicate."

Standards have evolved in the real-time Web space, from OAuth to PubSububbub, WebFinger and Salmon (as documented here), but that's not to say there aren't still heated debates over these standards, or even which version of standards should be supported. (See this article for a discussion of OAuth 2.0)

"I try to be a practical person, and when I hear about a family of specifications, it sounds like a family of work," said Dare Obasanjo of Microsoft. "There is clearly a place where we have a common pain that we can work on. There is a bunch of shared pain, and the way you have to get real-time service is to work on APIs, and that is a clear starting point for standards. Pubsubhubbub can help solve that problem, but I get concerned when you have to implement certain specs to solve that problem."

"These specifications we agree on should be useful on their own," answered Slatkin. "When you implement a specification like HTML, you are not buying into an ideology."

As the real-time Web's protocols are debated and deployed, so too does the application of these services. Google Buzz and Facebook have received scrutiny for their aggressiveness in converting assumed private data to public, and Netflix recently canceled an algorithm development contest thanks to concerns of assumed privacy violations.

"When talking about privacy, right now, unfortunately, the social networking market is failing, and they have little incentive to encourage user privacy," said Obasanjo. "I am waiting to see when people find what they thought were private updates as part of trending topics on Google and Bing. Users and companies are in conflict."

Obasanjo gave the example of Twitter needing its users to be public in order to drive value into the system. After all, if users were all private, there would be no trending topics, and thus it is Twitter's best interests for updates to be public. "There is a factor that if a user wants to be private, it subtracts value from the system," he said.

Beyond these concerns, known benefits of the real-time Web are scratching the surface of what could be done with more expanded to real-time data from other sources, it was argued. Slatkin forecast a time where you could query supply chains for inventory and purchase locally instead of from Amazon.com, turning economies of scale on their head. Scott Raymond of Gowalla talked about intersecting real-time Web technologies with geodata to show trending locations and the hot parties of the moment, by decaying the relevance of checkins over time. Jack Moffitt, CTO of Collecta, said a development environment for new tools and applications that leveraged zero latency was becoming "very interesting".

"All these guys are working on realizing the potential right now, working on real-time data," Kirkpatrick said. "Brett Slatkin said it was important people focus on the unforseen future that systems we worked on to support undiscovered use cases - things are going to get real crazy real soon."

Web-wide adoption of RSS and Atom standards has eliminated the problem of publishers providing their data, and tools like Pubsubhubbub are working to get data from one site to another faster. "Polling doesn't scale and you need a push notification to deliver it. It's possible we will have multiple winners, and we have to consider privacy considerations that people won't want their data available to everyone," said Moffitt.

The element of real-time is being layered across the Web, and it seems to be happening even if developers aren't completely in agreement over the tools needed to optimize the experience or if the debates on privacy versus public data are solved. And there's a lot of room for real-time to grow outside of the statusphere and to more traditional markets. The question is can developers provide solutions that don't have users running to the FCC?

March 01, 2010

Twitter Unleashes the Firehose to Seven New Partners

As Twitter promised back in December at LeWeb, the company has now extended access to the firehose, the full feed of public tweets, to a full array of partners - going beyond the initial deals with search leaders Yahoo!, Google and Microsoft. The partners announced today include personal favorites Twazzup, Collecta and Kosmix, taking Twitter more into the realm of platform rather than destination.

Twitter, through its continued rise in traffic, use and the public consciousness, has seen the data flowing through its system grow increasingly valuable, as the company aims to live up to its lofty goal as the pulse of the planet. This rich data set, now open for partners to tap into, can potentially rival traditional search engines in terms of value, especially where real-time reaction and sentiment analysis comes into play.

Kosmix, behind the personalized newspaper service, Meehive, and extended topic pages on its own search tool, wrote to say, "Having full access to the entire tweet streams will give us the chance to surface all relevant and interesting information on what is buzzing on the Web for any given topic."

Financial deals were not disclosed by Twitter, or by its partners, but the company looks to have found a way to recognize revenue and gain income off of the public stream. In this age of real-time data, the faster and more relevant, the better. That Twitter is letting more of its secret sauce spread across the Web will make it harder for other sites, namely Facebook, to look as nimble.

Disclosure: Kosmix has previously done business with Paladin Advisors Group, where I am managing director of new media.

December 29, 2009

Collecta Delivers Real-Time Search for MySpace

Any time I hear the word MySpace, my inner geek coughs and the eyes tend to roll a little. But to other geeks, who recognize the social network's 75 million users are posting a significant amount of increasingly rich media every day, the ability to harness this flow of updates and find information sounds like a true tech challenge worth pursuing. Collecta, a real-time search engine best known for indexing Twitter search, blog posts, photos and videos, this morning has introduced a site-specific search engine just for MySpace, pushing their discovery engine toward what CEO Gerry Campbell called "a different vibe and message than any other service."

MySpace, part of News Corp., while languishing in comparison to the juggernaut of Facebook and the geek hipness of Twitter and others, has become something of a hangout for creative artists and consumers, giving the site a tremendous amount of images, videos and text flowing through the network. In addition to this data, one of MySpace's hallmarks (or quirks) has been its mood features, and a highly entertaining way of language, with all sorts of misplaced capital letters, caps lock and exclamation points. The result is a flow that Campbell called "monstrous" and "rich".

In 2009, as mentioned in the wrapup of my predictions post entered at the beginning of the year, real-time became legitimized. In a conversation I had with Campbell yesterday in advance of this announcement, he said that "people are beginning to expect data is fresh and hot," adding "We see this as the fabric of the Web." Collecta's mission is to not only grow traffic on their main site, at Collecta.com, but to enable site-specific searches for other brands to bring real-time search to their content. The first trial was with branded identi.ca search, and the second is with MySpace, which you can find at http://myspace.collecta.com/.


MySpace Search, Powered by Collecta

"We have, since the inception of Collecta, said our destination site is important to us, but equally as important to our strategy is to make sure others who can take advantage of the real time platform can feed into us, or in other cases, they can start with a finite set of data, and we can turn on a full, streaming, bells and whistles site, to tap into the vibrancy and excitement of the community," Campbell said yesterday.

Differentiating in a world of multiple real-time search engines can seem difficult to the typical visitor. Lumped in with OneRiot, Twitter Search, Topsy and others, Collecta is trying to present more than just the newest results, with Twitter dominating, but to also present hot topics in context. The site's front page, echoed on their MySpace specific search, highlights a photo, a story, an update, and a comment, from the real-time Web.

The hot topics on MySpace become especially interesting when mood is involved - giving the results a very emotional feel. In response to the failed airplane bombing on Christmas, the response on MySpace was very guttural, as moods displayed just how MySpace's members were taking the news. In more positive news, you can see reactions to movies or music, such as "Avatar" (See MySpace search for Avatar Movie) and get moods displayed alongside the messages, ranging from shocked, to inspired, rejuvenated and impressed.


MySpace Moods and Updates Around the Avatar Movie

Lest we get too caught up in the frequent hype around Twitter, the site's messaging traffic is still measured in the tens of millions per day, and Campbell reported that MySpace's flow is at a much higher volume. The difference is even more dramatic when you consider how many retweets are counted in that number.


MySpace Updates and Moods on the iPhone

To separate the signal from the noise on sites like MySpace and Twitter, Collecta is working hard to also include concentrated content from traditional news sources and blogs - working to maintain the integrity of a single blog post that may have dozens of retweets. One approach the company has taken is to have a human editor work to curate the data, and then try to match that activity through algorithms.

While many of us may not have ever registered MySpace accounts, or logged into our long-since dormant accounts, the information and moods flowing through the new Collecta-powered MySpace search is very interesting, and a good proofpoint for the real-time search engine pointing to a new data set. Check it out here: http://myspace.collecta.com/

July 10, 2009

Real-time Search: What's Most Important Now, Not Most Accurate

This afternoon, at TechCrunch's Real-Time Crunchup event, representatives from many of the innovators in the real-time search space had a quick round table aimed at furthering the discussion, framed by a question by moderator Erick Schonfeld, who said that some on the panel may believe real-time search is defined by Twitter Search, while others believe it is "everything on the Internet, but with a freshness or recency component". And while many different companies, including the standard-bearers, like Google and Microsoft, are looking to take on this new challenge, how they are doing it differs greatly.

Danny Sullivan, author of Search Engine Land, said, "We need better definitions of what it is, so as consumers and users, we understand what we are interacting with. Through Twitter and a few other services, you have the option to publish in a few seconds. Maybe you call it social sharing search." (He also posted a summary of the players last night)

Some of the participants could be defined by how large a percentage of their data was initiated through Twitter, and how they worked with the data, including filtering.

"I would define it as what are people saying in real time about my topic," said Gerry Campbell of Collecta. "It's not what is most important, but it's what is in real time now."

This bifurcation of the "one right answer", often championed by the existing search leaders, versus what's most right "now" is helping to separate the old school search engines from this new breed. But don't think that the more-established companies are taking this lying down.

Google's Matt Cutts, who has been at the company since 2000, said "we have always talked about freshness of content." He relayed a history of his time at Google, saying they once had a "war room" of how they could refresh their search index as frequently as a month. By 2003, the company had moved from monthly updates to daily updates, and a few years later, in 2007, integrated the company's Blog Search product into its main search results. "We have rearchitected our system to be as recent as possible," Cutts said.

As updates flow in at an ever-increasing pace from all corners of the Web, search engines have the daunting task of getting accurate responses out there, while ignoring off-topic or harmful data, such as spam. And those who manage to get the formula right will have a serious leg up over those who don't filter well, making their results more noise than signal.

"Drinking from the firehose is a ticking time bomb," said Kimbal Musk of OneRiot. "Even by filtering 90 percent of what is going on with the Iran election, you're still only going to get a tiny slice, and a good portion of that is spam. If you don't filter content, you are going to get more and more spam." He later added, "If you stick to Twitter alone, you will have a spam-filled and biased data set."

With Twitter's API getting to a point where more and more companies are relying on it as their engine and data source, each is working of a common data set, and how they interact with the information will make the difference. And yes, Microsoft or Google may give you one result that is most accurate, but not for this moment, and not with any kind of impact from your friends or in terms of how that data is being interacted with in real time.

Sean Stutcher of Microsoft clearly stated this information is becoming more relevant, saying, "The sentiment around a link could be changing, and that might become very relevant to a user."

In an isolated search world, where an index is an index and the right answer is the right answer, that might not matter. But in real-time, it could matter immensely. As each of these companies works through their user interfaces, their data sets, and improves filters and social aspects, it should be very interesting to see how they separate from the pack and help define their goal.