Showing posts with label Outage. Show all posts
Showing posts with label Outage. Show all posts

December 20, 2010

Don't Fear Being Fireballed or Slashdotted if on Blogger

It can be difficult to plan for an unanticipated spike in Web traffic. One of the most amusing parts of reading Macolyte John Gruber's Daring Fireball blog is the near certainty that at least once per day, one of the poor saps he links to goes down due to the resulting crush of curious onlookers. It happens so frequently that it's practically second nature to see him update a post saying the link has been "Fireballed", opting instead to point to Google's cache of the original. This phenomenon, of course, is not a Gruber exclusive. For years, otherwise innocent bloggers have been Slashdotted, TechCrunched or even Scobleized, seeing their sites crack under pressure. But from what I can tell, if you're hosted on Blogger, this isn't an issue at all. The company was just recognized for perfect uptime during a survey by Royal Pingdom, and I've seen my own site spike in traffic following controversial posts with no ill effects.

Blogger was not always known for a stellar uptime record. In late 2006, the service practically had to write "a novel" about continued outages. In fall of 2007, I railed against Google for ignoring users during another major outage. But services mature and times change.

Royal Pingdom's Downtime Report. Blogger Scores a Zero.

With the backdrop of Tumblr's highly-publicized downtime of more than a day earlier this month, Royal Pingdom also spotted occasional outages at WordPress, TypePad and Posterous, each of whom looked great compared to Tumblr's unfortunate blip, but not as spotless as the often overlooked Blogger.

Getting Slashdot, Scoble and Spiegel Bumps Simultaneously


While escaping system-wide downtime is a major win in itself, that doesn't speak to the services' ability to scale under pressure. Gruber recommends the WP Super Cache Plugin for WordPress, letting users serve more than a page a second, as it's self-hosted WordPress blogs that often get crushed with one of his links. But I've been linked from Daring Fireball a few times and lived to tell the tale. Same for Slashdot, Techmeme, Hacker News and the Huffington Post, all of which are capable of sending solid and sustained traffic on good days.

Techmeme, Hacker News and Scoble Don't Bring Down Blog

Similar to how Google is well known for its Web site never going down, I have to assume Blogger has mastered Google's distributed architecture, and makes no single blog, or its hotspotting, an instrument for failure. This blog didn't slow or go down when tens of thousands of visitors dropped by this July to tell me I was an idiot for switching from iPhone to Android. It also didn't go down when I questioned the usefulness of Google Wave last year, or when the Huffington Post liked my recap of David Kirkpatrick's Facebook Effect.

Not Even Daring Fireball Could Take the Blog Down

Even the largest services have occasional downtime. Facebook was down last week, intentionally, due to a code push ahead of schedule. Twitter has had its share of bumps, though the site has gotten much better in the last year, and even Facebook subsidiary FriendFeed went down this weekend. But Blogger users are sitting pretty knowing their blogs are going to be up and responsive, even under pressure.

October 06, 2010

Congratulations Foursquare on Well-Deserved Downtime

Over the last two days, the popular location based service and king of check-ins, Foursquare, has been suffering from intermittent outages related to database overload and storage troubles. And this is wonderful news. Not because I wish them ill will, but it seems this is a necessary strain practically all successful companies go through at some point as they achieve rapid user growth. How users react to what are hopefully temporary bumps, and how the service reacts after the fact can help determine if the company will continue to be embraced, like a Twitter, or dumped, like a Technorati.

Foursquare's rise from its SXSW 2009 debut to near pervasiveness among the digerati, becoming synonymous with the location check-in movement, despite big players like Facebook and Google participating, and more direct competition like Gowalla and Mytown, is a testament to the service's viral nature, including game mechanics, close ties to friending and following, and a casual attitude that eliminated the dark feeling of insecurity many people once had about announcing their locations, and replacing it with something one not just accepts, but expects.

Downtime Hits All Folks, Even Facebook (from Tonight)

Foursquare seems to be in the position to choose their own destiny - to forge ahead as an independent company, flush with cash from a summer VC round, or to court willing suitors at a later date, hopefully at a higher valuation than deals they famously turned down in a public negotiation for control. While two days of shakiness are no doubt unwelcome by the company or its addicted userbase, it seems a rite of passage that the biggest Web services go through in a circuitous route to the top.

Twitter is by far the most famous for continued bouts with downtime, closely tied to scalability issues, and its fail whale mascot is the icon of downtime that all other interruption pages wish to be when they grow up. The company's status blog at http://status.twitter.com/ has actually seen 15 different posts since September 1st, contrasted with only 10 posts from the company's official blog. (http://blog.twitter.com/)

But the blue bird and white whale on a teal background are not the first to document downtime doom. Amazon Web Services promotes a service health dashboard, while Salesforce.com displays a system status at Trust.Salesforce.com. Google provides a status dashboard for Google Apps, with dedicated pages for Blogger status and Feedburner status and others. Even pokey old Web 1.0 dinosaur Hotmail has a status dashboard.

Common to all these services? Previous downtime issues that hit the community hard. Truth is, until there is a major downtime incident, most companies do not offer insight into their status. They may feel like they are teflon, and they can't be taken down, or they simply didn't think about what could happen if something went wrong. It takes a public shakeup that calls the service's stability and viability into question to push for improved system integrity and redundancy, and despite best efforts, some never quite solve the problem.

Back in 2007 and 2008, Technorati was as much known for its Technorat Monster escaping as anything else. This cutesy downtime message foretold the demise of the service, which had to strip out much of its functionality simply to stay up, and was pretty much abandoned by all but a core of users who watched the company evolve its business model and reduce their technical demands. And so far as I know, Technorati never launched a status page. They didn't do what the many other successful companies (above) did when push came to crash. Foursquare, it looks like, is working hard to push this recent event and will come out of this stronger because they are following the proven path - of ownership, late hours, and public transparency.

Tonight may suck for the company's employees, who are taking a public flogging that cans of Red Bull and pizza won't fix. But they should wear the fatigue as a badge of pride. In time, they could look back at this week's events with warm and fuzzy feelings, remembering the outage, which resulted from "too many checkins", as a clear tell they were graduating from Web curiosity to Web titan.

August 17, 2009

Progress Being Made On Twitter's Network After 10 Days of Attacks

After an initial wave of attacks brought the popular micromessaging service to a standstill, as well as a host of other social networking sites, on August 7th, the Twitter team has been fighting to turn the tide and restore access for the service's customers and developers to near-normalcy. Ten days into the attacks, they have not disappeared, according to postings from Twitter, but it looks like the company may have turned a corner in dealing with the additional stress from the distributed denial of service.

In a post to the company's Twitter Development Talk forum titled "Platform Status Update" entered at 4:30 this afternoon, Twitter's Ryan Sarver said:
"Over the past ten days we have been dealing with a lot of stress on our network and that has caused a number of our partners to be knocked offline for extended periods of time. This is obviously not something we want to happen and the platform and ops teams have been working hard throughout that time to address the needs of our ecosystem while protecting the system as a whole. We have made some great progress today in tuning the system to a point that should allow our partners to operate as they were before the recent issue began."
He added "the system is still under stress", and asked developers who might be continuing to have issues to work with them and provide information around what IP addresses their machines use to access the Twitter API, the method they are using to make requests, and any other information - including data as granular as their host operating system, browser, cookies and network connection.

Ryan's comments followed a note from earlier in the afternoon by Alex Payne, the company's platorm lead, where he also said the company continues to "fend off" the attack, saying that without specifics from developers (like those listed by Ryan), they could not adequately troubleshoot any issues.

While the majority of Twitter users are not experiencing the widespread outages that we saw a week and a half ago, it is still not uncommon for applications utilizing the Twitter API, and specifically OAuth, to see failures. There's also no doubt the ongoing attacks have led to a fair share of fail whale sightings. (I've encountered a few myself) Twitter posted a note to the company's status blog Sunday acknowledging the problems, and said they are "working to resolve it".

It has been tempting at times to poke at Twitter for its struggles during this period, but as I mentioned on August 7th, when the troubles first arose, the company is being much more visible and fast to respond to developers and to provide status, as they show signs of maturing. Their operations teams look to have been busy around the clock for the last two weeks, and they will need to keep fighting - as it looks like this battle may go on for some time.

August 07, 2009

Twitter's "Harsh and Cold" Honesty Tells Devs No ETA for Fixes

The much-discussed distributed denial of service attacks over the last two days which hobbled Twitter and also impacted a number of other sites, including Facebook and LiveJournal, has also had a tremendous negative impact downstream, bringing many third party applications and services to a halt. With the sluggishness and non-responsiveness stretching into what will be the third calendar day tomorrow, Twitter told developers, heading into the weekend, that not only was the attack ongoing, and not lessening, but they had no ETA for anything to get better.

While many developers are understandably sympathetic toward Twitter's troubles, even as the popular microblogging service's network architecture is being analyzed for weaknesses, the strain many of them were feeling without receiving answers started to be on display this evening as hours went on without an update - while at the same time, key Twitter employees happily updated that they were going out for sushi or enjoying tonight's Underworld concert.

One developer, in the company's Google Groups forums, wrote:
"There's been lots of posts from Twitter saying they plan for better communication. Okay, cool. But it's Friday night and I am pulling my hair out with no new information. I am sure they are doing all they can... but many third-party apps have been broken for 2 days now. And just as Twitter is getting bombarded with emails from developers... developers, too, are getting bombarded with complaints/questions from customers."
After the same developer noted that employees had gone out to catch dinner, he exclaimed frustratedly, "While food is important, would be helpful to also update devs what's going on, too, before a night on the town?!?!?!" This led to another developer's stating he too agreed with the outrage, adding:
"Thanks for your outrage. I thought i was one of the only few that have issues with how things are being handled."
But shortly after the conversation threatened to get ugly, Twitter's Chad Etzel posted an update, around 8 p.m. Pacific time, catching up the community, and he didn't have much good news.

His major points:
  • "The DDoS attack is still ongoing"
  • "the intensity has not decreased at all"
  • "interaction with the site and with the API will continue to be shaky"
  • *There is no ETA on fixing any of this* (repeated twice)
He told developers that the situation at Twitter will "continue to be rocky" and may actually get worse, but also said that the company's operations team will be working "around the clock" this weekend, adding, "as much as you want it to be fixed, we want it to be fixed more."

With past outages, Twitter has had a mixed bag in terms of communicating to the developer community. Had the initial complaints gone unanswered, it's likely Twitter would have transitioned further away from being the victim to being seen as part of the problem. Yes, applications are broken. Yes, access is still slow, and as the site has become a major piece of infrastructure in today's social Web, its downtime is more than trivial. But this isn't a situation of the company's employees partying on a sinking ship. It sounds like the team is strained, and being realistic about what will continue to be a long process of restoring normality.

June 02, 2009

TweetStats Down More than 24 Hours As Twitter Attacks Cache Issues


TweetStats: Closed Since Sunday Afternoon

On Sunday, we mentioned Twitter had run into a bug that masked the display of third party clients on the service, erroneously reporting all updates as coming from "Web", whether they were, or if they were instead from the mobile interface or any of the growing array of applications that interact with the microblogging leader. Now into Tuesday, the bug hasn't yet been fixed, and in the meantime, popular statistics tracker TweetStats has been shut down - incapable of operating correctly with bad data pouring in.

Twitter's API team is telling developers that the "from Web" issue is the result of large growth in the company's database, thanks to a recent increase in API developers and their registering applications to work with OAuth. The database object reportedly grew to a size incapable of being cached, dramatically impacting performance.

Still a small company, rather than push for Twitter's engineers to come in over the weekend to resolve the issue, the company said they chose the "quick solution", opting to "disable source parameters", according to postings on their Twitter API forum here.

Initial guidance was that the issue would be resolved "likely early in June 1 workday" (sic), but more issues have cropped up. The company followed up with a second note saying, "Due to problems with other *critical* code we've had to delay deployment of this fix until tomorrow."


TweetStats Reports The Trouble Sunday

TweetStats shut down mid-day Sunday after the issue had impacted the statistics gathering service dramatically. And while Twitter has again offered a new date of resolution, the developer, Damon Cortesi, has said he'll find a way to get rebooted, with or without help from Twitter.


TweetStats Hopes for Best, But Prepares...

He noted late Monday night, "If not fixed tomorrow AM, will re-open and deal with the consequences."

Jesse Stay has more around the details of the API issue here: Where is Twitter’s Emergency Response System?

May 08, 2009

Every Piece of the Infrastructure Carries Potential to Fail

Though it may end up being a temporary blip, at this moment FriendFeed is down, following a scheduled outage at Twitter this afternoon. And while that's not really news, it comes on the heels of many discussing the potential for failure that third-party URL shorteners bring to the Web. For every fan of TinyURL or bit.ly, there are others who say relying on another service to be a go-between between the user and the intended data is just begging for trouble. But the truth is that in a network, when there are multiple items with potential to fail between the user and the data, any one of those pieces in many cases can bring the entire system down to its knees.
  • Storage can fail.
  • Servers can fail.
  • Networks can fail.
  • Routers can fail.
  • Lines can be cut.
  • Services can close down.
  • Users can delete images or pages.
It happens, and until we control all aspects of the system, there will be outages.

On Wednesday, in the middle of testing a third-party Twitter service, I linked to the Guardian using a URL shortener called tr.im, required to get the service to work. Later that night, tr.im failed, and it broke all links that were being used.


The conversation (in Google cache)

In response, Paul Buchheit, co-founder of FriendFeed, with a long history at Google, Microsoft and Intel prior to his latest efforts, referenced the break, calling it "another reason why url shorteners are annoying."

But FriendFeed itself has a URL shortener, called ff.im, which it uses when sending updates to Twitter. Paul added in the thread, "Except ff.im of course :)"

But guess what? Because FriendFeed is down (for now), also down are the ff.im links, making them as likely to fail as any other third party shortener. I could rant up and down saying that FriendFeed and ff.im should be served from different data centers, or offer better redundancy, but I won't. Nobody loves downtime, and FriendFeed by and large has had a fantastic track record of staying up. But as they become a more integrated part of the ecosystem, they too will get more opportunities to fail and need to take the same safeguards to protect the infrastructure as do all the other players.

Things will fail. We will live, but we know that there is no such thing as a fully redundant failsafe machine. Every hop delivers the potential to turn into a skip, and not in a good way.

March 13, 2009

Nobody Can Hear You Scream If Your RSS Feed Is Dead

I didn't make any blog posts on Thursday, after a full Wednesday which included a visit to Google headquarters to meet with the Google Reader team. And even though I made a few posts on Friday, the first day of the SXSW conference, many people still think I'm on a temporary hiatus, thanks to a tag-team failure between FeedBurner and Blogger, who have significantly impacted many users by zeroing out their feeds, stopping their posts from getting out of their domain. It looks like I should have spent more time in Mountain View after chatting up the Reader team, to see just what the heck is going on elsewhere on campus.

As the resurgent Kent Newsome of Newsome.org noted today, in his post, "FeedBurner & Blogger Conspire to Assassinate My Joy", the XML file that Blogger generates to distribute RSS feeds was completely wiped out - and I have been impacted as well. No matter if I had 20 or 2,000 posts historically, the file reports it has zero kilobytes, and no amount of trouble-shooting thus far has been of any help.

In this world of RSS-enabled services, a failure of this level means that my posts aren't getting out to the previously-mentioned Google Reader. They aren't populating social networking sites like FriendFeed, Socialmedian and Facebook. And that will no doubt cripple visits for those affected.

As always, I at first considered the issue was my fault. I noticed FriendFeed this morning wasn't picking up my posts, and tried to delete and add my blog six ways from Sunday, but with no success to speak of. While Kent recounts his searches for real tech support, I was in the air during much of the day, and came back to find I was not the only one hearing the silence from Google, as others assumed silence from me.

Complaints are wide-spread in Blogger's help forums - from RSS Feeds are EMPTY today when publishing via FTP. to atom.xml is empty and simply empty feeds.

I'm concerned on quite a few levels here. FeedBurner and Blogger are among the most mocked Google products because of their presumed neglect. Yet they are two of the major foundations for my blog. Also, I am concerned because we are seeing this issue as a weekend starts and a major tech conference is starting up. There's a huge chance this issue won't be resolved if Google is asleep at the wheel. There's nothing I can do except accelerate my move to Wordpress. This doesn't make me feel all that lovey-dovey with Blogger at the moment.

December 15, 2008

Blog in the Dark Much?

Just before 7:30 this evening, as we were putting the twins down to sleep, the lights fluttered and went out. They whirred to life again, twice, but soon dropped again, and we've pretty much been in the dark for the better part of two hours. No TV. No WiFi. Not even the background noises of the refrigerator and heater. With temperatures in the low 40s outside, our home is cooling, and we've unpacked the flashlights and candles we could find.

Did you know you can combine a C battery and a D battery in a flashlight, and it will still work? By necessity, I found out tonight that it does.

Preferring to be constantly connected, my iPhone 3G is keeping me sane. The laptops are fairly useless, but not my ubergadget. It still lets me post updates to Twitter and FriendFeed, browse bookmarks, and read e-mail. Nobody told it to shut down, after all.

Losing power is really no big deal, for the short term. Everyone is safe, and if this goes longer, we could pluck our twins from the cooling crib, and warm them up ourselves. But it's got me thinking about being better prepared for something bigger. It's now clear we need more batteries. And the whole hubbub about Twitter being a good news hub during emergencies doesn't hold too well if you lack power. And it doesn't translate well to the small screen. If a neighbor has discovered the source or reach of this outage, I haven't seen it. My network is too diverse and too noisy to get data on local happenings.

So for now, we're a little disconnected. It's dark. It's getting colder. And we don't have answers. I'm lucky to have the iPhone 3G around, but that aside, the infrastructure holding our power and Web together looks pretty flaky - even in Sunnyvale, smack dab in Silicon Valley proper.

Update: Power was restored shortly after 10 p.m., having been out for just under three hours. No cause has yet been determined.

October 21, 2008

Two Features Every Gmail User Must Utilize

By Mona Nomura of Pixel Bits (FriendFeed/Twitter)


Reading Search Engine Journal Loren Baker's Gmail horror story brought back my Gmail nightmare.

Way back in 2005, I tried logging into Gmail as usual. But Gmail kept redirecting me to this odd error screen with the message: "Sorry... account maintenance underway" and would not let me sign in. (Above image taken from my 2005 Japanese blog.)

After trying (and failing to log in) for two full days, I contacted gmail-maintenance@google.com, and even posted in Google Groups, but did not receive a resolution or even a response. I tried logging in twice a day, everyday, for four months, and finally my persistence paid off. Out of the blue my access had been restored and I haven't had problems since. But to this day, I still have no idea how or why my account was under maintenance -- for four months.
  • Yes, I know GMail is free.
  • Yes, I know GMail is still in beta.
  • Yes, I am aware I should not be complaining... but it's... Google.
Even if it's free and in beta, Google isn't supposed to... break. As embarrassed as I am to admit this, I quickly got over the trauma, and continue to use Gmail. But when Google nightmare stories catch my eye, it brings me back to 2005, and the panic of when I couldn't access my e-mail, compelling me to go out of my way and remind my friends the same thing could happen to them.

Fortunately, Gmail has two great backup features that takes only a few seconds to set-up. My peers were extremely thankful I shared, so hopefully they'll help you too. :)

Two Features Every Gmail user should have enabled:

E-mail forwarding.

I created a backup e-mail account for my main e-mail, and have a copy of everything sent to my inbox to my backup. To set this up:
  1. Create a back up e-mail (ie: mye-mailaddress.backup@gmail.com).
  2. Settings
  3. Forwarding and POP/IMAP -> Forwarding -> Foward a copy of incoming mail to "mye-mailaddress.backup@gmail.com" -- or whatever your backup e-mail address is.
Gmail's "Send mail as:"

Gmail enables adding custom 'From' addresses for free. learn more here.
So in case my main e-mail is disabled for one reason or another, I can always send e-mail as my e-mail address from my backup. For free. Pretty neat.

With the above, I have some peace of mind, though I truly hope I will never ever get locked out of my account again. Bonus: check out techradar.com's "40 Brilliant Gmail hints, hacks, and secrets" it may have some more useful tips.

Have any of the nightmare Gmail stories happened to you? Is Gmail your primary e-mail address? Do you have any preventative tips or tricks I don't know?

Read more by Mona Nomura at Pixel Bits

August 11, 2008

GMail and Apple's MobileMe Holding an Outage Contest

Apple's replacement for .Mac, MobileMe, has been roundly mocked for its spotty uptime since rollout last month, drawing the company's CEO, Steve Jobs to apologize for the lack of quality in an internal memo. But even following an internal reorganization and the public thrashing, users, including me, were unable to access their e-mail for a good portion of the afternoon - even as the company's MobileMe Status page shows no updates since the end of July.

Not to be outdone, the most popularly cited alternative to MobileMe, Google's GMail, has also suffered outages this afternoon, locking its many users out of their e-mail, again, including me.


At the beginning of the issues with MobileMe Mail, Apple famously said the outages were only impacting a small 1 percent of users, despite widespread complaints throughout the Web. Today's outage, which Apple reported lasted about a half hour, cited only that "MobileMe members were unable to access MobileMe mail", so that indicates a full outage.

GMail, on the other hand, says, "We’re sorry, but your Gmail account is currently experiencing errors," without going into detail as to how widespread the issues are. GMail even goes the extra mile to promise "your account data and messages are safe." A discussion sparked by Shey Smith on FriendFeed shows the outages don't appear to have hit everyone.


The 1-2 punch of the outages has made discussion of the downtime the top conversation starters on Twitter, even higher than the Beijing Olympics or the Russia/Georgia skirmish. People must really hate having their e-mail interrupted!

July 31, 2008

TweetStats Shows Impact of Instability on Top Tweeters' Activity

Much of the impact of Twitter's frequent downtime has been anecdotal. Amid some saying they're leaving the service for greener pastures, or developers pulling up their stakes in the Twitter community, statistics show that some of Twitter's most prominent and active users have dramatically reduced their activity on the site over the last two months.

The drop in total tweets by virtually every top user who was measured could be part technical - due to their simply being unable to login, or psychological, a result of lower activity and lower conversations which became a self-fulfilling prophecy. While none abandoned the site altogether, what could have been an "up and to the right" activity graph has largely stalled, and in many cases, reversed.


My own activity, rising month by month after I finally started using the service in January, stalled in June, and is still well below what it likely would have been had stability not been impacted.

Using TweetStats, a site which can show your total tweeting activity, who you most frequently message, and which hours and days you use the service, I polled ten top Tweeters to see how their June and July activity compared with April and May. Here's what I found:



Chris Brogan / @chrisbrogan
April and May Tweets: 2,896
June and July Tweets: 1,070
Change in Tweeting: Down 63%


Corvida Raven / @corvida
April and May Tweets: 2,669
June and July Tweets: 1,065
Change in Tweeting: Down 60%


Danny Sullivan / @dannysullivan
April and May Tweets: 1,281
June and July Tweets: 551
Change in Tweeting: Down 57%


Dave Winer / @davewiner
April and May Tweets: 1,535
June and July Tweets: 527
Change in Tweeting: Down 66%


Drew Olanoff / @drewolanoff
April and May Tweets: 2,131
June and July Tweets: 909
Change in Tweeting: Down 57%


GeekMommy / @geekmommy
April and May Tweets: 6,030
June and July Tweets: 1,419
Change in Tweeting: Down 76%


Jason Calacanis / @jasoncalacanis
April and May Tweets: 1,017
June and July Tweets: 562
Change in Tweeting: Down 45%


Leo Laporte / @leolaporte
April and May Tweets: 363
June and July Tweets: 237
Change in Tweeting: Down 35%


Robert Scoble / @scobleizer
April and May Tweets: 3,579
June and July Tweets: 746
Change in Tweeting: Down 79%


Michael Arrington / @techcrunch
April and May Tweets: 1,587
June and July Tweets: 1,079
Change in Tweeting: Down 32%

Across the board, Twitter's issues cut activity to the site by about half or more for some of the most visible users of the site. Others, like Kevin Rose of Digg (TweetStats) and Pete Cashmore of Mashable (TweetStats) saw only a less than 20 percent reduction in their Twittering activity between the two time periods. While there's no doubt many people, like Steve Rubel and Allen Stern, wish discussion of Twitter's problems would just go away, the impact it had on the site over the lsat few months has been very real, and we're just now able to take a step back and measure its impact.

July 24, 2008

Twitter Finding New and More Creative Ways to Fail

Just when you thought it was safe to Tweet again, Twitter ran into yet another database problem, which not only resulted in sporadic "Fail Whale" sightings, but dramatically impacted the roster of those following one's updates, as well as those each of us were subscribed to. The latest snafu comes at a time when the growing number of microblogging addicts are seeking alternatives, moving to FriendFeed, Plurk, and increasingly, Identi.ca.

According to a post on the Twitter Status blog, the issue first showed a reduced number of followers on the service, and later, in order to solve the issue, Twitter went into a maintenance mode, warning of lower counts across the board. Amusingly, they claimed that some of the lower counts could be due to the removal of spammers, but in my experience, it's been more than 9 hours since the problem was first identified, and I've seen the number of people I'm following drastically cut, from more than 1,500, down to "only" 672, less than half.


On Monday, when I said "The Talk About Rules for Social Following Is Getting Out of Hand", I had taken a screenshot of my current Twitter ratio, at 1,534 to 1,441, after having worked for a good part of the previous week with Twitter Karma to get my ratio synchronized. Just a few days later, that data is carved to 672 and 1,236, prompting some to try and refollow me, and even more to flock to identi.ca.

Twitter's gotten a lot of abuse on this blog in the past few weeks, as we've gone over issues with developers, uptime and changes to the API, but every time I think they've captured the market on a single route to failure, they find another way.

The team's employees are talking a good game about getting this resolved, but seriously, Twitter, why should we believe you now?

See also:

Why Does Everything Suck: The Nightmare Twitter Scenario May Be Upon Us
Profy: So You Thought Nothing Could Be Worse Than Fail Whale? Now Get Your Followers Back

July 21, 2008

TweetMeme Returns Following Months-Long Twitter-Forced Outage


The on again off again cold war that Twitter has been having with its development community has been the subject of much discussion over the last few weeks, especially with the news of reduced unauthenticated API calls, and the new integration of Gnip. But even as Twitter is appearing to get its footing, significant damage has already been done to many services that relied on the microblogging service to survive. One of those was the popular link tracker, TweetMeme, which returned to the Web over the weekend, after months of the service being unavailable, not thanks to developers' neglect, but Twitter's restrictions.

TweetMeme launched in January, gaining significant coverage in the blogosphere, including an article in TechCrunch, who gushed, "The killer Twitter-tracker just arrived and its name is Tweetmeme". But by May, Twitter, under incredible pressure, started disabling developers' access to Jabber and XMPP services, which knocked the service off the Web.

See: Tweetmeme Down Due to Twitter Jabber Problems

At the time, the downtime was expected to only last days, but it turned out to be months.

Service founder Nick Halstead, also the author of Fav.or.it, wrote in a comment on this blog Friday, "Our side project http://www.tweetmeme.com which was the first twitter URL tracker has now been down for months because we were offered the use of the XMPP feed and by the time we had implemented they pulled it. We will not bring it back up again or put development effort into it unless these kind of restrictions are a thing of the past."

The XMPP firehose has famously been limited to only four partners - FriendFeed, Zappos, TwitterVision and Summize, plus Gnip this last week. And Tweetmeme couldn't play on the uneven field, shutting down. But as of yesterday, Halstead reported his team had a work-around, essentially piggy-backing on the search capabilities of Summize itself, now owned by Twitter. (Confused yet?)



Tweetmeme is back in operation now, aiming to show the most popular shared links on Twitter, highlighting the biggest stories on their front page, like Techmeme does, and showing them in order of appearance on the Tweetmeme river, just as Techmeme's river does.

Now that Tweetmeme is back in action, the questions remain - will traffic return, remembering the site's out there, and can it deliver relevant results worth following, as Techmeme has proven it can? And will following Summize's lead be good enough, or will Twitter change the rules again? Hard to know, given the microblogging giant's inconsistencies. That's why many developers are bailing on Twitter altogether.

June 18, 2008

twitAbit Debuts as New Service to Escape Twitter Downtime

Many Twitter users have a love/hate relationship with the service. They love what it does, helping people communicate in real time, from the Web or their mobile phones, but they hate that it hasn't scaled to meet demand. In its place, a new crop of services is rising to work around the downtime. The latest, debuting today, is called twitAbit, which leverages store and forward capabilities to ensure that Twitter fail doesn't ensure your own fail.

The Twitter "fail whale" is well known and a great number of users are looking for a way out. Some have left Twitter. Some are just using it less. Others have moved on, to Plurk, to FriendFeed, or Pownce. But leaving Twitter comes at a high cost for those who have invested time in building relationships, and in some cases, thousands of followers. Even despite the many outages, the vast majority of Twitter's user base has largely stuck it out, hoping for better times.


But if those better times don't come right away, twitAbit is prepared. The service offers a simple form, asking for your user name and password, what you are doing, and a link. It appears to be a project of betaworks, and was announced on the switchAbit Web site, which RSS and blogging guru Dave Winer announced back in May.

At the time of posting, Winer promised Flickr to Twitter functionality, and a second Twitter application, most likely twitAbit, Although it's not 100% clear, he has spoken of a need for a decentralized Twitter, and this could be the first step.

Also: You can see Winer's first "tweet" from twitAbit back on June 13th.

June 11, 2008

Google Blogger FTP Publishing: Out for 12+ Hours

I hadn't planned on making my blog a sounding board for all products that have had significant downtime, but this one has certainly hit close to home.

Starting yesterday evening, around 8 p.m., I have been completely unable to add new posts to the site, making it appear that I am asleep at the wheel. The culprit? An issue with Google's Blogger service, which has blocked the ability to post via FTP.

This is not the first time Google's Blogger has had a outage of significant length here, and also, not the first time they have completely ignored a throe of user complaints and support requests on their site.

At a time when Wordpress and other platforms are gaining significant momentum, and can tout "5 minute" upgrades, the temptation to move, assuming the site structure and comments are retained, is extremely high.

I'd have thought Google's acquisition of Blogger via Pyra Labs would have provided the team with significant experience in growing a scalable, trouble-free infrastructure, but from conversations I've had with people close to the team, it seems that the most infrastructure-focused employees at Blogger had stars in their eyes around Google's other products, and Blogger has suffered from neglect.

Something is Broken indeed.

Update: Finally acknowledged on Blogger Status and FTP is starting to flow. But this is ... bad.

See Also:

Google Blogger: Something is Broken!
http://tinyurl.com/5sejhy

Thomas Hawk Cites Blogger Outage on FriendFeed
http://tinyurl.com/3p96y8

From August 9, 2007:
Google Ignores Users During Major Blogger Outage
http://www.louisgray.com/live/2007/08/google-ignores-users-during-major.html

June 09, 2008

SiteMeter Stats Sputter to a Stop, With No Reason Given

It seems that outages are the new black.

After a weekend filled with stories on Amazon downtime, a brief Disqus blip, continued Twitter troubles, and many sites straining to take on increased crowds swelling to catch the latest from WWDC, I was surprised to see my blog statistics engine, SiteMeter, get in on the act. Since 11 this morning Pacific time, data has been almost completely stalled, not logging visits, and the company's blog doesn't give any reason for the slowness.

I'd like to blame Scoble, or blame Steve Jobs, but I don't think they're the cause.


SiteMeter is one of the most widely used statistics trackers in the blogosphere. And while I could put up with occasional outages from a free product (See: Mark Evans: The Wonderful World of Web 2.0 Whining), I'm one of those who wanted to support the site's developers, paying $89 a year to gain a premium version of the service last year, which gave me expanded access to a wider array of reports.

I'd like to say I don't check with SiteMeter throughout the day out of curiosity, but I'd be lying to you for sure. I love stats. I even made a dashboard widget for Mac OS X that shows me the day's activity, letting me just drag my mouse to the bottom right corner to get caught up. Except, today, I was surprised to see I was extremely unpopular. Not only was the total visit count much lower than I had anticipated, but it said absolutely nobody had checked in in the last hour. And since this morning, I've seen no updates at all.

SiteMeter's seen issues like this in the past. They operate not from one mega-database, but instead, each of its individual servers runs on its own database. When one has a hiccup, only those users on that single server show issues. I expect that's likely what's going on here, and just maybe, with luck, the total statistics will catch up overnight.

Now, we'll see just how much my going dark for about 36 hours over the weekend will have hurt me. With the company's blog not giving any hints as to what's happening, hope is all I have. Maybe it's time to check in with my FTP server and download my logs.

louisgray.com Experiences 100% Uptime During WWDC Keynote

Today, many of the popular sites aimed to deliver minute-by-minute updates to Apple CEO Steve Jobs' keynote, as well as some social networking sites extremely popular among the technology elite, slowed to a crawl under the rush of traffic from Mac and iPhone fans hoping to get a glimpse of Cupertino's latest products. Sites as diverse as TechCrunch to Twitter buckled under the pressure, while others, like MacRumors Live, Engadget and FriendFeed, maintained stability, gaining praise.

I am happy to report that louisgray.com enjoyed 100% uptime during this rush. Here's how we pulled off the enviable feat, all without removing services, reducing features, or requiring the offloading of some traffic to partners:

1. I Posted Absolutely Nothing At All

After much advance study, I realized that one of the major issues behind some of these sites who ran into trouble was in their offering of interesting content. Whether through rumors of new products in advance of the conference, live feeds during the conference, or reaction to the conference's announcements, each of the sites had attracted a population of users disproportionate to the norm, resulting in traffic spikes well above average, invariably causing slower access times or even downtime.

To avoid such a fate, not only did I not promise anything, but I didn't post anything, not just the morning of the keynote, but in the preceding 24 hours, essentially throwing potential visitors off the scent. I believe that this strategy, delivering an overall reduction, both day over day and week over week, in terms of total visitor traffic and page views, left the site with considerable headroom, and reduced chance of service interruptions.

2. I Made No Modifications to the Infrastructure

For several years now, louisgray.com has been hosted via FTP on Register.com hosted servers, powered by Google's Blogger engine. After considering many options available in advance of the WWDC keynote, it was determined the best course of action was again, nothing. Given the site's near 100% uptime over the last few years, despite significant year over year growth, the prevailing bias was to hold off on any significant software or hardware purchases which could cause complexity.

3. I Made No Advance Promises to Uptime

Murphy's Law dictates that anything that can go wrong will go wrong. Many were surprised as to Twitter's advanced promise of significant uptime during the keynote, after such a recent spotty track record. By promising 100% uptime for louisgray.com, I knew that I would, in turn, be placing myself in the line of fire for overzealous hackers and an overcaffienated faction from the Mac army, ready to take my site down like so many others.

Conclusion

I think there's no other option except to congratulate myself for delivery of 100% uptime during a time of considerable stress for tech media giants. Where they zigged like moths to the flame, I zagged away from the noise, bravely hiding in the corner, cowering in fear. You can count on louisgray.com to deliver the 100% uptime during such mega-tech events, both now and in the future thanks to our unique strategy of reverse traffic optimization. I hope we can count on your support.

June 06, 2008

Disqus' Downtime Reminds Us of Woes for Data In the Cloud

I am a happy Disqus customer. Implementing Disqus comments on this blog, enabling people to track their conversations around the Web, show personal custom avatars and thread conversations, has been among the better things I've done with the blog. Since installing Disqus, total comments have increased, I can get a better sense of my most frequent participants and they can connect one to one. My Disqus comments, and those of others, can even be shared on FriendFeed and other lifestreaming services.

My enthusiasm has not been unanimous across the blogosphere. Some have been concerned that Disqus' hosting the comments on their own site reduces the control a blogger has on this critical element of their site. Others say that Disqus effectively "steals" the SEO value of those comments, robbing you of the Google juice that's yours.

And to date, I've defended Disqus in every way. I'm not an SEO nut, so I can shrug my shoulders at these so-called issues. Until today.

Starting last night, I was surprised to find my e-mail empty of Disqus comments flowing to my in box. Checking the blog, I found many heartfelt comments on the passing of our dog yesterday. But Disqus wasn't sending me the updates. I logged in to the service, and ensured my preferences were set to notify me, and they were.

This morning, the situation is much worse. No comments are showing. The Disqus widget on the right side of the blog is missing. And every Disqus comment that every person posted on any Disqus-powered site is gone. This highlights the concern many have had on trusting the cloud and putting your data in the hands of others. It's always good to make a copy, especially if you don't know their infrastructure, or the company doesn't have a decades-long track record.

I trusted Disqus to host my comments, to run the show, to power my blog and to take on the challenging task of being my connection to my audience. Now, they're down hard. Their blog hasn't been updated to say what's going on, and the last update we got from Disqus' Daniel Ha is that he was playing poker 10 hours ago, via Twitter. I just hope he didn't bet the future of Disqus on a pair of 3's.

In this time where users are turning their data over to the cloud and trusting the underlying Web services, downtime can be a killer. The second half of responding to downtime? Transparency. And right now, Disqus is failing at both.

See also: FriendFeed discussion on Disqus downtime.

February 24, 2008

My Double Standard for Web Services

I don't play fair. I admit it.

The kind of miscues and errors that would get big headline, keyboard-pounding rants from me when the biggest of Web services fall short might instead get a pass if I know the service is run by a small handful of developers, or if I'm on a first name basis with the author. Instead of joining a chorus of complainers about why a service doesn't act the way I wanted it to, or implying they are unresponsive or nefarious in some way, I give them the benefit of the doubt.

Part of me wonders if this is just due to my own personal biases, or if I should expect companies that operate to the masses to perform at a higher level. Just as you would expect to get better service from a paid relationship than a free one, does it follow that a company with hundreds of thousands of users should be more tightly honed than one with a few dozen or a few hundred?

I was thinking of this over the weekend as on Saturday, I logged into AssetBar, and found, to my surprise that none of my feeds were updated. Peeking over at Google Reader, I knew that blogs were still being posted to, news was still being written, and keywords were still being discovered by search engines. But AssetBar lied to me and said I had nothing to read. A shame!

Given how fond I am of their service, and its potential, I could have jumped up and down, shaking my fist. But I didn't. Instead, I lobbed a quick note to the site's developers and said there had to be a glitch somewhere. No big deal. And sure enough, AssetBar posted a note to their blog saying they were updating the servers, which had caused my issue.

But if it were Google Reader who had gone hours without updates, there's no doubt I would likely have said something, and many others would have stood alongside me, calling them out. Just see our reactions when this type of thing has happened before:
Google Reader Down Overnight?
Google Reader Glitch Deletes Feeds: Blogosphere Weeps
Ack! Google Reader Update Wipes Out History
Now is that entirely fair? Probably not. Poor Google Reader team. I know they work hard and do a great job. But I also know that when it comes to smaller services just getting off the ground, like AssetBar, FriendFeed, LinkRiver, ReadBurner or RSSMeme, if they blow up something, or a key feature goes bump in the night, I'll likely give them a pass.

After all, ReadBurner did go down hard on February 1st, and took the entire site history away. (See: ReadBurner Down) When it did, I jokingly posted, Forget Twitter Issues... ReadBurner is Down!, and in the same post, gave Alexander Marktl praise for taking the opportunity to eliminate duplicates and add new features.

I never would have let Google get away with that, or Microsoft, YouTube, Apple, you name it. The big guys are held to higher standards, and always will be. It comes with the territory. That might not be fair, but that's the way it is.

February 21, 2008

Google FeedFetcher and FeedBurner Miss Each Other Again

This hasn't exactly been FeedBurner's best week.

First, the RSS syndication engine dropped all-time statistics due to a code error, and didn't quickly respond to user complaints, leading to questions from around the blogosphere, mine included. Making the problem worse was that the company's blog, once quite active, hadn't been updated in three months, most notably called out by Mashable.

Seemingly, both those issues were resolved, first, with the restoration of all-time stats, and second, seeing FeedBurner update their blog with a post, Hello? Hellooooo?, where they outlined their recent activity. First on their list? "Full integration with Google", no doubt much harder than it sounds.

But now, the very next day, it looks like Google's FeedFetcher, which reports how many RSS subscribers a blog has from Google Reader and iGoogle, didn't update FeedBurner. And around the blogosphere today, statistics are undoubtedly plunging. For example, louisgray.com saw subscribers more than cut in half, from 606 yesterday to 241 today, and ParisLemon plummeted, from 550 yesterday to 288 today. It's not the first time this has happened, but past instances have always led to promises of improved integration, and they're clearly not there yet.

It's enough to make me curious how RatingBurner is going to handle this data, once they synchronize their stats tomorrow.