July 02, 2012
Last Friday, in what has become a not-so-unusual occurrence for the cloud services provider, Amazon's Elastic Compute Cloud went dark. The event, lasting two hours, happened just weeks after a power outage knocked the company's US-EAST-1 region offline for roughly six hours.
Both failures originated from Amazon operations based in Northern Virginia. Ars Technica detailed the earlier event, which was the result of primary power, primary backup power, and secondary backup power failures. Amazon explained that the issue began with a cable fault, disconnecting their primary power source. Shortly thereafter, the primary backup generator failed due to a faulty cooling fan. The Secondary backup was also inoperable due to an incorrectly configured circuit breaker.
Amazon had since promised that circuit breaker configuration would become part of their auditing process, but the message was met with some valid skepticism by Ars.
So, the breakers are fixed, but it's hard to imagine there won't be other problems in the future.
Surely enough, another power related event knocked out the US-EAST-1 region. This time, operations were affected by a major storm that left roughly 400,000 people without electricity. The issue took down websites Instagram, Pinterest, Heroku and Netflix for 2-3 hours.
Netflix is a prominent user of Amazon Web Services and is fully aware that the cloud provider is not infallible. Last April, their website famously stayed online during a major EC2 outage that took down Reddit, Quora, Hootsuite and Foursquare among others. Following that event, Netflix explained how their service stayed online during the Amazon outage.
Why were some websites impacted while others were not? For Netflix, the short answer is that our systems are designed explicitly for these sorts of failures. When we re-designed for the cloud this Amazon failure was exactly the sort of issue that we wanted to be resilient to.
Unfortunately, Friday's event took down the video streaming site as well. As of now, the Netflix tech blog has not posted a breakdown of the event. PC Mag received a vague explanation from a Netflix representative, saying the downtime was the result of a "rare technical issue that our engineers fixed."
The recent events demonstrate how fragile some portions of the Internet can be. They also act as a wake-up call to services relying on Amazon. HP, Microsoft and most recently Google, have entered the public cloud game, offering alternatives to EC2. If reliability continues to hinder the cloud giant, these competitors may be more than willing to tempt some of its current customers away.
Frank Ding, engineering analysis & technical computing manager at Simpson Strong-Tie, discussed the advantages of utilizing the cloud for occasional scientific computing, identified the obstacles to doing so, and proposed workarounds to some of those obstacles.
Read more...
The private industry least likely to adopt public cloud services for data storage are financial institutions. Holding the most sensitive and heavily-regulated of data types, personal financial information, banks and similar institutions are mostly moving towards private cloud services – and doing so at great cost.
Read more...
In this week's hand-picked assortment, researchers explore the path to more energy-efficient cloud datacenters, investigate new frameworks and runtime environments that are compatible with Windows Azure, and design a unified programming model for diverse data-intensive cloud computing paradigms.
Read more...
05/10/2013 | Cleversafe, Cray, DDN, NetApp, & Panasas | From Wall Street to Hollywood, drug discovery to homeland security, companies and organizations of all sizes and stripes are coming face to face with the challenges – and opportunities – afforded by Big Data. Before anyone can utilize these extraordinary data repositories, however, they must first harness and manage their data stores, and do so utilizing technologies that underscore affordability, security, and scalability.
04/02/2012 | AMD | Developers today are just beginning to explore the potential of heterogeneous computing, but the potential for this new paradigm is huge. This brief article reviews how the technology might impact a range of application development areas, including client experiences and cloud-based data management. As platforms like OpenCL continue to evolve, the benefits of heterogeneous computing will become even more accessible. Use this quick article to jump-start your own thinking on heterogeneous computing.