HPC in the Cloud


Dedicated to covering high-end cloud computing
in science, industry and the datacenter

Language Flags

Amazon Bit the Dust... Again


Last Friday, in what has become a not-so-unusual occurrence for the cloud services provider, Amazon's Elastic Compute Cloud went dark. The event, lasting two hours, happened just weeks after a power outage knocked the company's US-EAST-1 region offline for roughly six hours.

Both failures originated from Amazon operations based in Northern Virginia. Ars Technica detailed the earlier event, which was the result of primary power, primary backup power, and secondary backup power failures. Amazon explained that the issue began with a cable fault, disconnecting their primary power source. Shortly thereafter, the primary backup generator failed due to a faulty cooling fan. The Secondary backup was also inoperable due to an incorrectly configured circuit breaker.

Amazon had since promised that circuit breaker configuration would become part of their auditing process, but the message was met with some valid skepticism by Ars.

So, the breakers are fixed, but it's hard to imagine there won't be other problems in the future.

Surely enough, another power related event knocked out the US-EAST-1 region. This time, operations were affected by a major storm that left roughly 400,000 people without electricity. The issue took down websites Instagram, Pinterest, Heroku and Netflix for 2-3 hours.

Netflix is a prominent user of Amazon Web Services and is fully aware that the cloud provider is not infallible. Last April, their website famously stayed online during a major EC2 outage that took down Reddit, Quora, Hootsuite and Foursquare among others. Following that event, Netflix explained how their service stayed online during the Amazon outage.

Why were some websites impacted while others were not? For Netflix, the short answer is that our systems are designed explicitly for these sorts of failures. When we re-designed for the cloud this Amazon failure was exactly the sort of issue that we wanted to be resilient to.

Unfortunately, Friday's event took down the video streaming site as well. As of now, the Netflix tech blog has not posted a breakdown of the event. PC Mag received a vague explanation from a Netflix representative, saying the downtime was the result of a "rare technical issue that our engineers fixed."

The recent events demonstrate how fragile some portions of the Internet can be. They also act as a wake-up call to services relying on Amazon. HP, Microsoft and most recently Google, have entered the public cloud game, offering alternatives to EC2. If reliability continues to hinder the cloud giant, these competitors may be more than willing to tempt some of its current customers away.

Most Read Blogs

Aspen

Feature Articles

CometCloud: Using a Federated HPC-Cloud to Understand Fluid Flow in Microchannels

The ever-growing complexity of scientific and engineering problems continues to pose new computational challenges. Thus, we present a novel federation model that enables end-users with the ability to aggregate heterogeneous resource scale problems. The feasibility of this federation model has been proven, in the context of the UberCloud HPC Experiment, by gathering the most comprehensive information to date on the effects of pillars on microfluid channel flow.
Read more...

CERN, Google, and the Future of Global Science Initiatives

Large-scale, worldwide scientific initiatives rely on some cloud-based system to both coordinate efforts and manage computational efforts at peak times that cannot be contained within the combined in-house HPC resources. Last week at Google I/O, Brookhaven National Lab’s Sergey Panitkin discussed the role of the Google Compute Engine in providing computational support to ATLAS, a detector of high-energy particles at the Large Hadron Collider (LHC).
Read more...

Avoiding Scientific Computing Bottlenecks in the Cloud

Frank Ding, engineering analysis & technical computing manager at Simpson Strong-Tie, discussed the advantages of utilizing the cloud for occasional scientific computing, identified the obstacles to doing so, and proposed workarounds to some of those obstacles.
Read more...

Sponsored Whitepapers

Best Practices in Big Data Storage

05/10/2013 | Cleversafe, Cray, DDN, NetApp, & Panasas | From Wall Street to Hollywood, drug discovery to homeland security, companies and organizations of all sizes and stripes are coming face to face with the challenges – and opportunities – afforded by Big Data. Before anyone can utilize these extraordinary data repositories, however, they must first harness and manage their data stores, and do so utilizing technologies that underscore affordability, security, and scalability.

Exploring the Potential of Heterogeneous Computing

04/02/2012 | AMD | Developers today are just beginning to explore the potential of heterogeneous computing, but the potential for this new paradigm is huge. This brief article reviews how the technology might impact a range of application development areas, including client experiences and cloud-based data management. As platforms like OpenCL continue to evolve, the benefits of heterogeneous computing will become even more accessible. Use this quick article to jump-start your own thinking on heterogeneous computing.

Sponsored Multimedias

Newsletters

Stay informed! Subscribe to HPC in the Cloud email Newsletters.

HPC in the Cloud Update
HPCwire Weekly Update
Digital Manufacturing Report
Datanami
HPCwire Conferences & Events
Job Bank
HPCwire Product Showcases


ISC

HPC Job Bank


Featured Events



  • June 16, 2013 - June 20, 2013
    ISC'13
    Leipzig,
    Germany

  • June 17, 2013 - June 18, 2013
    Forecast 2013
    San Francisco, CA
    United States




HPC in the Cloud Conferences & Events