Zum Inhalt springen

Auch Netstream hatte im November kein Glück

1 Antworten1'109 AufrufeGestartet von tosci ·

tosciThemenstart5'731 Beiträge
#1
Wie es der Titel schon sagt, auch bei Netstream lief wohl einiges schief - habe davon bis auf den heutigen Newsletter eigentlich gar nichts mitbekommen. Aber es ist immer wieder interessant wie die angeblich redundanten und sicheren Systeme grausam abschmieren:

Am 25. November 2010 waren bedeutende Netstream Dienste für mehrere Stunden offline. Nach dem Ausfall unseres primären Speicherlaufwerkes fiel kurz danach auch das sekundäre Speicherlaufwerk aus. In der Folge kam es zum schwersten und längsten Systemausfall in der dreizehnjährigen Geschichte von Netstream.

Wir werden umfassende Vorkehrungen treffen, damit ein solcher Vorfall nicht noch einmal passiert. Detaillierte Informationen zu den ergriffenen Massnahmen werden wir Ihnen in einem der folgenden Newsletter präsentieren.

Wir entschuldigen uns für alle Probleme, die durch den Ausfall verursacht wurden. Den Rapport der Grossstörung finden Sie unter:

Major Outage

14:35 Our systems detect a failure on the primary storage devices. The system automatically switches successfully to the secondary storage group

14:40 Investigations on the failure of the primary storage device starts.

14:55 Secondary storage group fails

14:55 Begin of outage

15:25 Engineers On-Site

15:40 Escalation to Storage supplier Dell

16:10 Escalation to the highest severity level for Dell

17:00 Issue has still not been localized. Analysis is ongoing

17:30 Issue has still not been localized. Analysis is ongoing

18:00 A Storage controller has been replaced

18:30 Additional Storage Hardware has been replaced.

19:00 Issue has still not been localized. Analysis is ongoing

19:30 Issue has still not been localized. Analysis is ongoing.

20:00 Our engineers do analyze every server before booting up the service

21:00 Service status is getting back to normal

22:00 All services are back online

Although all systems are built fully redundant and have an average availability of at least 99.99% this outage has produced the longest and broadest outage for services in the 13 year history of Netstream. We apologize for all problems caused due this outage. We will deeply analyze the root cause of this issue and do whatever is needed with both financial and human resources in order to avoid that such kind of outage will never ever happen again. Furthermore we will increase our effort for communications in cases of emergencies.

Betrifft Hosted Exchange, NetVS, Hosted PBX, Virtual PBX, Wholesale VoIP, Primary Webcluster, Hosted Sharepoint, Netstream AG, partial xDSL, mail.2wire.ch

smid12'144 Beiträge
#2
Leider kann man den wirklichen Notfall nur selten wirklich testen.

Seite 1 von 1 · 2 Beiträge in diesem Thema