Showing posts with label Service Alerts | Updates. Show all posts
Showing posts with label Service Alerts | Updates. Show all posts

Friday, May 4, 2012

SERVICE ALERT

calan servers are back up.

I will post an update once we have had a chance to review the hosting sites' logs

SERVICE ALERT

The calan servers are experiencing issues.
The team is addressing.

We will update you by 7:45AM EST

Wednesday, February 22, 2012

Report on Service Issues 2.15.2012


This report is intentionally written in layman’s terms. It is after all, being written by a layman, with respect to these more complex networking issues. We are available to discuss with our technical support the occurrences with anyone who would like a call in order to obtain any better understanding.

Here goes.

Tuesday at approximately 11:30PM our servers were completing a re-indexing to help retain operating efficiencies.

During this approximate two hour window, that this process takes, a NAS (Network Attached Storage) device failed. This device controls multiple network storage devices. It is identified as “enterprise level”, a designation reserved for devices that are as close to bullet proof as possible.

The device failed. Not only did it physically fail, as unlikely as that is to happen. It did so at the very moment that our servers were re-indexing.

The device had safe guards and a fail over that came on line. However the resulting impact to our db server was a corruption in the index. The index is the Dewey decimal filing system that helps the hard drives find and access data quickly.

This corruption to our index started an unseen internal loop where the server was trying to repair itself and dumping huge log files until the drive actually filled up, a second issue was created and at around 4AM on what would now be Wednesday, the server shut down.

An Emergency repair, an actual term for an internal SQL process, was implemented. At the same time our most recent backup, the db just before the 11PM incident was being reloaded to a second server. The restore of our backup to a second drive completed first, but we waited for the emergency repair sequence to complete. The primary server was still struggling; it finally completed the emergency repair, but was still not re-indexed. The re-index was now going to take more than the typical two hours as it was completely lost.

After waiting perhaps a bit too long, this solution was abandoned and plan B, the cut over to restored backup was implemented. The system operated reasonably well on this backup drive. Actually better than anticipated, only adding insult to the decision to wait for plan A to complete.

A decision was made to run on the reserve servers till the weekend. Operations were slow, but having missed several hours of operation already, we made the decision that access to slow data was better than access to no data at all. We limped into the weekend, where we felt the impact of completely pulling everything off line would be far less inconvenient.

As this was unfolding, the technical side of the issue was being addressed. A new NAS and upgraded NIC (Network Interface Controller) cards were ordered. The act of obtaining these hardware pieces and the necessary firmware upgrades required to make them fully compatible with our existing system took about 48 hours. This was a contributor to when we could make a full hardware replacement, conduct a total re-index and bring our primary server back on-line. This work was completed by late Sunday morning, with a couple of tweaks added late Monday night.

The system did not lose data. The backups did restore and the system was operable, to a degree despite the catastrophic hardware failure. The length of time in getting back on line and the limited performance until the full swap out on Sunday was part hardware related and a business decision, as mentioned.

The failed NAS has been replaced. New NIC cards have been installed, the db was fully re-indexed.

We realize the negative impact of the downtime. The failure of this NAS device is I am told very rare. The fact that it failed at such an inopportune moment in time is, well just a fact. There is no really explanation, nor excuse. It failed. It is replaced.

We took every step we could to protect the data. Immediately took actions to replace the bad hardware and complete the firmware upgrades. The decision to allow access to the system on Friday rather than go fully off line to conduct the needed repairs was completely ours.

To the best of our knowledge everything has been done to prevent a re-occurrence.
We have already scheduled a few calls with some of our sites to discuss this occurrence in greater detail. If you would like to schedule such a call please feel free to contact me directly.

Again our apology for the impact we had on your business operations.

Monday, February 20, 2012

Service Update: Weekend Maintenance


Work on both re-indexing and hardware issues were preformed over the weekend.

Given the holiday, I have asked for a call with the hosting and development team tomorrow to gather information on the specifics of what happened last week and what steps have been taken to prevent a reoccurrence.

I plan to post a more complete update by end of day tomorrow.

Friday, February 17, 2012

Service Update

We have identified the cause of the temporary slow down.

Until we take system off line for a full re-Index Saturday and then hardware fix Sunday we are being subjected to constant internal errors.

These error fill something called the "error log". As it becomes bloated the system slows.

We just dumped the logs and we are back to the early AM position. 

The index is corrupted. It will keep generating errors till we completely rebuild on Saturday AM. So errors will keep coming. 

We will monitor and dump as it gets bad. It is like bailing the row boat as you stroke to shore. 

You can stay afloat but until you get to dry land and can pull the boat completely out of the water for repairs progress is slow as you are bailing the entire time.

We are doing everything we can to avoid having your users lose total access again during the workweek.

SERVICE ISSUE

We are still seeing "after shocks " from the collapse earlier this week.
We were running and now are experiencing tremendous slow down in response.
We are on line with hosting firm.

Will provide update once we see why it is destabilizing.
This is related to our getting to the Saturday and Sunday scheduled work.

We are trying to hold the site open till then.

Thursday, February 16, 2012

Administration Update on Performance Issues


In addition to the hardware fixes this Sunday the 19th.

We will also be implementing a system sweep for bad records and deleted company records.
The 55M records in just one of our tables will reduce by approximately 25%.
This Table size was the primary reason the total re-index was taking so long to run.

We are unsure if it will manifest in actual improvement in daily performance but we want you to know that we are leaving no stone unturned in address all possible opportunities to prevent this kind of performance failure occurrence and improve daily operation.

We will be re-indexing early Saturday AM so service will be interrupted for several hours very early EST.

SERVICE UPDATE

Your Users can access the site
We are not fully indexed as we hoped but you have access
It will likely be somewhat slow

SERVICE ALERT UPDATE

As promised your 10AM EST update
We are going to plan B
Anticipated restore time is 30/45 minutes

BLACKOUT WINDOW UPDATE

OK stress levels are clearly off the chart for me.

The hardware repair work that will take the servers off-line will be:

THIS SUNDAY FEBRUARY 19th between 6AM till NOON EST.

SERVICE UPDATE

The servers are still re-indexing
They are re-indexing themselves, we have no direct work to do but watch.
It is consuming all the resources
They normally take 2/3 hours for this effort
They have been running since 4AM EST

The hardware failure caused all the index to be lost
We are going to switch to a backup if it does not clear by 10AM EST
That restore will take approximately 30 minutes
But we know that will put everyone in creep mode
But you will at least have access.
I will advise no later then that time of what twill occur

SERVICE UPDATE

calan service continues to run re-index
Process was started at 4AM EST
Some users have accessed and worked
But the majority are likely being denied access as the servers hold nearly all of the resources for its own use
It is progressing
We can not predict when it will complete

Maintenance Window BLACKOUT of SERVICE


As all are painfully aware, we suffered a major hardware failure late Tuesday night that resulted in a five hour interruption of service on Wednesday. With only marginal service restored for the balance of the day.

This morning we are still cleaning up the collateral mess that the failure created.

While our hosting site has fixed the original issue they plan to take additional steps to bolster those repairs and provide additional protection against such an occurrence happening again.

On SUNDAY MARCH 19 from 6AM till NOON EST CALAN WILL BE OFF-LINE.

This unplanned interruption is to allow for the existing hardware to be replaced and upgrades installed.

Please advise your users of this window.

The hosting site suggests repairs may go much faster but given the desire to see this work completed as soon as possible without any additional surprises we are blocking a six hour window.

Our apologies for this inconvenience.