calan servers are back up.
I will post an update once we have had a chance to review the hosting sites' logs
Showing posts with label Service Alerts | Updates. Show all posts
Showing posts with label Service Alerts | Updates. Show all posts
Friday, May 4, 2012
SERVICE ALERT
The calan servers are experiencing issues.
The team is addressing.
We will update you by 7:45AM EST
The team is addressing.
We will update you by 7:45AM EST
Wednesday, February 22, 2012
Report on Service Issues 2.15.2012
This report is intentionally written in layman’s terms. It is
after all, being written by a layman, with respect to these more complex
networking issues. We are available to discuss with our technical support the occurrences
with anyone who would like a call in order to obtain any better understanding.
Here goes.
Tuesday at approximately 11:30PM our servers were completing
a re-indexing to help retain operating efficiencies.
During this approximate two hour window, that this process
takes, a NAS (Network Attached Storage) device failed. This device controls
multiple network storage devices. It is identified as “enterprise level”,
a designation reserved for devices that are as close to bullet proof as
possible.
The device failed. Not only did it physically fail, as
unlikely as that is to happen. It did so at the very moment that our servers
were re-indexing.
The device had safe guards and a fail over that came on
line. However the resulting impact to our db server was a corruption in the
index. The index is the Dewey decimal filing system that helps the hard drives
find and access data quickly.
This corruption to our index started an unseen internal loop
where the server was trying to repair itself and dumping huge log files until the
drive actually filled up, a second issue was created and at around 4AM on
what would now be Wednesday, the server shut down.
An Emergency repair, an actual term for an internal SQL
process, was implemented. At the same time our most recent backup, the db just
before the 11PM incident was being reloaded to a second server. The restore of
our backup to a second drive completed first, but we waited for the emergency repair
sequence to complete. The primary server was still struggling; it finally
completed the emergency repair, but was still not re-indexed. The re-index was
now going to take more than the typical two hours as it was completely lost.
After waiting perhaps a bit too long, this solution was abandoned
and plan B, the cut over to restored backup was implemented. The system
operated reasonably well on this backup drive. Actually better than anticipated,
only adding insult to the decision to wait for plan A to complete.
A decision was made to run on the reserve servers till the
weekend. Operations were slow, but having missed several hours of operation already,
we made the decision that access to slow data was better than access to no data
at all. We limped into the weekend, where we felt the impact of completely
pulling everything off line would be far less inconvenient.
As this was unfolding, the technical side of the issue was
being addressed. A new NAS and upgraded NIC (Network Interface Controller) cards
were ordered. The act of obtaining these hardware pieces and the necessary
firmware upgrades required to make them fully compatible with our existing
system took about 48 hours. This was a contributor to when we could make a full
hardware replacement, conduct a total re-index and bring our primary server
back on-line. This work was completed by late Sunday morning, with a couple of
tweaks added late Monday night.
The system did not lose data. The backups did restore and the
system was operable, to a degree despite the catastrophic hardware failure. The length of time in getting back on line and the limited
performance until the full swap out on Sunday was part hardware related and a
business decision, as mentioned.
The failed NAS has been replaced. New NIC cards have been
installed, the db was fully re-indexed.
We realize the negative impact of the downtime. The failure
of this NAS device is I am told very rare. The fact that it failed at such an inopportune
moment in time is, well just a fact. There is no really explanation, nor
excuse. It failed. It is replaced.
We took every step we could to protect the data. Immediately
took actions to replace the bad hardware and complete the firmware upgrades.
The decision to allow access to the system on Friday rather than go fully off
line to conduct the needed repairs was completely ours.
To the best of our knowledge everything has been done to
prevent a re-occurrence.
We have already scheduled a few calls with some of our sites
to discuss this occurrence in greater detail. If you would like to schedule
such a call please feel free to contact me directly.
Again our apology for the impact we had on your business
operations.
Monday, February 20, 2012
Service Update: Weekend Maintenance
Work on both re-indexing and hardware issues were preformed
over the weekend.
Given the holiday, I have asked for a call with the hosting
and development team tomorrow to gather information on the specifics of what
happened last week and what steps have been taken to prevent a reoccurrence.
I plan to post a more complete update by end of day
tomorrow.
Friday, February 17, 2012
Service Update
We have identified the cause of the temporary slow down.
Until we take system off line for a full re-Index Saturday and then hardware fix Sunday we are being subjected to constant internal errors.
These error fill something called the "error log". As it becomes bloated the system slows.
We just dumped the logs and we are back to the early AM position.
We will monitor and dump as it gets bad. It is like bailing the row boat as you stroke to shore.
You can stay afloat but until you get to dry land and can pull the boat completely out of the water for repairs progress is slow as you are bailing the entire time.
We are doing everything we can to avoid having your users lose total access again during the workweek.
SERVICE ISSUE
We are still seeing "after shocks " from the collapse earlier this week.
We were running and now are experiencing tremendous slow down in response.
We are on line with hosting firm.
Will provide update once we see why it is destabilizing.
This is related to our getting to the Saturday and Sunday scheduled work.
We are trying to hold the site open till then.
We were running and now are experiencing tremendous slow down in response.
We are on line with hosting firm.
Will provide update once we see why it is destabilizing.
This is related to our getting to the Saturday and Sunday scheduled work.
We are trying to hold the site open till then.
Thursday, February 16, 2012
Administration Update on Performance Issues
In addition to the hardware fixes this Sunday the 19th.
We will also be implementing a system sweep for bad records
and deleted company records.
The 55M records in just one of our tables will reduce by
approximately 25%.
This Table size was the primary reason the total re-index
was taking so long to run.
We are unsure if it will manifest in actual improvement in
daily performance but we want you to know that we are leaving no stone unturned
in address all possible opportunities to prevent this kind of performance
failure occurrence and improve daily operation.
We will be re-indexing early Saturday AM so service will be interrupted for several hours very early EST.
SERVICE UPDATE
Your Users can access the site
We are not fully indexed as we hoped but you have access
It will likely be somewhat slow
We are not fully indexed as we hoped but you have access
It will likely be somewhat slow
SERVICE ALERT UPDATE
As promised your 10AM EST update
We are going to plan B
Anticipated restore time is 30/45 minutes
We are going to plan B
Anticipated restore time is 30/45 minutes
BLACKOUT WINDOW UPDATE
OK stress levels are clearly off the chart for me.
The hardware repair work that will take the servers off-line will be:
SERVICE UPDATE
The servers are still re-indexing
They are re-indexing themselves, we have no direct work to do but watch.
It is consuming all the
resources
They normally take 2/3 hours for
this effort
They have been running since 4AM EST
The hardware failure caused all
the index to be lost
We are going to switch to a backup if it does not clear by 10AM EST
That restore will take approximately 30 minutes
But we know that will put
everyone in creep mode
But you will at least have access.
I will advise no later then that time of what twill occur
SERVICE UPDATE
calan service continues to run re-index
Process was started at 4AM EST
Some users have accessed and worked
But the majority are likely being denied access as the servers hold nearly all of the resources for its own use
It is progressing
We can not predict when it will complete
Process was started at 4AM EST
Some users have accessed and worked
But the majority are likely being denied access as the servers hold nearly all of the resources for its own use
It is progressing
We can not predict when it will complete
Maintenance Window BLACKOUT of SERVICE
As all are painfully aware, we suffered a major hardware
failure late Tuesday night that resulted in a five hour interruption of service
on Wednesday. With only marginal service restored for the balance of the day.
This morning we are still cleaning up the collateral mess
that the failure created.
While our hosting site has fixed the original issue they plan
to take additional steps to bolster those repairs and provide additional
protection against such an occurrence happening again.
On SUNDAY MARCH 19 from 6AM till NOON EST CALAN WILL BE
OFF-LINE.
This unplanned interruption is to allow for the existing
hardware to be replaced and upgrades installed.
Please advise your users of this window.
The hosting site suggests repairs may go much faster but
given the desire to see this work completed as soon as possible without
any additional surprises we are blocking a six hour window.
Our apologies for this inconvenience.
Subscribe to:
Posts (Atom)