Showing posts with label ITIL. Show all posts
Showing posts with label ITIL. Show all posts

Saturday, 29 November 2014

Managing Your DataCenter Physical Infrastructure - Part 3

4.       Change Management – It’s the process of making changes in environment with methods or procedure with lowest possible impacts on business. Every change requires preparation, planning, simulation for expected results and verification. As it involves business, every change requires approval from stake-holders too
a.      Why: To minimize the unplanned downtimes
b. Challenges: executing movement/addition/changes without impacting availability or business, maintaining compatibility within eco-system of IT Infrastructure including spares while implementing upgrades or changes
c.      Solution:
                                                              i.      A plan which explains the work-flow for the changes to be performed
                                                            ii.      A proper change management, with business approved downtimes to perform change and to avoid business impacts like off-hours or weekends. A redundant infra or resources would be required to perform change without downtime, which adds to CAPEX and OPEX both
                                                          iii.      HCL or Hardware Compatibility Lists should be referred along with Inter-portability matrix to verify compatibility of components with other components post upgrade e.g. Storage firmware level of Brocade or Cisco SAN switch with Host HBA card or Storage Controller card should be verified before performing firmware upgrade. There would be a separate article to give you some over-view on Change Tech Plans in DataCenter.
                                                           iv.      It is a best practice to maintain even with firmware level when upgrading infrastructure.
                                                             v.      Some examples of Changes required in a datacenter can be relocating the server, patch-panel re-wiring, firmware or software upgrades, addition of resources or components, replacement of faulty components etc. Configuration changes are also considered under change management.

Putting Pieces Together: Most of the Business organization takes the following initiative as their strategy to meet above
·         Implementation of Incident Management process
·         Defining and Measure Availability Targets
·         Monitor and Plan future Road-Maps for Long Term Capacity Plan
·         Implementation of Change Management process

Special Note:
·         It may be difficult to find single tool to manage complete or all layer of equipment in your datacenter i.e. rack space, power/cooling, server, storage, network and application. Usually Datacenter are hybrid and have multi-vendor solution. Integration or coordination between solutions by these vendors may be difficult. 
·         20% Improvement during Designing Phase can remove 80% of the problems which may be caused by them in future, now this statement doesn’t cover only the DataCenter but the process or methods and procedures as well.
·         As observed, there may be overlap of responsibility between different departments like facility and IT, which may cause conflicts. Management decisions would be required for defining roles in such scenarios. It’s the technology, process and people, which together keeps the business up and running
·         Even though Physical Infra or Enterprise Management systems (EMS) sounds similar to Building Management system (BMS) due to components being managed or monitored under them i.e. space/power/cooling, but the focus of both are different. Physical Infra focus on availability of business, while BMS focus on comfort and safety. Integration of both EMS and BMS could be very costly.

Any more questions? please write back or comment here. There are more things to share.. 

Request you to join my group on Facebook & LinkedIN with name "DataCenterPro" to get regular updates. I am also available on Tweeter as @_anubhavjain. Shortly I am going to launch my own YouTube channel for free training videos on different technologies as well. 

Happy Learning!!

Managing Your DataCenter Physical Infrastructure - Part 2

2.       Availability Management – Identifying Availability & Reliability requirements against actual performance  and if required introduce improvement to meet and sustain quality of service
a.       Why: Once availability is defined, SLA should be monitored to analyze the potential downtimes from impact of any individual component or entire system
b.   Challenges: Metrics Reporting, Raising alarms, Planned Downtimes  and Continuous Infra Improvements
c.       Solution:
                                                              i.      A Tool which reports uptime/downtime of infrastructure or service, while identifying the cause of downtime, providing the time-stamp or duration and time it took for recovery. It’s a best practice to configure tools to provide  Pro-active warning
                                                            ii.      A Tool, which doesn’t require special training or expert e.g UPS or Battery health, temperature or humidity, Disk or Power status etc.
                                                          iii.      Unplanned downtime can lead to false alerts. Using Maintenance mode in system units e.g.VMware ESXi, Storage Array Controllers, UPS or Blade Servers etc; A usual mistake often happens is to rollback all changes made to put a system into maintenance, hence need to be done with caution. It is suggested to use tool which provide alert if any condition is left uncorrected post maintenance.
                                                           iv.      To make improvements, the very first step is identify its need along with Risk involved in bringing up the change. Now FMEA techniques may not be known to everyone, hence its is suggested to use a tool which does risk assessment. You can even use health reports as you reference points to bring in improvement. Corrective Measures to Mitigate these Risk will become your Improvement Plan. Note: Improvement should be continuous. Some examples can be: power consumption, Disk full status, Performance reports, Cluster Loads etc.

3.       Capacity Management – providing IT resources as when required at right cost.
a.    Why: Current and Future requirements keep changing and need to monitoring and addressed
b.   Challenges: Asset Management with on-going changes (monitoring, recording, tracking), providing capacity as when required, Optimizing capacity for more ROI and better management, Incremental Scalability
c.   Risk involved: unplanned downtime if resources over-utilized e.g. Power, CPU/RAM in ESXi, Network Bandwidth etc
d.     Solution:
                                                              i.      A tool that performs centralize monitoring for current usage and alerts upon potential over-load of resources e.g. Power & Cooling Monitoring Systems for DataCenter by Emerson/ APC /Schneider-electric, Network Bandwidth Monitor, vCOPs for VMware Clusters, Storage Performance Indicators, HP Insight Managers for HP Blade enclosures or Servers, Brocade Fabric Managers etc.
                                                            ii.      Capacity requirements are tend to miss or not considered during implementation. Hence a tool is required for Trending Analysis , which also alerts on Threshold violations or over-loads. This tool should be referred before going ahead with Future procurements or new deployments.
                                                          iii.      A poorly designed datacenter may requires more resources (server, storage, network, space, power) (High CAPEX) and hence would cost more to operate (High OPEX). Analysis of requirements should be done during designing of datacenter or even during new deployments. It is good to implement Six Sigma DMADV techniques if possible.
                                                           iv.      Weekly/Month/Quarterly reviews on current capacity and usage trends will forecast the future required scalability too. Ideally while designing a new datacenter, every single component (server, storage, network, virtualization platform, space, power) are designed with such a scale that they can either bear 30% incremental capacity (scalability) and should be operable for atleast next 3 years. Even support contracts are considered in the same manner.
                                                             v.      In terms of DataCenter; Location, Power (input, socket), Cooling, Rack Space and Cabling are the major requirement and consideration in terms of Capacity Management; while in terms of Network, Network ports, bandwidth, VLAN, IPs etc. can be considered. In terms of Servers  (Physical or Virtual) , CPU/cores, RAM, Cluster etc can be considered; while for Storage, type of connectivity (FC, NFS, iSCSI, FCoE, DAS), Space required, IOPS, Backup, Recovery & Redundancy options (RAID, snapshots, Clones, replication) etc should be considered.
                                                           vi.      Usually Life Cycle or Capacity Manager track the inventory and usage of their assets, which is a best practice as well. It is also suggested to visualize the impact on capacity with every new deployment.

                                                         vii.      Note that TCO & ROI need to be considered and is a deciding factor when it comes to business decisions. 

Continue to Read..

Part 1: Managing Your DataCenter
                             Part 3: Managing Your DataCenter

Managing Your DataCenter Physical Infrastructure - Part 1

We have been talking a lot about DataCenter, how to Manage it? Every DataCenter has a physical infrastructure involved in it, which, by definition, includes:

        i.            Power
      ii.            Cooling
    iii.            Racks and physical structure
     iv.            Cabling
       v.            Physical security and fire protection
     vi.            Management systems
   vii.        Services


To manage all these layer, ITIL framework has been in use, considering all aspects of it. Now,  ITIL (Information Technology Information Library)  is Not a Standard but Framework; and you implement pieces which are relevant to your business. ITIL consist of two aspects; one is Service Support Process, where focus  is on End User and other is Service Delivery Process, where focus is on Business owners. They include following management process under them:

  1. Service Support process - Incident, Problem, Change, Release, Configuration Managements
  2. Service Delivery process - Service Level, IT Service Community, IT Financial, Capacity, Availability


Now, I wont be covering the in-depth study about ITIL, that’s a different discussion all together. While Managing DataCenter only following process:- Incident, Availability, Capacity and Change Management. All process are inter-related via process flows
A brief description about each process w.r.t. DataCenter is mentioned below:

1.       Incident Management – bringing business back to normal with minimum impact on business
a.    Why: Monitor events alarms of physical infra such as of hardware, network, power etc
b.     Challenges: identifying location, owner of incident, identifying severity, and executing action plan to fix the problem
c.     Solution:
                                                              i.      System level view of inter-related components will give overview of location along with impact of individual components.
                                                           ii.      Its best practice to mention system owners in inventory/asset management tracker. E.g. Blade enclosures OA gives you option to mention its rack details along with Point of Contact details. UID lights are also helpful in remotely identifying the equipment with the help of local hands and feet support. Note that responsibility is often shared so as to avoid single point of failure. ARCI (Accountable, Responsible, Consultant, Inform) matrix should be defined and available as reference in case of any Incidents. Its also helpful when you need approval to apply break-fix solution or driving change.

                                                          iii.      It’s a good practices to use system defined alerts (High Medium Normal) along with Business SLA to as reference points while defining the prioritization of Incident. 

Continue to Read..

Part 2: Managing Your DataCenter
                             Part 3: Managing Your DataCenter