IsvaraIsvara
The Guide
Executive Summary/Introduction

The reason some customers run their infrastructure via complaint-based operations is because the operations team has no other means by which to measure success. They have not defined the acceptable performance of their infrastructure and have no benchmark for “good”. Solving this challenge is one of the goals of this whitepaper.

Does troubleshooting mean all hands-on deck?

If a troubleshooting event means all hands-on deck, that indicates that you don’t have the process or data required to triage a problem and engage specific specialists, so you engage everybody. Do you have a troubleshooting process that is followed by all teams (including network, storage, server, OS, application, etc.)? Does that process end with Root Cause Analysis (RCA)?

As part of RCA, do you set up alerts so the same issue can be detected faster if it happens again? Without an alert configured, the RCA is not complete.

Do Help Desk support tickets often require escalation?

If Help Desk simply passes issues through to the next level, you need to look at why.

Help Desk is your first line of defense. They do not go as technically deep as the specialists higher in the support framework. Equip them with simple dashboards so that they can handle complaints by proving:

  • Is the problem caused by the infrastructure not serving the VM well?

  • If yes, which part of the infrastructure? Is the problem at the CPU, memory, disk, or network layers?

  • If not, how can we prove this convincingly to the application owners?

Is proving the cost effectiveness of private cloud a challenge?

The commoditization of infrastructure means your private cloud is being compared with public cloud platforms like Amazon AWS, Microsoft Azure, and Google Cloud.

If your private cloud is not demonstrably cheaper and better, or if you cannot measure cost on a per-workload basis at all, the business may question the value of the private cloud. One of the primary reasons for running a private cloud, alongside privacy, compliance, and security, is cost-effectiveness. This cost equation must also include the cost of the staff and facilities required to operate the platform.

Do you worry about running out of capacity in the private cloud?

The expectations of your customers of frictionless consumption of IT services to build and run their applications means that IT cannot exist as blocker to protect the capacity of the infrastructure. Public cloud, which private cloud is often compared against, has the perception of being able to scale endlessly. Public clouds benefit from economies of scale in this respect, so an effective capacity management process is of paramount importance for a private cloud, particularly when considering the lead times of purchasing and provisioning new hardware. This whitepaper will help to give you confidence around capacity by understanding the current capacity of your private cloud and enable you to forecast future consumption against current and future capacity.

Do you struggle with over-provisioned VMs?

This is an indicator that you are operating in a ‘system builder’ function and not as a service provider. As a system builder, you are touching and customizing individual VMs. You size them and argue with the application teams, who are the customers/consumers of the infrastructure. As a result, you are busy as there are many applications and you are outnumbered.

If you are operating as a service provider of private cloud to the business, you should not be “in the way” of the business. You should be using an effective pricing model to drive the right behavior. Does a public cloud provider block customers from buying a 40 CPU VMs when they only need 2 CPU? Of course not.

This does not mean there is no value in “right-sizing” workloads and helping to drive efficient consumption by your internal customers. The tools described in this whitepaper can help you understand how to do that too.

What does “good” look like?

Underpinning a private cloud is effective and proactive management of performance, capacity, and cost within the infrastructure, backed up by SLAs that explicitly define what level of service your customers can expect.

SLAs are a sign of matured, cloud-like operations. SLAs for the service you provide, which is likely the software-defined-data center and its associated components, must be complete, correct, and accurate

Figure 1: Overview of Service Level Agreements

Complete means you have SLAs for performance – and potentially for compliance -, not just availability (which is the most common form of SLA in the world of IT). Performance and compliance are crucial components of the overall service you are providing. There is limited value in ensuring a workload is available if its performance is so poor that the application running on it is unusable, or if an environment is non-compliant leading to a security incident or a data breech.

Correct means the SLA is measured on each paying VM, and not at the infrastructure level, because ultimately the measure for success is not the health of the infrastructure platform, but the health of the workloads which are running the applications. Correct also means you are using the right metrics to track the health and performance of your service.

Accurate means the measurement must be measured/collected every 5 minutes. Longer intervals than this don’t provide the granularity you need to catch problems. Shorter intervals for collection causes impacts at the infrastructure layers for the collection, processing, and storage of the additional datapoints.

To support your infrastructure and operations teams in ensuring the service meets the defined SLAs, you will have SLA Leading Indicators, which provide you a forward-looking prediction of how a service is tracking towards its SLA. You will also have Key Performance Indicators (KPIs) which are a useful way of condensing the numerous metrics which give valuable insight into the performance of a workload (or cluster) across the different resource types, into a single score which makes performance health more visible and aids proactive troubleshooting.

This is a journey and maturity in these areas must be built step-by-step. This whitepaper is written looking from the top down as this provides a better conceptual understanding of what we’re trying to achieve and provides context for a subsequent focus on some of the lower-level details. It’s a bit like showing someone a house: you start with the conceptual-level information like how big the house is and what rooms and features it has. However, also like a house, to embark upon the journey and build maturity in these areas, you need to start from the bottom up.

Service Level Agreements

One of the key differences between a virtual infrastructure platform and a Private Cloud is the SLA. A cloud provider can state that they have the best technology, the best staff, the most innovative processes, the most industry certifications, etc., to prove that their product is the best, but none of that carries weight because it’s not contractual. If you are providing a service to your business, your SLA is the definition of what that service actually is. The SLA enables operations teams to hold themselves accountable to their customers because the SLA carries financial penalties. Once the SLA is defined, only then will customers want to know how it will be delivered. This is where those elements mentioned earlier, like processes, architecture, certifications etc., come into play.

Anatomy of an SLA

So, what’s in the SLA? The SLA, being a contractual document, will contain service definitions and descriptions that inform the customer on all sorts of details about the service. It should provide a thorough description of what the service is that your customers are receiving from you. As an example, the SLA for a VM within a Private Cloud service may contain description of the VM, it’s resource size and hardware capabilities, as well as details about backups, disaster recovery, patching schedules and maintenance windows, and other features and capabilities that your business has decided a virtual machine (VM) running within its organizational boundary must have.

Much of this content will be specific to your business and the service that you have designed.

An SLA also contains another fundamental piece: the Service Level Objective.

← Back
The Guide
Home
Next
What does “good” look like?