Translate

March 28, 2019

Blame prevention through devops

This post is a follow up of my previous post The question that takes away all blame.
Blameless postmortems, or blameless RCA’s are supposed to be the new-normal in devops organisations, but all too often we see that first the team and sometimes the person to blame is sought, and then we tell them to ‘fix it’.
You might’ve noticed that I wrote devops in all lower case in this post's title. I did that on purpose.

Devops sanse capital 'd'


Even though you typically see DevOps instead of devops. I think that DevOps implies that it’s about Development and Operations engineers working together. In my opinion, devops is about combining the responsibility of the development team and the responsibility of operations team, turning them into the responsibility of a single team. This would make devops a matter responsibility and less of an organisational concern.

The organisational aspect would then be in the form of ‘Product Teams’, responsible for a product with a Product Owner that is accountable for that product.
Something for another day. This article is about blameless post-mortems and root cause analysis driven through devops-tinted glasses.
Blameless post-mortems are something that comes more natural in environments of shared responsibilities. Environments where the same people are responsible for both the quality of the product as well its usage. Environments where devops is considered the combined responsibility of development and operations within a single team.

Silos vs Accountability

I have observed in a number of organisations that one of the main reasons these organisations are considering the move towards devops is based around the concept of shared responsibility. It is the idea that silos prevent this sharing of responsibility. It is a misconception though. Silos don't prevent shared responsibility, although culturally they'll probably inhibit the sharing of responsibilities. What the real problem is, is the lack of accountability in a siloed organisation. Or quite the opposite; Too many persons are accountable for different/conflicting objectives.

In a siloed organisation, each silo is primarily responsible and even accountable for its own output, its immediate contribution to the process of product delivery, but not the full process of delivery itself. Meaning that when the outcome of the process is of the unwanted kind (caused an incident), either one of the silos’ outputs caused the problem (who is to blame?) and when there is no single silo to be blamed, nobody can be held accountable.

This can go as far that a sales team is responsible for selling a product. A signed contract is considered a success. The development team is responsible for changing the product. The release of the change into production is considered a success. The operations team is responsible for 'running' the product and it is considered to be successful when there are no incidents. This would be for SaaS vendors. For more traditional software companies, i.e. those that require implementations of a product at the customer site, the operations team is part of the customer's organisation or at least the operations accountability is typically with the customer. It'll be more complicated, because there will likely be an implementation team that is successful when the product is implemented according to the contract sold.
Success of the product is defined as selling/changing/operating/implementing. With different persons accountable for each of these successes, you see that conflicts are imminent. So when a problem happens anywhere in the delivery, each of the accountable persons will elaborate that they're not to blame, because they are successful. Actually, sales and development were successful, and operations and implementation were given something that prevented them from being successful. Contract was signed based on the availability of missing features at the time the implementation project reached completion. Future releases are feature complete and functionality fully tested. But it's unmanageable, not performant and definitely not secure. And can't be implemented as integrations with other systems not available until project end.

Siloed organisations are structured around tasks, competencies and expertise. By centralising capabilities, they can be shared across products. Siloed organisations are build on shared service centres. Reason behind these structures is cost reduction through utilisation optimisation. I'm not a fan, see: Perish or Survive, or being Efficient vs being Effective.

Output vs Outcome

It’s the difference between output and outcome that often drives ‘blaming’ in a post-mortem.

In siloed organisations, each silo’s focus is on output, its output. The silos are in many cases the result of centralising the responsibility for specific aspects of the delivery process, with a lack of accountability for the full process. Specialists are responsible for doing their 'thing' is efficient as possible.

Often responsibility is mistaken for accountability, so these task-optimised teams, teams of experts, are held accountable for what they deliver, which is not the outcome of the process, but output of their effort. Because they are the experts, they perform their task for different products, i.e. they participate in various delivery processes. And are held accountable for the number of tasks that they completed in total.

The problem therefore is in that the silos are operating truly independent of each other. Each silo services several product delivery processes. Because servicing only one (product delivery) process, would mean that a significant amount of time the silo would be idle. Since the silos are there to optimise resource utilisation, idle time is undesired. Idle time is considered wasted time by many non-Lean'ers. In Lean wasted time is time spend on something that is not immediately needed for the delivery of a product.

Every silo will work hard to meet its numbers. Meet its targets. And when the target is the number of tasks performed instead of the number of products delivered, we're doing a lot and contributing nothing.

Within a context of blameless post-mortems. In a context where blaming should be prevented, we need to make sure that responsibility is shared on the (product delivery) process outcome. Accountability is set to manage that outcome. Meaning that development and operational responsibilities are both defined to contribute to the outcome of the process. Something devops shines at.



Thanks once again for reading my blog. Please don't be reluctant to Tweet about it, put a link on Facebook or recommend this blog to your network on LinkedIn. Heck, send the link of my blog to all your Whatsapp friends and everybody in your contact-list. But if you really want to show your appreciation, drop a comment with your opinion on the topic, your experiences or anything else that is relevant.

Arc-E-Tect


The text very explicitly communicates my own personal views, experiences and practices. Any similarities with the views, experiences and practices of any of my previous or current clients, customers or employers are strictly coincidental. This post is therefore my own, and I am the sole author of it and am the sole copyright holder of it.

March 21, 2019

The question that takes away all blame


Blameless postmortems, or blameless RCA’s are supposed to be the new-normal in devops organisations, but all too often we see that first the team and sometimes the person to blame is sought, and then we tell them to ‘fix it’.
It’s for a large part a remnant of the siloed organisation and the culture that stems from it. And it is a matter of asking the wrong questions. This is, I would say, about 80% of the reason why we are unable to prevent incidents from recurring.

The cool part is in the ‘asking the wrong questions’ thing.
Now, before going into it, let me emphasise that postmortems, RCA’s are about preventing incidents to occur again. You want to know the root cause, because you want to prevent something like the incident to ever happen again. Stop reading if you disagree.

‘What?’ is wrong

All too often, when we are dealing with the aftermath of an incident, we wonder ‘what caused this incident?’. Which is a valid question, but not one that is very valuable. The issue I have with this approach is that when we have an answer to this question, we think we found the root cause of the incident. Which we haven’t.

The answer to the question ‘What caused the incident?’ is typically something technical. There was not enough memory in the server. There was a bandwidth problem. There was a bug in the software that resulted in the incident.

There is something very satisfying in the answer to the question ‘What caused the incident?’, it is in the gratification you get from knowing where the weakness in the product is. It’s in the software, in the hardware configuration, in the network infrastructure. Because when we know where the weakness was, we know who is responsible for the weak component. It’s the development team, automation team, the network team. And when we know who is responsible for that component, we know who to blame and who to tell to go fix it.

In about all cases I’ve been involved in RCA’s, putting the blame on somebody was not about punishing that person, it was about identifying who should fix the problem.

The problem is in ‘responsibility’, because the person being held responsible is not necessarily the person that is accountable for the incident. Often, especially in a siloed organisation there is no one accountable.

Although it is important to understand what went wrong, and what caused the impact, we need to realise that this is not the same as understanding what caused the incident. We’re not at the root-cause just yet. But we want to, because this investigation it painful. Colleagues are to blame, and the responsible persons must be called to justice. They must be told that we can never ever feel that impact again. And so, we make sure that next time the impact will be bearable. We increase memory in the server, increase the bandwidth in our network, fix the bug in our software. All holes are plugged. Ready to go.

If only we had addressed the root cause, it all would be honky dory.

‘Why?’ is right

The question that should be asked is not so much about what caused the incident, it’s about why the incident could occur in the first place.

That’s a tough question to answer. Why was there a bug in the software? Why was there not enough bandwidth? Why was there not enough memory in the server?

And that’s only the first ‘Y’.

A very common, tried-and-tested, effective way of identifying the ‘real’ root cause of an incident is by applying the 5-Y method. In this approach you ask 5 times ‘Why could the previous answer happen?’. Experience has taught us that going 5 levels deep will get you to the root cause of the problem, sometimes less, hardly ever more than five levels deep.

Let’s assume that the incident was due to insufficient bandwidth and let’s start asking ‘Why?’

  1. Why was there not enough bandwidth? Too many customers accessed the newly released API.
  2. Why did too many customers access the new API? Because we announced it prematurely in our global newsletter.
  3. Why was it announced in our global newsletter? Because the marketing manager wasn’t aware that the API was to be released following the ‘soft-launch protocol’.
  4. Why was the marketing manager not aware of the fact that the release was to follow the soft-launch protocol? Because she was not in the meeting in which it was decided to follow the soft-launch protocol.
  5. Why wasn’t she in the meeting in which it was decided that the API was going to follow the soft-launch protocol? Because she was on vacation and didn’t appoint a delegate.

Now we know why the incident could occur. Not what caused it, but why it could be caused. Making sure that the marketing manager or a delegate is attending meetings in which product launch strategies are decided will prevent this incident to occur in the future.

Of course, the above is only an example, but it shows that by asking ‘What?’ the solution would be a costly technical solution and by asking ‘Why?’ the solution is better meeting attendance.

Another important conclusion you might’ve drawn is that asking ‘What?’ only involves technical people. Further leading the path to a solution down the costly technical path. Whereas the ‘Why?’ question requires all parties involved in the delivery of the product (the API) to attend the postmortem. Getting to the bottom of the incident’s cause, requires a multi-disciplinary team. Just like delivering a product requires many disciplines.

It makes no sense to think that creating a success is requiring many disciplines, but when it results in a failure, to prevent it, only requires a single discipline. There is no difference between delivering something that works and something that doesn’t. Not from a product delivery perspective.

Product Owner or Problem Owner?

The question is of course: Who would go through all this trouble and assemble all these people that are involved in delivering a product into the hands of our customers? It’s the one that is held accountable for the incident. More importantly, it’s the person that is held accountable for the fact that the incident doesn’t occur again.

The Product Owner would be my preferred role, ownership of the product implies ownership of the success of the product and all challenges that come with it.

There is a follow-up story that you can find here.


Thanks once again for reading my blog. Please don't be reluctant to Tweet about it, put a link on Facebook or recommend this blog to your network on LinkedIn. Heck, send the link of my blog to all your Whatsapp friends and everybody in your contact-list. But if you really want to show your appreciation, drop a comment with your opinion on the topic, your experiences or anything else that is relevant.

Arc-E-Tect


The text very explicitly communicates my own personal views, experiences and practices. Any similarities with the views, experiences and practices of any of my previous or current clients, customers or employers are strictly coincidental. This post is therefore my own, and I am the sole author of it and am the sole copyright holder of it.

January 2, 2019

Cloud Native Enterprises - Rapid elasticity

Which don't have a lot to do with Cloud Native Apps but everything with truly embracing the paradigm shifts the Cloud has brought IT within the realm of businesses.

Read the Introduction first.

After reading the introduction to these posts you know what a cloud infrastructure is and what cloud native applications are, what about cloud native enterprises. Well these are enterprises that adhere to these same 5 characteristics. These enterprises, or organisations in general, cannot be modelled according to traditional enterprise models because of their specific market, competition, growth-stage, etc. These enterprises need to be, for all accounts, be cloud native in order to grow, succeed and be sustainable. Interestingly, but not surprisingly they require The Cloud and Cloud Native Applications.
In coming posts I will address every essential characteristic of The Cloud as defined by NIST from a perspective of the Enterprise. Unlike most cases, I will post these within the next 7 days and I certainly do hope before coming weekend.
  • On-demand self-service. When online services and core systems really seamlessly integrate.
  • Broad network access. When customers, partners and users are distinct groups treated equal.
  • Resource pooling. When synergy across value chains makes the difference.
  • Measured services. When business resources are limited.
Rapid elasticity. When business is extremely unpredictable.

Very few organisations can say that year round they experience the same amount of business. Almost every organisation experiences something like a 'season'. There are always highs and lows in an organisation's load. This can be 'Black Friday/Cyber Monday' for retailers, tax-season for the IRS or its equivalent in your country, or for example hotels during the summer holidays.
But not only in sales, there are highs and lows, often very predictable, but in production companies there are peaks in order fulfilment.

In order to deal with these fluctuations in 'business load', organisations need to be elastic. The more elastic an organisation is, the better it will be able to handle the fluctuations. Mind, this is not the same as being agile. Being agile is being able to change direction as needed, in a timely manner. Being elastic is being able to change scale as needed. Also in a timely manner. Arguably, being elastic is more difficult to accomplish than being agile.

Elasticity is a matter of scaling up, and down. The more elastic an organisation can be, the more efficient it can do its business. Efficiency here is a matter of spending just enough. Another difference between elasticity and agility. The former being about efficient in using resources, the latter is about being effective in resource utilisation.

From the Cloud we know that elastic means scaling up, or down of the IT infrastructure in order to manage fluctuating load. Thus not having to worry about under-utilisation of IT infrastructure, thus paying for infrastructure that is not being used. And at the same time, not having to worry when load increases and the stability of the environment isn't compromised due to over-utilisation of that same infrastructure.

In the Cloud Native Enterprise, we project these traits onto the organisation itself. Onto the business.

Traditionally, scaling organisations are hard to accomplish. Extremely hard in fact. For one there is always the challenge of resource availability. Where resources are of course human resources. The most common way of addressing this is through contractors. The hiring and firing processes around contractors are far more flexible than for permanent employees. Thus bring some elasticity to the organisation. But getting the 'bodies in' is only part of the challenge. Getting the 'right bodies in' is another aspect. Adding personnel with the right skills is possibly even harder. First you need to find them, next you need to validate that you found them. The hiring process is cumbersome and the further you need to stretch the organisational elastic band, the harder it becomes. The amount of effort needed is not increasing linear, but almost exponential.
Organisations need to be able to scale up quickly. Which often also means that the HR department needs to be elastic as well. Again, the same challenge apply.
In an economy with an up-beat, there are more companies that are looking for the same scalability and resources become even more of a challenge to find.
I haven't mentioned the strain on the organisation itself due to growing too large. More management layers need to be introduced in order to be able to grow and manage this. And this is where scaling down becomes a problem. Adding more management layers in the hierarchy to handle the size of the workforce, is relatively easy compared to removing these layers again.

The more seasonal a market is, the more of a challenge elasticity becomes for an organisation. Tax-offices will be overloaded with calls to the help-desk as the deadline for submitting your tax-forms nears.

At one point I was involved in a court-case as one of the expert witnesses where my client succumbed at its own success. The client was in a business where 98% (!!!) of its revenue was generated in 5 consecutive days of the year. It couldn't handle the increased business as it couldn't predict what that load would be. Especially since in the months prior to their 'peak' they took over their largest competitor.

Organisations can't scale their workforce to the point where we can talk of elasticity. Scaling up and down is a matter of weeks. Where elasticity is the ability to scale up and down in a matter of days, hours or even minutes and less.
This is where the Cloud comes in. It is not so much the elasticity of Cloud infrastructure that plays a role, but the ability of an organisation to utilise the Cloud such that it can scale its business efficiently. Cloud infrastructure is an important part, where it comes to handle IT load. But the true business elasticity comes from handling at a business level the fluctuating load.
Extreme automation of business processes allows an organisation to reduce its dependence on specific business knowledge of its employees. Everybody can press a button or enter data when the 'system' does the heavy lifting in applying business rules and execute repetitive tasks. Understand that by following this paradigm, the reliance of the organisation on less-educated employees is reduced to a bare minimum, and it can focus its efforts towards attracting higher educated employees. Employees that require less managerial guidance which allows for a more flattened hierarchy. Flat hierarchies are more elastic.
Automated processes can in addition, benefit from the technological elasticity of the Cloud. Thus having a double edged sword.

Elastic organisations are not trivial. Far from it in fact. Main reason being that traditionally, organisations scale with their business by either increasing its workforce through short term contracts, i.e. contractors, which can be hired and fired as needed. Or by attracting permanent employees and assume a positive attitude towards the future. Or laying off personnel in advance when the future is perceived less positive. It is how organisations are used to manage the seasons.

The Cloud Native Enterprise on the other hand is far from traditional. At its core it will embrace technology resources not to replace human resources, but to complement them. It invests in senior, higher educated experts, to reduce the reliance on management. In effect, pay more to less, in order to reduce the need to scale the workforce.

Concluding

And so the Cloud Native Enterprise is the enterprise where IT is part of the business when it comes to delivering business products. Allowing a business to grow and shrink as needed, just ahead of time and with the flexibility of a rubber band.


Thanks once again for reading my blog. Please don't be reluctant to Tweet about it, put a link on Facebook or recommend this blog to your network on LinkedIn. Heck, send the link of my blog to all your Whatsapp friends and everybody in your contact-list. But if you really want to show your appreciation, drop a comment with your opinion on the topic, your experiences or anything else that is relevant.

Arc-E-Tect


The text very explicitly communicates my own personal views, experiences and practices. Any similarities with the views, experiences and practices of any of my previous or current clients, customers or employers are strictly coincidental. This post is therefore my own, and I am the sole author of it and am the sole copyright holder of it.

December 30, 2018

Digital Transformation to Survive in '19

This is a cross-post from an article I published on LinkedIn.
Today I came across a quote and realised the enormity of it. The quote inspired me to write this article. It is an excerpt of some preliminary text from the book I started writing in late 2017 and am still working on while typing this article. A book for which I am still trying to find an appropriate working title.

“If the rate of change on the outside exceeds the rate of change on the inside, the end is near.” - Jack Welch


The enormity of this quote is in the fact that unless an organisation can adapt to a changing world, one that is changing at an ever increasing rate, the organisation is doomed to seize to existing. Having worked with and still working in organisations that are realising that change is needed in order to survive, especially in the long run. I often encounter leadership that is acting more like a 'management-team' than a 'group of leaders' welcoming change and promote that change. A core cause of problems when implementing the required changes.

Finances to drive change

These organisations financial outlook typically show the necessity of change, but all too often those same finances insufficiently show the required pace in order for that organisation to survive.

The change I am referring to is often initiated under the name 'Digital Transformation'. A transformation many organisations find themselves in. Unfortunately, it is often met by the organisation's cynicism as it is perceived as yet another restructuring project. And in more cases than not, rightfully so. Any organisation that is looking for a sustainable means of staying competitive, will need to be able to continuously reinvent itself. It will need to adopt a practice of continuous change. As it is the only way to become more capable than the competition and excel on a never levelled playing field.

Understanding this need to adopt a mindset where 'to last' is replaced by 'for change' in the organisation's DNA, is the key to find the right sense of urgency. A sense of urgency that is a pre-requisite for a successful digital transformation.

Not just your average Agile Transformation

The 'Digital Transformation' is more than just an 'Agile Transformation', which typically takes place in the IT department. The agile transformation is limiting itself to a different way of working. Ranging from adopting an iterative approach in software development to empowering the team that develops the software also support the application in a production environment; DevOps. Embracing the build-measure-learn paradigm, these teams are not so much 'more efficient', instead they are 'more effective'. Increased effectiveness results in a reduced risk of 'bad investments'. It means being better able to act 'just-in-time'. This also increases the opportunity to be more efficient. And it is efficiency that leads to more productivity. 'More for the same' instead of 'the same for less'.

A digital transformation is one that impacts the full organisation, not just the IT department. It allows the organisation to leverage (new) technology to be more effective. Applying technology to increase effectiveness of an organisation. Improving the organisation's sustainability is what defines the difference between future success and future failure. The reason for this is that increased effectiveness allows an organisation to change at a higher rate 'on the inside' than it experiences change 'on the outside'. A prerequisite of surviving in the long run in a competitive market.

The full story

Shorter product development cycles, a typical objective of agile transformations, are only a tiny part of the full story. An essential part, but tiny in comparison to what else is needed. A digital transformation encompasses the restructuring of an organisation. With the intention to establish change (over time) organically. Change that allows the organisation to grow, or shrink, as required. One that as a trait, constantly reduces the 'chain-of-command' to its bare minimum. One that embeds 'build-measure-learn' in its core processes and where its leadership relentlessly applies fact-based decision making. Where measures of success are defined at the start of every initiative and that continuously searches for means to reduce the risk of 'not knowing'.

Although the transformation is mainly one that impacts the organisation and its processes. It is a transformation that almost religiously relies on the (digital) technology available in the market. Cloud, AI, Data Analytics etc allow us to work with all relevant information (facts) at our fingertips. But the introduction of these technologies are not what makes the digital transformation such a challenge. Harnessing them is. And since we are at the start of the evolution of these technologies, being able to change with them. As the new techhnologies become ever more sophisticated and powerful, they will set the boys apart from the men. But mind you, it will most likely be the boys that will come out on top.

Concluding

The digital transformation, therefore, is one that transforms an organisation from a mindset of 'to-last' to a mindset of 'for-change'. From one that finds its future in a foundation of bedrock, to one that finds it future in being able to surf the waves. Traditional stability is focused on not having to change in a changing world, whereas adaptability is focused on being able to change with a changing world.



Thanks once again for reading my blog. Please don't be reluctant to Tweet about it, put a link on Facebook or recommend this blog to your network on LinkedIn. Heck, send the link of my blog to all your Whatsapp friends and everybody in your contact-list. But if you really want to show your appreciation, drop a comment with your opinion on the topic, your experiences or anything else that is relevant.

Arc-E-Tect


This article is reflecting my own, personal, opinion and does not reflect in any way the views, ideas or opinions of any organisation I work or worked for. It is based on my own personal experiences and research I conducted by myself.

September 17, 2018

How Charming is Laziness?

Long time ago, while I was still a student at the PolyTechnic in Enschede, The Netherlands, I would justify my choice to study computer science, by proclaiming that laziness leads to efficiency. Which of course makes no sense I now know, because it leads to effectiveness. Penny-pinching leads to efficiency. You can read all about it in this post Perish or Survive, or being Efficient vs being Effective. But there's this little concept that results in more efficiency and laziness as well. It's therefore a charm. It is called automation and closely related to that third time.

Third Time Is Automated Principle

One of the key principles at Amazon AWS is that everything must be automated. It's not just that everything should be automatable, but it should be automated.
Whether or not it's an urban legend, but word has it that when you create something that requires manual action, you're out looking for a new job. For a company like Amazon AWS it is clear why automation is such a huge thing. And it is clear as well why they have such focus on API's. API's are how automation across the board is facilitated. But most companies are not Amazon or any of the other cloud providers. Most of you my dear readers are more likely to run your systems on Amazon AWS, Google Cloud or Microsoft Azure, then run those applications for your customers. Probably, your IT landscape is in size not even close to Amazon. Probably comparing to Amazon AWS as the Netherlands compares to the rest of the EU. In most if not all of the dimensions you can think of.
Arguably, the principle of "Automate Everything" doesn't apply to you. I'll leave it up to you to think of one or more arguments why automation is not something you should hold dearly.

Challenge to you: put in the comments a good reason why you think automation is not necessarily needed. I'll make an effort to counter your argumentation as a reply to your comment.
But read on first.

The benefits of automating processes are many. Irrespective of the kind of processes. An automated process is infinitely more likely to be repeatable than any manual process. This results in higher quality since errors will either be made consistently, and can be fixed, or will consistently not occur at all. How compelling is that?
Although the automated and manual version of a process might take the same amount of time to be executed, the automated process allows a person to work concurrently on something else that cannot be automated. And automated processes do not rely on the availability of a specialist to execute the process. So automation makes you and your organisation more scalable.

Not automating processes, even IT processes doesn't make sense. Still, when it comes to IT, we hardly do this. Why?

The situation at Amazon


I've come across many situations where things weren't automated. Worse yet, they could be automated. They were not automatable.

For me this was always an interesting fact to find out. For one, we're in IT and IT is all about automation. In fact, in Dutch we refer to IT as Automation (Automatisering - Dutch). The paradox is that we apply IT all over the enterprise to automate business processes, but when it comes to the IT processes themselves, automation is very likely the last thing on our minds. And when you think of it, that doesn't make sense at all. Walk into a room full of IT people, and just pose the statement that it is hard to understand why we're so good at automating business processes, yet we don't have automation in our own processes. And you'll see at least 90% of in-agreement-nodding heads, the remaining 10% are too flabbergasted with the realisation that this is a true statement. Same statement in a room full of non IT people and the first thing you find yourself doing is explaining why this is.

The fact that IT people are not automating their own processes is tough to explain, and I for one do not have such an explanation. There is an explanation though, for why Amazon AWS's processes are all automated. It's because one of the core principles by which they do IT is "Automate Everything". At a dinner party with Werner Vogels (Amazon's CTO) I was invited to, being the Chief Architect of a FinTech startup, I asked Werner (all attendees were told that we were on first-name basis), how it was possible for a huge company like Amazon AWS to live by these rigorous principles. Everything is an API, Automate Everything, and a few more. His reply was that there were two main reasons why it works. 
  1. Senior management all the way up to Jeff Bezos, were openly behind these principles. In fact many of them were mandated by Jeff Bezos himself. 
  2. Everybody in the organisation experienced for themselves the validity of these principles.
I asked him to clarify that second reason. According to Werner Vogels, Amazon is dedicating a lot if not all of its time to make its processes impacting customer experience as efficient as possible without impacting customer satisfaction. 'A satisfied customer is a returning customer'. The effects of changing a process is made visible to the whole company, at all times. So compliance to the principles will result in changes to the processes and the effects would be visible. Therefore, the validation of the principles would be continuous. And according to Werner, nothing is as motivating to change your processes than to see the effects of your efforts.

Rest of the dinner I was milling this over, enjoying the food and talk to the other guests.

The situation in the 'real world'

Unfortunately, most of us work at real enterprises. And those two reasons Werner Vogels had given me why IT process automation worked for Amazon aren't obvious in the situations I have found myself in.
For example, the amount of automation in a process within IT is not a metric that anybody is held accountable for anywhere I've worked or consulted. Neither are the benefits of automation part of somebody's accountability.
IT process cost reduction (IPCR) is as far as I know not a KPI within enterprises. Nor is the time-to-market (TTM). The latter often does find itself in another incarnation on reports and dashboards, namely as MTTR, the Mean Time To Resolution. When it comes to MTTR, we do see them on reports, but the MTTR is in hardly any case part of somebody's accountability. Same goes for time-to-market. Although it is always mentioned for any improvement project, it is hardly measured or reported on. It's a project issue in most cases, meaning that an organisation is doing projects instead of delivering products.
The lack of real metrics and if you will KPI's means that from an accountability perspective, there is not a clear person that has an incentive to push automation. And if you read my blogs on a regular basis, you know that I'm big on accountability.
Visibility of the effects of changes to the IT processes is another challenge for organisations. We often find ourselves in organisations that don't have a culture to measure our IT processes. This is especially true for organisations that have been around for decades. Since IT process automation is not pushed, there is no incentive to find bottlenecks in processes or pinpoint areas that are up for improvement. Resulting in the situation where the effect of improvements are not visible in most cases.

The lack of accountability and the rather big hurdle to be taken by IT departments in order to automate result in a situation in which we are just not automating our processes, because it gets no priority on our backlogs or funding in our PRINCE2 budgets. Process automation is collateral. And it's not a matter of not being as large a company as Amazon. It's a matter of not being aware of how automation affects the bottomline. And of not being held accountable for impacting the bottomline.

Third Time Principle

Three times is a charm, or at least should be automated.

In organisations I worked previously, the principle of automation as in play at Amazon was a little bit more pragmatic. One that has worked well for me, is the principle of "Everything done a third time, will be automated", reasoning behind this is that if you do the same thing a third time, it is extremely likely you will do it a 4th, 5th and even more times. The time required to create the automation will be saved by executing the task over and over again.

Time invested is never gained, but can only lead to savings later on.
The Third Time Principle is a compromise, although one might argue that it is a Troyan horse. By adopting this principle, the argument that it creates too much overhead for mundane actions is off the table. Only those actions that are performed repeatedly are automated. Built in justification for the investment needed.

The Third Time Principle is easy to adopt. Since it is a compromise, it can be applied when the push to automation is bottom-up. We often see that engineers see the need for automation since that is where automation will solve a problem and effects are noticed. To develop the automation is often a matter of getting the time to do so. Automation is now competing with other requirements for development time, for priority on the backlog. We all know where the priorities will be. Applying The Third Time Principle will remove this obstacle. Engineers can justify the need for automation and the Product Owner can justify the priority of the automation stories on the backlog.

Continuous Delivery

When striving for Continuous Delivery (CD) and more so for Continuous Deployment (also CD), there is no other way than to automate everything. CD requires a rigorous regime to move every manual task to the left in a process defined western style (i.e. left to right notation). Product development follows the DTAP model. Development is followed by testing is followed by accepting is followed by production → DTAP. In CD we hold true to the paradigm that everything manual is done in D, and TAP are fully automated. Any manual action in T, A and P needs to be automated and the scripts for this are developed in D, because otherwise it can't be done.
So when striving for Continuous Delivery, everything in the product development cycle is to be automated.

Accountability

When process automation metrics are part of somebody's accountability, that person will also be mandated to automate as much as possible. This is for the simple reason that you can't hold somebody accountable for something they cannot influence. Therefore, when process automation is measured through some metric for which somebody is accountable, it is that person's prerogative to drop The Third Time Principle and instead adopt the Automate Everything Principle.

Accountability is a matter of top-down enforcement. It is also a key aspect of culture change. When we want to adopt a culture in which we allow ourselves to reap all benefits of automation, especially our IT processes, we can't get around the fact that we need to revisit the metrics by which we hold ourselves accountable.

Scalability


One aspect that I haven't really touched upon is scalability, although I did mention it briefly. Automation is a key aspect of scalability, not so much scalability on a technical level but organisational scalability. Manual actions always need a person to execute them. The more complex the activities become, the more experienced or knowledgeable the person needs to be. And before you know it, it requires a specific person within the organisation. Because you rely on this person, that person will become the person with all required knowledge, it's a vicious circle. Think about it for a second, and I'm sure that you can think of a process or an activity in a process where you rely on a specific person and you know that person by name. In fact, if you want the activity to be done, you will even call that person because she's among the busiest persons in the company.

When you want to keep activities simple, you need to automate them. The more you automate, the easier it is to keep the activity simple. Hardly anything is more simpler than to have a command like 'do_complex_activity.py' provided that the automation is done using Python. Anybody can run this command provided they have the rights to do so, and when done really properly, anybody can do it at any time, because all the complexity of who can do what when is taken care of by the automation code as well.

Concluding

In the end, you'll agree that automation is what we need to apply to all our processes. At least when it's the third time you're doing the same job. It will improve quality and consistency. It also means that we as a company can scale our organisation without the need to increase headcount. It will be important to monitor the benefits of automation, without metrics it will remain a matter of good faith on the short term, and it will not justify the investments on a longer term. By understanding the benefits of automation and getting insights on where automation has most impact.


Thanks once again for reading my blog. Please don't be reluctant to Tweet about it, put a link on Facebook or recommend this blog to your network on LinkedIn. Heck, send the link of my blog to all your Whatsapp friends and everybody in your contact-list. But if you really want to show your appreciation, drop a comment with your opinion on the topic, your experiences or anything else that is relevant.

Arc-E-Tect


The text very explicitly communicates my own personal views, experiences and practices. Any similarities with the views, experiences and practices of any of my previous or current clients, customers or employers are strictly coincidental. This post is therefore my own, and I am the sole author of it and am the sole copyright holder of it.

September 2, 2018

Cloud Native Enterprises - Broad network access

Which don't have a lot to do with Cloud Native Apps but everything with truely embracing the paradigm shifts the Cloud has brought IT within the realm of businesses.

Read the Introduction first.

After reading the introduction to these posts you know what a cloud infrastructure is and what cloud native applications are, what about cloud native enterprises. Well these are enterprises that adhere to these same 5 characteristics. These enterprises, or organisations in general, cannot be modelled according to traditional enterprise models because of their specific market, competition, growth-stage, etc. These enterprises need to be, for all accounts, be cloud native in order to grow, succeed and be sustainable. Interestingly, but not surprisingly they require The Cloud and Cloud Native Applications.
In coming posts I will address every essential characteristic of The Cloud as defined by NIST from a perspective of the Enterprise. Unlike most cases, I will post these within the next 7 days and I certainly do hope before coming weekend.
  • On-demand self-service. When online services and core systems really seamlessly integrate.
  • Resource pooling. When synergy across value chains makes the difference.
  • Rapid elasticity. When business is extremely unpredictable.
  • Measured services. When business resources are limited.
Broad network access. When your business hours are truly 24x7.

This is an interesting aspect of the Cloud Native Enterprise. Because many organisations are already 24x7 businesses. Especially in the online world, being always on is a requirement to stay in business. 

Network here is not referring to the the computer network on a technical level as we know it. According to NIST, one of the characteristics of the cloud is broad network access, which from the definition, this means that the cloud is always accessible through the internet. The computer communications network.
Within the context of the Cloud Native Enterprise I am referring to the communications network of the business. On the one side, this encompasses all parties a business has dealings with. Think customers, users. partner, employees etc. I will get back to this later on in this post. But it also encompasses all means through which this communication takes place. Think devices and associated channels.

So, when we look at the business that is truly cloud native, when we talk about broad network access, we mean that the business is accessible by its complete business network, via a variety of devices and channels.

The Cloud Native Enterprise exposes its services, all of its services, in the same way to customers as it will to users and partners, as well as employees. Every member of every group enjoys the same level of service and the same constraints. Distinctions are made using the concept of role-based access to services. Still all services are always accessible to all members of the network through the same channels. In addition, it will provide the same level of service through all channels on all devices. Ideally. Of course contextual limitations are to be taken into account.

Traditional Focus - Cost vs Value

Consider the more traditional enterprise, where customers are assigned a specific account manager. The account manager maintains the relationship with the customer. She has specific KPI's that need to be met and she is more or less free in determining how to achieve this. Users of the customer are not aware of the account manager, instead they interface through a help-desk with the enterprise which will consists of self-service functionality for the more mundane support, automated systems like chatbots and more complex interaction through a help-desk agent.
Partners of the enterprise on the other hand, will work with peers within the enterprise. Informal contacts are more prevalent and accepted. Where customers and users will need to follow the formal processes in order for the enterprise to be more efficient, partners will follow the informal communication lines in order to be more effective. We see in traditional enterprises that with customers and users, processes are cost driven. Partner oriented processes are value driven.

Cloud Native Focus - Revenue

In the Cloud Native Enterprise, all interactions are focusing on revenue. Key aspect here is the homogeneous approach towards interaction with stakeholders. Customers, employees, users, partners are all considered stakeholder of the enterprise. Access to the enterprise is homogeneous across the full network, and processes are optimised for revenue. Sometimes resulting in focus on efficiency, other times on effectiveness.
What you will see is that business scalability, both temporal and geographical, is addressed explicitly. Every stakeholder in the enterprise's network can access the enterprise 24x7 and from any location.
For every organisation this is a challenge to manage. But for the Cloud Native Enterprise the mere premise of the Cloud as an IT facilitator, it becomes a matter of survival.

Consistent User Experience

The Cloud is great for realising products and services that scale with your business. Not talking about elasticity here, but about the ability to have a single approach towards addressing your stakeholders' needs and demands.
Leveraging the characteristics of the Cloud, including its technical aspects of broad network access, means that it is possible to provide all users of all services and products with a consistent user experience. This entails the same experience provided to customers, partners, employees etc. This consistent user experience across all services will allow for a seamless transition for a user from one role to another.

Concluding

As with its technical counterpart, (broad) network access is a key characteristic for the Cloud Native Enterprise, as it will provide a homogeneous approach towards accessibility of products and services across the full breadth of the business network

(Special thanks to my colleague Mandeep for pointing out some much needed clarifications in the original post)

Thanks once again for reading my blog. Please don't be reluctant to Tweet about it, put a link on Facebook or recommend this blog to your network on LinkedIn. Heck, send the link of my blog to all your Whatsapp friends and everybody in your contact-list. But if you really want to show your appreciation, drop a comment with your opinion on the topic, your experiences or anything else that is relevant.

Arc-E-Tect


The text very explicitly communicates my own personal views, experiences and practices. Any similarities with the views, experiences and practices of any of my previous or current clients, customers or employers are strictly coincidental. This post is therefore my own, and I am the sole author of it and am the sole copyright holder of it.