Ticker

6/recent/ticker-posts

Ad Code

Responsive Advertisement

Here’s What Goes Catastrophically Wrong When You Don’t Follow Your IT Process

IT operations leader presenting a miniature scene of cascading server failure

An IT process is a repeatable set of steps, owners, controls, and evidence used to run and protect technology systems consistently. When people skip those steps often enough, the shortcut can start to feel normal even as risk accumulates underneath it.

That is how an ordinary deployment, access review, backup handoff, or capacity test goes catastrophically wrong and becomes a security breach, data loss event, or outage. The pattern has a name: normalization of deviance. Understanding it is the first step toward building IT processes people can follow under real operating pressure.

A Microsoft Azure Storage incident shows how quickly one process break can spread. In its final root cause analysis of the November 18, 2014 interruption, Microsoft explained that a configuration change departed from its standard staged deployment and validation process. The change triggered an infinite loop in Storage front-end roles and affected customers across multiple regions. The lesson is not that change is dangerous. It is that the deployment process, including staged validation and rollback readiness, is part of the system.

The same discipline applies to everyday configuration management. A technically valid change can still create an operational failure when it bypasses the controls that limit its blast radius.

I learned that distinction during my first real IT job at 16. I had just passed my CCNA with help from my dad and was hungry to get my hands on real technology. I was apparently the youngest person in Australia to hold it at the time. Configuring a live Windows server was a totally different experience from tinkering with my father’s machines at home. There was rigid process everywhere because the consequences were real: one unchecked assumption could affect customers, colleagues, and every system that depended on that server.

Engineer Dan Luu captured the recurring postmortem pattern plainly: someone decides a staging or testing step is inefficient, skips it because nothing is expected to go wrong, and then discovers why the step existed. High-profile breaches and outages are not always caused by an unknown technical weakness. Many begin with a known control that people stopped treating as mandatory.

Habitual violation of process leads to mistakes in all kinds of IT arenas. High-profile hacks, cloud computing outages, and downtime on major websites can happen because people do not follow the processes put in place to prevent them. They may not even care until the shortcut becomes a catastrophe. That makes computing security an organizational concern as well as a technical one.

Why Deviance Becomes the Norm

Approved IT process with a controlled deviation and owner review loop

Sociologist Diane Vaughan developed the concept of normalization of deviance while studying organizational failure. It describes the gradual acceptance of behavior that falls outside a defined standard. When a team breaks a rule and nothing bad happens, the absence of an immediate consequence can be misread as proof that the rule was unnecessary.

The shortcut then becomes a local norm. A new person notices the gap and reacts with alarm. Experienced colleagues acknowledge the problem, explain that it has always worked, and return to business as usual. After enough exposure, the new person learns to tolerate it too. When the next colleague arrives, yesterday’s critic is now the person explaining why the exception is acceptable.

In Dan Luu’s version of the common scenario, a new person joins, discovers bad IT processes, and asks how everyone can accept them. The old hands say they know and are concerned about it. Soon the new hire becomes accustomed to the same behavior and gives that answer to the next person. The insidious part is that people can carry the normalized practice elsewhere for the duration of their career.

That is what makes normalization of deviance so dangerous in IT. Systems often contain redundancy, retries, access boundaries, and experienced people who quietly compensate for weak process. Those protections can hide the cost of a shortcut for months or years. The team sees successful outcomes, while the margin for error keeps shrinking.

A healthy process does not ban every exception. It makes exceptions visible. The owner records why the normal path cannot be followed, identifies the added risk, approves a compensating control, and sets a time to return to the standard. That review loop keeps an urgent workaround from becoming permanent infrastructure.

Leaders also have to examine the process itself. If skilled people bypass the same step repeatedly, the control may be unclear, slow, or disconnected from the way the system now operates. The response should be to investigate the friction and redesign the control without losing its protective purpose. A usable standard gives people a safe path for ordinary work and a governed path for genuine exceptions.

1. Hackers Steal Your Customers’ Data

Security controls for encryption credential rotation multifactor authentication and access review

In 2013, Buffer was compromised and attackers used connected social accounts to publish spam. Buffer responded publicly, revoked tokens, restored service, and documented security improvements. The incident is a useful reminder that customer trust depends on more than the application interface. It depends on how credentials and access paths are protected behind the scenes.

For a social sharing tool like Buffer, trust is crucial. Customers trust the app with login information and access tokens so it can post to connected accounts. When Buffer was hacked, customers experienced a flagrant violation of that trust as compromised accounts posted spam. Buffer was extremely open about the security breach, sending regular updates about the app’s functionality and explaining what had happened.

A security process fails when controls such as encryption, credential rotation, multifactor authentication, least privilege, and access review exist as recommendations rather than operating requirements. Each individual shortcut may look harmless. Together they determine what an attacker can reach and how useful stolen data will be.

Important measures implemented after the incident included added encryption for OAuth access tokens and changes to API calls. The broader lesson is basic common sense, but common sense still needs accountability. Security can be de-prioritized because it is hard to build and is not immediately urgent. Teams table security to tackle pressing issues, get accustomed to sweeping deviant behavior aside, and eventually stop recognizing it as deviant.

Deviant behavior: Treating security controls as optional

Security work competes with launches, incidents, customer requests, and technical debt. Because the benefit of a preventive control is an event that does not happen, it is easy to delay the work. A missed review becomes two missed reviews. A temporary broad permission remains in place. A service account is never rotated because it has not caused a visible problem.

The answer is not to tell people to care more. Design the control into the work. Require an owner and due date, automate evidence collection where possible, and make exceptions expire. The process should show which systems are in scope, what evidence proves completion, and who must review a failure.

Incident reviews should trace both the technical path and the operating path. Ask which decision was made, what information the operator had, why the control did not stop the action, and whether similar exposure exists elsewhere. That produces improvements the next person can use, instead of a warning that depends on everyone remembering the last incident.

What you can do

  • Require phishing-resistant multifactor authentication for privileged and remote access, using CISA guidance as a baseline.
  • Encrypt sensitive data in transit and at rest, and manage keys separately from the protected data.
  • Rotate secrets and revoke unused credentials on a defined schedule and after role changes.
  • Run periodic access reviews that name the reviewer, retain evidence, and track remediation.
  • Use a repeatable network security management process so control checks are assigned and visible.

Buffer’s incident communication also offers a process lesson. Clear ownership and frequent updates help a company coordinate recovery while customers decide what action to take.

2. You Literally Lose Your Data

Encrypted backup chain of custody with offline storage and a verified restore test

In 2010, the UK Financial Services Authority fined Zurich Insurance £2.275 million after the loss of an unencrypted backup tape containing personal information for about 46,000 policyholders. The regulator found that Zurich UK had failed to oversee the transfer adequately and did not identify the loss for roughly a year.

The regulator’s concern was not only the lost data tape. Zurich did not have adequate controls in place to prevent the data being used for financial crime. The company certainly knew those controls mattered, but the routine transfer had fallen off the radar. Everyone became accustomed to the weak handoff until the big blunder made the risk visible.

The tape was lost during a routine transfer, which is exactly why the case matters. Routine work is where normalization thrives. A familiar handoff feels low risk, so teams stop verifying custody, encryption, inventory, or receipt. The process appears to function until an asset cannot be located and nobody can prove where control was lost.

Deviant behavior: Backing up without proving recoverability

A backup is only one component of recovery. Teams also need protected copies, defined retention, custody records, isolated or offline storage, monitored job results, and restore tests. Without those controls, a dashboard can show successful backups while the organization remains unable to recover after ransomware, deletion, corruption, or physical loss.

Endpoint-only copies create the same false confidence. If the device is stolen, encrypted by ransomware, or damaged, the copy may disappear with the original. Critical data needs a recovery design that accounts for independent failure domains and the time the business can tolerate without the system.

Recovery requirements should be explicit. The recovery point objective defines how much recent data the business can lose, while the recovery time objective defines how long the service can remain unavailable. Those targets determine backup frequency, storage architecture, test scenarios, and escalation. Without them, a team cannot tell whether its backup process matches business need.

What you can do

  • Follow the CISA ransomware guide by maintaining offline or otherwise isolated backups and protecting backup systems from ordinary administrator compromise.
  • Encrypt removable media and backup data, with controlled key access and documented custody.
  • Record each handoff, confirm receipt, and investigate missing evidence immediately.
  • Run scheduled restore tests that verify data integrity, application dependencies, and recovery time.
  • Use a client data backup process to assign checks, evidence, exceptions, and follow-up actions.

The goal is not to perform backup activity. It is to produce a recoverable service. Restore testing turns that outcome into evidence and exposes gaps while there is still time to fix them.

Historical tape processes included details that could feel irrelevant, such as verifying the right equipment, labeling tapes, confirming the backup type, and setting recurring reminders to perform a differential backup. Modern storage changes the mechanism, not the principle. Media identity, retention, custody, isolation, monitoring, and restore evidence still have to be checked.

3. Your Server Crashes When You Need it Most

Load test dashboard with capacity threshold failover status and monitoring checks

In 2011, demand for Target’s limited Missoni collection overwhelmed the retailer’s website. Customers encountered delays and outages at the exact moment the campaign generated the most attention. Capacity failures are especially damaging because success creates the load that reveals the weakness.

Italian designer Margherita Maccapini Missoni created a highly anticipated limited-edition line for Target. It produced long lines outside brick-and-mortar stores and caused the website to come crashing down. For one of the nation’s largest retailers, it was a little embarrassing to be unprepared for a rush. The incident also created a big customer-retention problem at the moment the brand had the most attention.

Traffic is only part of the problem. A high-demand launch tests application capacity, database behavior, caching, third-party dependencies, queue limits, observability, incident coordination, and customer communication at the same time. If the team treats load testing or failover preparation as optional, the launch itself becomes the first realistic test.

Deviant behavior: Skipping maintenance and resilience tests

Maintenance and resilience work are easy to postpone because production is already running. A team may defer patching, capacity review, failover testing, or status-page preparation to avoid disruption. Each delay feels rational in isolation. Together they create a system that is stable only under yesterday’s conditions.

Tests must resemble the event the team expects to survive. A simple request-rate test may miss authentication spikes, cache misses, database locks, large payloads, or a slow third-party service. Define the user journey, ramp pattern, success threshold, and abort criteria before the test begins, then connect every finding to an owner and remediation date.

The better comparison is Paper magazine’s preparation for an expected traffic spike around a widely shared cover story. Its engineers planned for high demand, tested infrastructure, used caching and content delivery controls, and watched the system as traffic arrived. Preparation did not guarantee that nothing could fail. It created evidence about capacity and a plan for responding when assumptions changed.

When Paper magazine released its Kim Kardashian cover, the engineers knew the servers risked melting under demand. They anticipated the event, tested load capacity, and strictly adhered to processes that kept the site functional. That is the difference between discovering limits through a planned test and discovering them through customers during a launch.

What you can do

  • Run a documented server maintenance process with owners, change windows, validation, and rollback criteria.
  • Use a current load-testing tool such as Grafana Cloud k6 to test realistic traffic patterns and service-level thresholds.
  • Exercise failover and recovery paths instead of assuming they work because they are configured.
  • Monitor customer-facing health independently from the infrastructure being monitored.
  • Create and maintain an external status page using a defined incident communication process.

Strong process adherence is not blind obedience to a static checklist. It is a reliable way to perform the known controls, record what happened, and improve the process when evidence changes.

Bottom Line: Battle the Culture of Complacency

Human error is inevitable, but preventable failures do not have to be. Checklists, approvals, automated reminders, evidence fields, and monitoring reduce the chance that a missed step remains invisible. The value comes from the operating culture around them, not from boxes alone.

It is easy to get complacent when technology appears to be doing the work for us, but small exceptions can have huge repercussions. Automation tools can address repetitive checks, yet they do not remove the need for judgment and ownership. Checklists help hedge against human error by constantly reminding teams about values, expected behavior, and the moment when a deviation needs attention.

“Just ticking boxes is not the ultimate goal here. Embracing a culture of teamwork and discipline is.”

Atul Gawande’s point in The Checklist Manifesto is especially relevant to IT. A checklist creates a deliberate pause where experts confirm the small number of conditions that must be true before the work continues. It makes deviation visible and gives the team a shared language for stopping unsafe work.

Deploying checklists can fundamentally change the mindset teams bring to their work. They reinforce the culture behind adhering to processes and restrain the natural instinct to normalize deviance. That discipline matters most when everything appears to be running just fine, before a high-profile hack or outage puts the company’s name in the headlines.

Process Street is a single Compliance Operations Platform with Docs and Ops capability areas plus built-in AI. Docs keeps governed policies and procedures connected to the work. Ops turns those instructions into assigned workflows with approvals, evidence, conditional routing, and audit trails. Built-in AI helps teams monitor execution, handle exceptions, and improve processes without separating governance from operations.

The best IT process is not the longest or strictest. It is the one that makes the safe path clear, practical, observable, and easier to follow than an undocumented shortcut.

IT Process FAQs

What is an IT process?

An IT process is a repeatable set of steps, owners, controls, and evidence used to run and protect technology systems consistently. Examples include access reviews, change deployment, backup recovery, incident response, and server maintenance.

Why do people stop following IT processes?

People often stop following a process when repeated shortcuts do not create an immediate incident. The team begins to treat the absence of failure as proof that the control is unnecessary, even though the underlying risk is accumulating.

How can a team improve IT process adherence?

Make the process easy to execute, assign owners, require evidence at critical steps, review deviations, test recovery controls, and update the process after incidents and near misses. Measure whether controls work, not merely whether tasks were checked.

The post Here’s What Goes Catastrophically Wrong When You Don’t Follow Your IT Process first appeared on Process Street | Compliance Operations Platform.

Enregistrer un commentaire

0 Commentaires