Hey People
When you have ever stared at a multi-tier app in Azure and requested your self, “Is that this really going to outlive a zone outage?”, you aren’t alone. In session MAIS23 of the Microsoft Azure Infra Summit 2026, Bhavya, Aditya, and Chaya from the Azure Resiliency product staff walked us by the brand new Resiliency in Azure experiences (previously Azure Enterprise Continuity Middle) and confirmed easy methods to cease treating resiliency as a per-resource checkbox and begin treating it as an application-level end result.
Most of us have lived this story. An app is “within the cloud”, unfold throughout IaaS VMs, PaaS databases, an app service plan, and a shared Azure Firewall managed by another staff. Then a zonal blip hits, and immediately no one can reply the straightforward query: was this app imagined to be zone resilient or not?
The session opened with a buyer situation known as Zava, a fast-growing insurance coverage firm operating a claims app at 99.9 % availability that simply misplaced greater than $40,000 in income in a single week due to zonal outages. That’s the price ticket the audio system placed on the issue, and it traces up with the patterns I see each week.
Right here is why this issues to IT professionals:
- You lastly get a single pane to see zonal resiliency posture throughout IaaS, PaaS, and shared companies.
- Resiliency targets are set on the software stage, not buried inside every useful resource blade.
- You get tailor-made Azure Advisor suggestions plus an Azure Copilot guided move that emits remediation scripts.
- You may run zone-down drills powered by Azure Chaos Studio with out stitching collectively 5 totally different instruments.
- Restoration plans orchestrate failover in an outlined order, with on-demand readiness checks earlier than the following actual outage.
Briefly, much less guessing, much less spreadsheet bookkeeping, and much more confidence that the app will behave the way in which you advised the enterprise it will.
The staff has rebranded Azure Enterprise Continuity Middle to Resiliency in Azure. It’s a unified answer that covers infra, information, and cyber resiliency in a single place. At this time the main target is zonal resiliency, with regional catastrophe restoration (and correct RPO/RTO targets) on the roadmap.
The central idea is the service group. A service group is a logical software unit that may span subscriptions and useful resource teams. You add the VMs, databases, app service plans, Redis caches, and different Azure assets that make up an software, and from that time on, resiliency operations work in opposition to the entire app, not one useful resource at a time.
There are two views you’ll spend most of your time in:
- Useful resource resiliency, a zonal configuration abstract throughout the (roughly 20) useful resource varieties supported immediately.
- Service group resiliency, the identical abstract however pivoted to the appliance stage, so you may prioritize the apps that want consideration first.
The audio system have been trustworthy about scope. Targets immediately are a easy intent (“this service group needs to be evaluated for zonal resilience”). As soon as extra pillars like regional DR ship, targets will develop to incorporate RPO and RTO targets. I recognize that they didn’t oversell it.
As soon as a service group exists, the workflow has three large constructing blocks. Each solves an issue I guess you could have hit.
- Targets and proposals. You assign a zonal resiliency purpose to the service group, and Azure Advisor surfaces tailor-made suggestions for the assets inside it. Two particulars I favored:
- The view exhibits value implications earlier than you flip the change. Some Azure companies haven’t any value delta for zone redundancy. Others do. You see it inline, not in a separate calculator tab.
- There’s an Azure Copilot guided remediation move that walks you thru the advice and, on the finish, emits a script. That script accounts for resource-type nook instances (SKU modifications, redeploys, and so forth) and is supposed to be run by your automation pipeline.
You too can exclude a useful resource with a motive (“not important, zonal redundancy not required”) or manually attest a useful resource when your individual customized answer already gives resiliency that the platform can not auto-detect. That escape hatch is necessary, as a result of actual environments all the time have a couple of bizarre instances.
- Software-centric restoration plans. As a substitute of failing over one useful resource at a time, a restoration plan orchestrates your entire app. It auto-detects current options (Azure Web site Restoration for VMs, for instance), permits you to group and order the assets for failover, and excludes assets which are already configured for prime availability (no level failing them over if they didn’t go down). You may run an on-demand readiness verify any time the app construction modifications, so you discover configuration drift earlier than an outage finds it for you.
- Zone-down drills powered by Azure Chaos Studio. A zone-down drill template identifies the service group assets, pre-populates the fitting native faults per useful resource kind (suppose a Redis cache fault, a VM scale set shutdown, and so forth), bundles in id and permission checks, monitoring, and the restoration plan you already constructed. Once you execute, you decide the area and the goal zone, the drill runs a pre-validation verify, injects the fault, runs failover, then reprotection and failback, and tracks all of it as a single job within the execution report. Per-resource metrics allow you to visualize the precise downtime every part skilled. If a local fault will not be what you need, you may override with a customized runbook.
That final level is the half I believe lots of people miss. A drill isn’t just fault injection. It’s fault injection plus failover plus reprotection plus failback, all measured and attestable in a single place.
Again to Zava. They wanted to reply three questions: what’s our present zonal resiliency posture throughout these Azure companies, what ought to we prioritize in opposition to our 99.9 % goal, and the way can we validate that we are going to really carry out throughout an outage? Resiliency in Azure solutions all three with out forcing the platform staff to jot down a 200-line PowerShell script.
Use instances that needs to be in your shortlist:
- Regulated workloads (insurance coverage, healthcare, monetary companies) that have to proof drills for compliance. The notes and guide attestation options have been clearly designed with auditors in thoughts.
- Apps with combined estates, the place a central platform staff owns shared companies (firewalls, id) and app groups personal every thing else. Service teams may be parented to reflect that org construction.
- Apps with customized resiliency options that the platform can not detect. Guide attestation retains the dashboard trustworthy with out forcing you to refactor.
- Sport-day rehearsals. The pre-built zone-down template means you may run a significant drill in a day as a substitute of standing up a customized Chaos Studio experiment from scratch.
The trustworthy tradeoff: zone redundancy will not be free for each service, and never each useful resource kind is in scope but (round 20 immediately). Plan accordingly, exclude what will not be important, and attest what is roofed by one thing else.
Right here is the trail I might tackle a Monday morning:
- Open the Azure portal and seek for Resiliency. You’ll land on the Resiliency in Azure web page that replaces the previous Enterprise Continuity Middle.
- Create a service group. Add assets instantly, or add useful resource teams if every useful resource group is already an software boundary in your surroundings.
- Assign the zonal resiliency purpose to the service group.
- Overview the abstract tiles. Exclude or manually attest the assets that want it.
- Stroll the Advisor suggestions. Use the Copilot guided move to generate a remediation script and run it by your automation.
- Construct an application-centric restoration plan, group and order the assets, run an on-demand readiness verify.
- Create a zone-down drill from the template, validate id, monitoring, and faults, then execute the drill in a non-production zone first.
Catch the total Microsoft Azure Infra Summit 2026 session playlist right here: https://www.youtube.com/playlist?record=PLjt5SKzX1iI8con7FJDB56G6hHqxGm7ki
Cheers!
Pierre Roman
