Zero-downtime deployments with slots

ยท 2 min read

Deploying straight to production means a cold start and no quick way back. Deployment slots let you deploy, warm up and check a release first, then swap it live in seconds.

Deploying directly to a production App Service has two problems. During the deployment, the app restarts and the first requests hit a cold process. And if the release is broken, rolling back means another deployment, while users see errors.

Deployment slots fix both. A slot is a separate, live copy of your app, such as staging, with its own hostname, running on the same App Service plan. They're available on the Standard tier and above, and also for Azure Functions.

The flow

  1. Deploy to the staging slot. Production keeps serving traffic untouched.
  2. Warm it up and check it. Hit the staging URL, run smoke tests, check the health endpoint.
  3. Swap. App Service warms up the staging instances, then switches the routing so staging becomes production. Users don't see a restart.
  4. If something's wrong, swap back. The previous version is still sitting in the other slot, already warm. Rollback takes seconds.

In Azure Pipelines

- task: AzureWebApp@1
  inputs:
    azureSubscription: 'my-service-connection'
    appType: 'webApp'
    appName: 'orders-api'
    deployToSlotOrASE: true
    resourceGroupName: 'orders-rg'
    slotName: 'staging'
    package: '$(Pipeline.Workspace)/drop/*.zip'

- script: curl --fail --retry 5 --retry-delay 10 https://orders-api-staging.azurewebsites.net/health/ready
  displayName: Smoke test staging

- task: AzureAppServiceManage@0
  inputs:
    azureSubscription: 'my-service-connection'
    Action: 'Swap Slots'
    WebAppName: 'orders-api'
    ResourceGroupName: 'orders-rg'
    SourceSlot: 'staging'

Sticky settings

When slots swap, their app settings swap too, by default. That's usually what you want for most settings, but not for everything. Some must stay with the slot, such as a connection string pointing staging at a test database, or a flag that turns off scheduled jobs in staging.

Mark those as deployment slot settings (the checkbox in Configuration, or slotSetting: true in infrastructure code). They stay put while the code moves.

Check this carefully before your first swap. A staging connection string that follows the code into production is a very bad afternoon.

Warm-up

App Service sends requests to the staging instances before swapping, so the first real user doesn't pay the startup cost. You can point the warm-up at a specific path, such as a health endpoint, with the WEBSITE_SWAP_WARMUP_PING_PATH app setting, and set WEBSITE_SWAP_WARMUP_PING_STATUSES to the status codes that count as ready.

Things to watch

  • Background work runs in both slots. Timer-triggered functions and hosted services in staging run too, unless you turn them off with a sticky setting.
  • Database migrations aren't swapped. Schema changes apply to the shared database before the swap, so they must work with both the old and new code.
  • Slots share the plan's resources. A heavy smoke test in staging competes with production for CPU and memory.

Takeaway

Deploy to a staging slot, verify it, then swap. Mark environment-specific settings as slot settings, keep migrations backward-compatible, and keep the previous version in the other slot as your instant rollback.