Free estimate
Menu

Articles

Release images with no critical or high CVEs, deployed blue/green

Key takeaways

  • Scan each image before any database migration or deploy, so a blocked image never leaves a changed database behind.
  • Block critical and high findings whether or not a fix exists; an accepted risk needs a written record with a review date.
  • Build on a hardened, shell-less base image, run as a non-root user, and pin every base image by digest.
  • Deploy blue/green behind a proxy, gated on health checks: full deploy cycles drop zero requests under a probe at one request per second.

A vulnerability scanner that only writes a report does not keep anything out of production. The scan has to sit in the pipeline as a gate, in the right place, with a rule nobody can argue with at release time. And the deploy after it has to be safe enough that teams do not batch changes to avoid it.

Our application images cannot reach test or production with a critical or high CVE, whether or not a fix exists. Each image is scanned before any database migration or deploy, and a new version takes traffic only after it proves healthy. This guide covers the order of the stages, the blocking rule, risk acceptance, the image itself, and the deploy.

Scan the image before any migration or deploy

Put the scan between the build and everything that changes a running system. If a migration runs first and the scan then blocks the image, the database has moved ahead of the code that runs on it. With the scan first, a blocked image leaves nothing behind but a failed pipeline.

# Illustrative stage order (GitLab CI syntax).
stages: [test, build, scan, migrate, deploy]

scan_image:
  stage: scan
  script:
    - trivy image --severity CRITICAL,HIGH --exit-code 1 "$IMAGE@$DIGEST"

migrate_db:
  stage: migrate
  needs: [scan_image]
  script: [./scripts/migrate.sh]

deploy_app:
  stage: deploy
  needs: [scan_image, migrate_db]
  script: [./scripts/deploy.sh "$IMAGE@$DIGEST"]

Scan the image by digest, and deploy that same digest. A tag can move between the scan and the deploy; a digest names one exact image. Move any convenience tag, such as latest, only after the scan has passed.

Keep the scan report as a pipeline artifact, stored beside the digest it describes. When someone asks what was checked before a release went out, the answer is one click away, with the scanner version and the date of its vulnerability database.

Block critical and high findings, fixable or not

The blocking rule is simple: a critical or high finding stops the image, whether or not the package has a fix. Most scanners offer a switch that ignores findings with no fix available, such as Trivy’s --ignore-unfixed. Our gate blocks either way.

An unfixed critical finding is still a reason to act. The usual actions are to move to a base image that does not carry the package, drop the package, or rebuild from a newer upstream. A gate that waits for a fix leaves the decision to the upstream project’s schedule.

Most scanners take severity from the CVSS score that the NVD and the package’s own advisories publish. Under CVSS v3, critical means a base score of 9.0 or higher, and high means 7.0 to 8.9.

Use the scanner’s severity as the rule, and keep the threshold the same for every image. A rule that changes per image or per release turns every finding into a negotiation.

The illustrative job above uses Trivy because its flags are easy to read; any scanner with a severity filter and a non-zero exit code works the same way.

Record each accepted risk with a review date

Sometimes a finding does not apply: the vulnerable function is never called, or the package is present but never loaded. The gate still needs a way to pass that image, and the reason must be written down where a reviewer will see it.

Our rule is that each accepted risk is recorded with a review date. A machine-readable statement such as OpenVEX lets the scanner apply the decision, and a register row beside it holds the owner, the reason, and the review date:

{
  "@context": "https://openvex.dev/ns/v0.2.0",
  "@id": "https://example.com/vex/app-0001",
  "author": "Example Team",
  "timestamp": "2026-10-02T00:00:00Z",
  "version": 1,
  "statements": [
    {
      "vulnerability": { "name": "CVE-0000-00000" },
      "products": [ { "@id": "pkg:oci/app" } ],
      "status": "not_affected",
      "justification": "vulnerable_code_not_in_execute_path"
    }
  ]
}

Prefer a justification code over free text, so tools can read the statement. Review each register row on its date. If the reason still holds, renew it with a new date; if not, remove the statement, and the finding blocks again until it is fixed.

Build on a hardened, shell-less base image

Our application image is built on a hardened, shell-less base image and runs as a non-root user. A smaller image carries fewer packages, so the scanner has less to find. An image with no shell gives an attacker who reaches the process far fewer tools to work with.

Use a multi-stage build. The build stage has the compilers, package managers, and shell the build needs; the runtime stage copies in only the built application.

# Illustrative: build stage with tools, runtime stage without a shell.
FROM example/node-build@sha256:<digest> AS build
WORKDIR /src
COPY package.json pnpm-lock.yaml ./
RUN pnpm install --frozen-lockfile
COPY . .
RUN pnpm build

FROM example/hardened-node-runtime@sha256:<digest>
WORKDIR /app
COPY --from=build --chown=1000:1000 /src/dist ./
USER 1000
EXPOSE 3000
HEALTHCHECK CMD ["node", "healthcheck.js"]
CMD ["node", "server.js"]

A shell-less image changes a few habits. Health checks and start commands use the exec form, as above, because there is no /bin/sh to parse a string. Debugging happens from outside the container, for example from a sidecar or a debug image, rather than by opening a shell inside it.

Run as a fixed numeric user, and check it in the pipeline: start the image and confirm it runs as that user and answers its health check before anything ships.

Pin base images by digest and review updates

Every base image in our application builds is pinned by digest, and updates arrive as reviewed merge requests. A tag such as node:24 points to a different image every time its maintainers rebuild it. A digest points to exactly one image, so the image you scanned is the image you ship.

A digest pin never changes on its own; something has to propose the update. A dependency bot such as Renovate does the proposing, opening a merge request with the new digest.

The pipeline builds and scans the result, and a person approves it. The update passes the same gate as any code change, so a new base image can never skip the scan.

For images built for several architectures, pin the digest of the multi-platform index rather than one platform’s manifest. The index digest still names one exact set of images, and each host pulls the one that matches it.

Deploy blue/green behind a proxy

Deploys are blue/green behind a proxy. The proxy owns the public port and forwards to one of two copies of the application, blue or green. A deploy starts the idle copy on the new image, waits for its health checks to pass, then points the proxy at it.

The new version must pass health checks before it takes traffic, and a failed deploy leaves the previous version serving. Nothing about the live copy changes until the switch, so a deploy that stops halfway changes nothing users can see.

#!/usr/bin/env bash
# Illustrative blue/green switch behind nginx.
set -euo pipefail
image="$1"
live=$(cat /srv/app/live)                 # "blue" or "green"
next=$([ "$live" = blue ] && echo green || echo blue)
port=$([ "$next" = blue ] && echo 3001 || echo 3002)

docker rm -f "app-$next" 2>/dev/null || true
docker run -d --name "app-$next" -p "127.0.0.1:$port:3000" "$image"

for i in $(seq 1 30); do                  # wait for health, or give up
  curl -fsS "http://127.0.0.1:$port/health" >/dev/null && break
  sleep 2
  [ "$i" = 30 ] && { docker rm -f "app-$next"; exit 1; }
done

sed -i "s/127.0.0.1:[0-9]*/127.0.0.1:$port/" /etc/nginx/conf.d/app-upstream.conf
nginx -t                                  # a bad config stops the deploy here
nginx -s reload                           # old workers finish their requests
echo "$next" > /srv/app/live
sleep 10 && docker rm -f "app-$live"

Make the health check mean something. An endpoint that returns 200 as soon as the process starts proves only that it started. A useful check also confirms what the application needs to serve a real request, such as a database connection, and answers quickly enough to poll every few seconds.

The graceful reload carries the switch. nginx starts new worker processes with the new upstream, and the old workers finish the requests they already hold before they exit. Requests in flight complete on the old version; new requests go to the new one. Keep the old copy running until the drain ends: switching back inside that time is one more reload.

Probe the switch with live requests

Measure a deploy the way a user feels it: send requests through the proxy for the whole cycle and count the ones that do not succeed. Full deploy cycles drop zero requests, measured with a probe at one request per second.

# Illustrative probe: one request per second through the proxy, log any non-200.
while true; do
  code=$(curl -s -o /dev/null -w '%{http_code}' https://app.example.com/health)
  [ "$code" = 200 ] || echo "$(date -Is) $code"
  sleep 1
done

Run the probe across complete cycles, from the start of the new copy to the removal of the old one. One request per second shows any gap of a second or more; raise the rate if your service needs a finer check.

Settings we use

SettingValueWhy
Blocking severityCritical and high, fixable or notA missing fix is a reason to change the image, not to ship it
Scan positionBefore any database migration or deployA blocked image leaves no changed database behind
Accepted riskRecorded with a review dateEvery exception is visible and expires
Base imageHardened and shell-lessFewer packages to scan, fewer tools for an attacker
Runtime userNon-rootA compromised process has less reach
Base image pinsBy digest; updates as reviewed merge requestsThe scanned image is the shipped image
DeployBlue/green behind a proxy, gated on health checksA failed deploy leaves the previous version serving
Deploy checkProbe at one request per second through full cyclesZero dropped requests at one request per second

Recommendations

  • Put the image scan before any migration or deploy, so a blocked image never leaves the database ahead of the code.
  • Block critical and high findings whether or not a fix exists, and keep one threshold for every image.
  • Record every accepted risk in machine-readable form with an owner and a review date, and let the scanner read it.
  • Build on a hardened, shell-less base, run as a fixed non-root user, and check both in the pipeline.
  • Pin base images by digest and let a bot propose updates as merge requests that pass the same gate.
  • Switch traffic only to a copy that has passed its health checks, and measure each deploy with a live probe.

References

If you want a pipeline that keeps vulnerable images out without slowing releases, see DevSecOps and CI/CD or get a free estimate.

Send us your pipeline and the scanner your program requires.

Get a free estimate