Skip to content

Deploy

The production install: your own S3, the security policies on. Work down the steps in order. Campaigns come afterwards, from a campaigns repo (Run a campaign). For a disposable one-node cluster instead, see Dev cluster.

1. Check the cluster

  • Kubernetes with Indexed Jobs and per-index retries: completionMode: Indexed, backoffLimitPerIndex, and a podFailurePolicy with FailIndex.
  • A CNI that enforces NetworkPolicy. The chart renders a default deny; a CNI that ignores it leaves every pod unrestricted, silently.
  • A StorageClass for the model cache. ReadWriteOnce (the default) pins every campaign pod to one node; with more than one GPU node, use a ReadWriteMany class (modelCache.accessModes).
  • GPU nodes with an NVIDIA GPU the wrapper image's CUDA build supports, the NVIDIA device plugin (nodes advertise nvidia.com/gpu), and a RuntimeClass for GPU pods (nvidia, or whatever the campaigns repo's runtime_class names).
  • An S3-compatible bucket that the campaign pods and the results proxy reach. Browsers never do: they read results through the web front at the results URL, https://<web front host>/results, which is written into every published manifest.
  • A IIIF source (Presentation 2 or 3 manifests, or plain image URLs) that campaign pods can reach, and its address range.
  • Public egress for the warm-up pods, which download models from the Hugging Face Hub. Campaign pods reach only DNS, S3 and the IIIF source.
  • Tools on the machine you deploy from: git, kubectl, helm, make and uv.

2. Install Kueue and Kyverno

The chart renders the queue objects and the policies, not the controllers that act on them. Install both from a checkout of the release you deploy:

git clone https://github.com/AI-Riksarkivet/htrflow-batch && cd htrflow-batch
git checkout <release-tag>
make install-kueue      # Kueue's Helm chart; safe to re-run
make install-kyverno    # Kyverno's Helm chart, in namespace kyverno

A Kueue installed from the upstream manifests must be moved to the Helm chart once; Helm does not adopt objects it did not create.

3. Prepare the bucket

The wrapper writes with credentials from a Secret in the release namespace that you create; no chart does. It has three keys:

  • credentials: an AWS ini file (pods mount it as a file):

    [default]
    aws_access_key_id = …
    aws_secret_access_key = …
    
  • S3_BUCKET: the bucket name.

  • S3_ENDPOINT: the endpoint URL, for anything but AWS itself.
  • S3_VERIFY_TLS (optional): false skips the check of the endpoint's TLS certificate, for a store whose certificate no client can verify yet. Campaign pods and the results proxy then send the bucket's credentials to whatever answers at that address, so treat it as a stopgap until the certificate can be verified. Absent, the certificate is checked.
kubectl create namespace <namespace>
kubectl -n <namespace> create secret generic htr-batch-s3 \
  --from-file=credentials=<credentials-file> \
  --from-literal=S3_BUCKET=<bucket> \
  --from-literal=S3_ENDPOINT=<s3-endpoint-url>

The results bucket stays private

The bucket needs no anonymous read, no bucket policy for browsers and no CORS rule. Browsers never talk to it: the results proxy reads it on the web front's own origin, with each logged-in person's own store keys.

  • Store accounts are the logins. Everyone who should see results needs an account on the store with read permission on the release's namespace prefix; run logs are under it too, at <namespace>/status/logs/. What an account may read is the store's decision, on every request. For a store that issues S3 keys (RustFS, MinIO, AWS), the login form takes the access key as the user name and the secret key as the password: set results.keyDerivation=none. For a store that derives keys from the account (HCP), keep the default hcp.
  • Create the session Secret. It holds the key the proxy seals login sessions with. Rotating it logs everyone out, without a restart: the proxy picks the new key up once the kubelet has synced the Secret, within about a minute.

    kubectl -n <namespace> create secret generic htr-session \
      --from-literal=key="$(openssl rand -base64 32)"
    

    Name it in results.sessionSecret (required). - resultsUrl is https://<web front host>/results. The results proxy answers under /results on the web front's address, so one origin serves the site and the results. The chart refuses a value that does not end in /results. - Writes must be able to overwrite. The wrapper rewrites keys such as progress.json, so the bucket has to accept a write over an existing key. Where the store needs it for that (HCP), turn versioning on and prune old versions.

S3 layout lists every key; Security explains the boundary.

4. Install the chart

helm install htr charts/htrflow-batch -n <namespace> \
  -f charts/htrflow-batch/values-prod.yaml \
  --set resultsUrl=https://<web-front-host>/results \
  --set results.sessionSecret=htr-session \
  --set results.keyDerivation=<hcp|none> \
  --set network.apiServer.cidr=<apiserver-address>/32 \
  --set network.iiifCidrs='{<iiif-source-cidr>}' \
  --set network.s3Cidrs='{<s3-endpoint-cidr>}' \
  --set network.clusterCidrs='{<pod-cidr>,<service-cidr>}' \
  --set network.web.ingressCidrs='{<client-cidr>}'
make psa-labels HTR_RELEASE=htr HTR_NAMESPACE=<namespace>
Value What to set it to
resultsUrl The results URL: https://<web front host>/results.
results.sessionSecret The name of the session Secret from the results bucket stays private.
results.keyDerivation hcp (default), or none for a store that issues S3 keys.
network.apiServer.cidr The kube-apiserver address as pods reach it. Further HA API servers go in network.apiServer.cidrs; with both empty, it is looked up from the cluster.
network.iiifCidrs Your IIIF source, and any host images: volumes point at.
network.s3Cidrs The S3 endpoint, on network.s3Ports (default 443).
network.clusterCidrs Your pod and service CIDRs, which the warm-up pods' public egress excludes.
network.web.ingressCidrs Your browsers' address ranges, not the node range. See Web front access.

values-prod.yaml turns on everything that enforces the trust boundary: the Kyverno policies, the image allow-list (the published images only), revision-pinned models, signature verification, and Pod Security restricted. An install that misses a value fails asking for it. Every other value, and its default, is in Chart values. resultsUrl need not resolve from inside the cluster: the web front reads progress through the results proxy, never through that address (View results).

make psa-labels sets the namespace's Pod Security labels, which Helm cannot set on a namespace it did not create. Run it after every install and upgrade.

5. Check it

kubectl -n <namespace> get deploy htrflow-web htrflow-results
kubectl -n <namespace> get localqueue
kubectl get clusterqueue
curl -s http://<node-address>:30800/healthz

htrflow-web and htrflow-results are 1/1 ready, the LocalQueue and ClusterQueue exist, and /healthz answers {"ok": true}. http://<node-address>:30800/ (behind an ingress, https://<web front host>/) shows the login page. Log in with a store account: the campaign browser is empty until the first campaign. The model-cache PVC may stay Pending until the first warm-up if its StorageClass binds on first use.

Next: Run a campaign. Its converter.yaml names objects this chart created, and they must agree:

converter.yaml Chart value Default
namespace the release namespace htr-batch
queue queue.name htr-batch
s3_secret s3.existingSecret htr-batch-s3
data_pvc modelCache.name htr-test-data
results_url resultsUrl none

Web front access

Every campaign, API and result answer needs a login (the results bucket stays private), but the page that asks for it is served to anyone who can reach the web front, so the network still decides who gets that far. In the default NodePort mode (web.nodePort, default 30800) network.web.ingressCidrs is who may reach it. The chart refuses every address, an empty list, or an entry wider than /8, unless network.web.allowPublicIngress=true says that is intended.

  • The Service keeps each client's own address (externalTrafficPolicy: Local), so list the browsers' ranges.
  • Only the node running the web front's pod answers on the NodePort.
  • A browser coming through an SSH forward to the node arrives from the node's own address: list that address as a /32.

Behind an ingress controller, set all four of web.service.type=ClusterIP, web.ingress.enabled=true, web.ingress.host and network.web.ingressFrom (selectors for the controller's pods, never an address range). web.ingress.className and web.ingress.tlsSecretName are optional. The pod then only sees the controller, so put the client allow-list on the controller, for example per Ingress:

web:
  ingress:
    annotations:
      nginx.ingress.kubernetes.io/whitelist-source-range: "192.0.2.0/24,198.51.100.0/24"

The controller must see the browser's own address (its Service with externalTrafficPolicy: Local, or the PROXY protocol), or the allow-list matches the wrong one.

The login limits failed attempts per client address, reading the address from X-Forwarded-For. The chart counts the hops it expects: two (the controller and the web front) when the web front's Service is ClusterIP and either web.ingress.enabled or network.web.ingressFrom is set, else one (the web front alone). So the controller must put the address it saw into that header, and must not believe one the browser sent. For an ingress-nginx controller that nothing sits in front of, leave use-forwarded-headers off (or turn compute-full-forwarded-for on), so the address the limiter keys on is the one nginx saw. Behind a further proxy or load balancer that appends its own hop, the count is off by one and the limiter keys on the wrong address.

Options

  • More volumes at once. The default queue quota admits one campaign pod (cpu 4, memory 8 Gi, 1 GPU). Raise queue.resources, and keep the campaigns repo's window at what it admits (Queueing):

    queue:
      resources:
        - {name: cpu, quota: 8}
        - {name: memory, quota: 16Gi}
        - {name: nvidia.com/gpu, quota: 2}
    
  • A private or gated model. Create a Secret with a Hugging Face read token and name it in converter.yaml as hf_token_secret. Only the warm-up Job sees it (The model cache):

    kubectl -n <namespace> create secret generic htr-batch-hf \
      --from-literal=token=<hugging-face-read-token>
    
  • Your own images. Add your registry to security.allowedImageRepos and pin the web image by digest in web.image (Releasing).

  • An existing model-cache PVC. Set modelCache.create=false, or adopt it into the release ("Adopting hand-applied resources" in the chart README). The cache is never evicted, and platform pods run as uid 1000: a volume a root pod wrote to first, on a plugin that ignores fsGroup, needs a one-time chown -R 1000:1000.
  • Without values-prod.yaml the chart's defaults leave the policies off and refuse to render that way unless security.policies.allowDisabled=true. Then nothing enforces digest pins or the image allow-list (Security → Trust boundary).

Cache source images

Campaign pods can keep every page image they download in a second, private S3 bucket. A volume run again, by a retry, another pipeline or another campaign, then reads its pages from there and asks the IIIF server for nothing. It is optional and off by default.

  1. Create the bucket on the same S3 store as the results, reachable with the same s3.existingSecret credentials, which need s3:GetObject, s3:PutObject and s3:ListBucket on it. Without s3:ListBucket, S3 answers a key that is not there with 403 instead of 404, and the wrapper logs the bucket once as unreadable. It must be a separate, private bucket, never the results bucket: that one holds results people read, and a pod told to cache there switches the cache off and logs why. Give it no public policy: nothing links to it, and source images can carry access rules the ALTO does not. The wrapper never creates it.

    aws s3api create-bucket --bucket images-batch --endpoint-url <s3-endpoint-url>
    
  2. Name it in converter.yaml in the campaigns repo. Every campaign rendered after that carries it (Campaign & Pipeline YAML):

    image_cache:
      bucket: images-batch
    
  3. Read the counts. Each volume's run log has one line for the cache, after its pages are through:

    [R0001203] image cache images-batch: 12 hits, 3 misses, 3 stored
    

    A hit is a page read from the bucket. A miss is a page downloaded instead: not there yet, or there but not a usable image. Stored counts the downloads written back. Fewer stored than misses means a download or a write failed, and the log says which above. The same counts are in manifest.json's image_cache, while bytes_fetched counts only what the IIIF server sent (S3 Layout).

A cached image is reused at whatever width first stored it, since the key has no width. An image stored by a run with a smaller MAX_IMAGE_WIDTH (a local run can set one) stays that size for every later run until its object is deleted; the volume's objects share one prefix (S3 Layout → Image cache bucket). Nothing is ever evicted; the bucket's own lifecycle rules can expire objects if wanted. What a hit, a miss or a cache error does inside the pod is in The Wrapper → Image cache.

On the devstack chart, s3.imageCacheBucket creates the bucket, private, next to the results bucket. Then set image_cache.bucket to the same name. The compose stack has no converter: set HTR_IMAGE_CACHE_BUCKET instead, and its init creates the bucket and its wrapper uses it.

Upgrading

helm upgrade htr charts/htrflow-batch -n <namespace> --reset-then-reuse-values [--set …]
make psa-labels HTR_RELEASE=htr HTR_NAMESPACE=<namespace>

Use --reset-then-reuse-values or a full values file, never plain --reuse-values, which keeps the old chart's defaults. Breaking changes are under "Upgrading" in the chart README.