Distributed htrflow

Pipelines plus campaigns

Enheten för AI-labb och datatjänster

Start from how htrflow works

pipeline.yaml — an htrflow pipeline, unchanged

steps:
  - step: Segmentation
    settings:
      model: yolo
      model_settings:
        model: Riksarkivet/yolov9-regions-1
  - step: Segmentation
    settings:
      model: yolo
      model_settings:
        model: Riksarkivet/yolov9-lines-within-regions-1
  - step: TextRecognition
    settings:
      model: TrOCR
      model_settings:
        model: Riksarkivet/trocr-base-handwritten-hist-swe-2
htrflow pipeline pipeline.yaml images/

One folder of page images in, one folder of ALTO and PAGE out, on one GPU.

From one machine to many

one machinemany nodes
where it runsyour machine, its GPUwhichever node has a GPU free — chosen for you
the pagesa folder on its diskfetched over the web — from a IIIF manifest or plain image URLs — by whichever node runs the volume
the modelsdownloaded to that diska shared cache every node mounts
the resultsa folder next to the pagesa bucket every node writes to and every browser reads from
when a machine failsyou start againthe volume restarts on another node and resumes from the bucket
how you start ita command on that machinea file in git — you never name a machine

"Volume" here is an archival volume — a bound unit of pages with a reference code such as R0001203, the batch one run works through — never a Kubernetes volume, which is a disk.

The campaign file

by reference code

pipeline: demo-v1
volumes:
  - R0001203
  - R0001204

by IIIF manifest

pipeline: demo-v1
volumes:
  - id: loc-mal2459400
    manifest: https://…/manifest.json

by image URLs

pipeline: demo-v1
volumes:
  - id: loose-scans
    images:
      - https://…/scan-0001.jpg
      - https://…/scan-0002.jpg

This is how you ask for a run: one file per campaign, in git — a pipeline to run, and the archival volumes to run it on. A reference code alone is enough: the platform turns it into the archive's IIIF manifest.

Rough architecture

You only open a pull request. The converter checks and renders it in CI; apply — run by Argo CD or the platform team, never by you — sends it to the cluster.

What the platform is made of

built in htrflow-batch

wrapperruns htrflow in each pod
convertervalidate · render · apply
web frontstatus page · run log
Helm chartsinstall the platform

projects we build on

htrflow
Kubernetes
Kueue
Kyverno
Universal Viewer
NVIDIA GPU stack
S3 store
Argo CDoptional

We built the four pieces on top. Everything below is an existing project we use as it is — htrflow included, driven as a library.

From campaign file to container

The work and the queue meet at the Workload: Kueue admits a Job's Workload, and only then does the Job create its pods.

Kueue

A job queue for Kubernetes. It decides when a Job may start, from a counted budget of GPUs.

Here it:

  • holds a campaign until its GPUs are free
  • lets higher priority go first
  • pauses and resumes campaigns

It can also:

  • preemption — stop lower-priority work to make room
  • fair sharing — divide idle GPUs by weight between teams

Why: without it, Kubernetes starts every pod it can, whoever asks first takes every GPU, and the rest pile up half-started.

The Workload

A Workload is what Kueue puts in the queue. It says what the campaign needs: two pods, one GPU each, so two GPUs. Kueue writes it when the Job is created. The pods start when both GPUs are free.

LocalQueue and ClusterQueue

LocalQueue — the door. It lives in a team's namespace, and a Job names it to get in line. It holds no GPUs of its own.

ClusterQueue — the pool. It holds the GPU quota, for the whole cluster. Many LocalQueues can point to one ClusterQueue, and share its GPUs.

ResourceFlavor

A kind of GPU. A flavor names nodes by their label — the GPU model NVIDIA's feature discovery writes on each node, A100 or L4 — and Kueue adds that label to the pods it admits, so they land on the right machines.

A quota per flavor. A ClusterQueue promises GPUs, cores and memory on each flavor, and tries them in order: a pod that may use either takes an A100 if one is free, an L4 if not.

GPU, cores and memory

# converter.yaml — the operator names sizes
sizes:
  small: { flavor: l4,   gpu: 1, cpu: 4, memory: 16Gi }
  large: { flavor: a100, gpu: 1, cpu: 8, memory: 32Gi }

# pipelines/demo-v1.yaml — the recipe picks one
size: large

You pick a size by name. The platform team decides what large means: how many GPUs, how many cores, how much memory, and which kind of GPU. validate refuses any other name.

Two pods of large need twice that. The campaign asks for 2 GPUs, 16 cores and 64 GB. It waits until the quota has room.

Cohort

Pools that lend. ClusterQueues in one cohort borrow each other's idle quota, within limits each pool sets — a busy team runs on a quiet team's GPUs, and the owner can take them back.

Here: planned. Today there is one pool.

One team, one repo, one queue

# converter.yaml — in the team's repo
namespace: transcription   # where its Jobs go
queue: transcription       # the LocalQueue they name

The converter writes both onto every Job; that LocalQueue points at the team's ClusterQueue.

A team is a repository: whoever can merge chooses what runs with the team's credentials.

A budget per team, lent out when idle. Urgency inside a team is priority, not another queue.

A team stays in its own queue

RBAC is the wall: each team's apply may write only its own namespace. Kyverno says what went wrong, in one sentence.

Why Kyverno

without a rule on the clusterwith Kyverno
an image by tagsomeone pushes a new build under the same tag, and every later run silently uses other coderefused unless pinned by digest — the image an ALTO names is the image that ran
a model without a revisionthe author re-uploads it, and the same pipeline gives different textrefused unless pinned to a commit — a pipeline id keeps meaning one set of weights
an image from anywherea merged file can run any code on our GPUs, with the bucket's credentialsrefused unless it comes from a registry we allow — optionally, signed by our own build
a Job sent by handanything that skips the converter skips its checkschecked anyway — every object sent to the cluster passes through

It enforces good provenance. When every image is pinned and every model has a revision, what each ALTO says produced it is true — and the cluster enforces that, not every campaigns repository on its own.

Kyverno — how admission works

Every object is checked before it exists — once the platform turns the rules on, since they ship switched off — and the same rules run in the pull request, so a bad pipeline usually fails there first:

models not pinned to a revision: Riksarkivet/yolov9-regions-1
— add revision: <40-character commit hash> under model_settings
(YOLO) or model_settings.model_kwargs (TrOCR and other Hugging Face models)

Kueue — how a campaign gets its GPUs

Kyverno checks twice: the Job before it is stored, so a bad one never reaches the queue — and each pod after admission, where image signatures can be checked.

Stop, remove, restart — all in git

stop

pipeline: demo-v1
suspend: true
volumes:
  - …

Running volumes stop, finished ones are kept, the GPUs go back. Delete the line to go on from where it stopped.

remove

git rm campaigns/demo.yaml

The Job is removed from the cluster. The results in the bucket stay — nothing cleans up S3 for now, so removing results is a manual step.

restart

git mv campaigns/demo.yaml \
       campaigns/demo-2.yaml

A new name runs the campaign again, and skips every page already in the bucket.

Every one is a pull request — reviewed and merged like any other change.

Suspend — Kueue's switch on the Job

Why a switch. While a Job is suspended its controller makes no pods: nothing holds a node, and a campaign starts whole or not at all. Kueue only decides when; the ordinary Job controller still runs the pods.

Pause is the same switch. suspend: true in the campaign file makes the apply set the Workload inactive, and Kueue suspends the Job: its pods go, the quota returns. Remove it and the campaign queues again; each volume resumes from the bucket.

Where a pod runs

The control plane decides, the nodes run. Kueue counts quota, not free GPUs: the scheduler places each pod on a node with a GPU free and the right labels, and a pod with nowhere to go waits Pending. You never name a machine.

htrflow in a pod

Every pod runs htrflow — your pipeline, unchanged — on one archival volume, page by page.

A campaign is a list of volumes, and one Job

campaigns/demo.yaml

pipeline: demo-v1
window: 2          # parallelism
volumes:           # completions = 4
  - R0001203
  - R0001204
  - R0001205
  - R0001206
  • A Job is Kubernetes' word for "run this to completion". An Indexed Job runs it N times, and hands each pod its number.
  • Pod number 2 reads line 2 of the volume list. That is the whole trick. No database, no controller of ours, nothing to keep in sync.
  • parallelism is how many run at once. completions is how many there are — and it is fixed the moment the Job is created.

window: how many volumes at once

campaigns/demo.yaml — you set it here, per campaign

pipeline: demo-v1
window: 2            # optional; the cluster caps it
volumes:
  - R0001203
  - R0001204
  # … six in all

Six volumes, window: 2. Two pods at once, each on its own GPU; when one finishes, the next volume takes its place. Nobody plans the waves — the Job keeps two indexes busy.

So window is the campaign's GPU count — not the number of volumes, not a speed setting. window: 1 is fine, and takes six times as long.

Asked for as a whole, capped by the cluster. Two GPUs free means it starts; one free means it waits. Above the cap it is clamped; left out, it gets the cap.

When there are not enough GPUs

A campaign starts only when all the GPUs its window asks for are free. Until then its card reads Queued — and a smaller campaign that fits may start before it.

priority: who goes first in the line

campaigns/demo.yaml

pipeline: demo-v1
priority: htr-interactive   # optional; default is htr-bulk
volumes:
  - R0001203

Three classes ship with the cluster. htr-interactive for a handful of volumes someone is waiting for, htr-bulk for the normal campaign, htr-idle for work that may wait for the gaps.

Priority orders the queue. Among the campaigns waiting, the higher class goes first; within a class, the older one. That is all it does.

It never evicts. A running campaign keeps its GPUs until its last volume is done, whatever arrives behind it. Preemption is deliberately off.

Results stream into a bucket, and the bucket is the truth

htr-batch/demo-v1/R0001203/
  page/0001.xml       PAGE XML, written first
  alto/0001.xml       ALTO — "this page is done"
  page/0002.xml
  alto/0002.xml
  …
  iiif.json           open it in the viewer, republished every ten pages
  progress.json       pages done and failed, live
  pipeline.yaml       the steps this run used
  manifest.json       written LAST — the only thing that means "done"
  • Resume is a list operation. A pod starts by listing its folder. Every page already done is skipped; the rest are fetched.
  • So a crash costs one page. Kubernetes restarts the pod, it lists, it carries on.
  • Provenance is in the files. Every ALTO names the image digest and the model revisions; manifest.json names every source URL.
  • The pipeline id is in the path. A better recipe writes beside the old results, never over them.

Where the state lives

holdsif it is lost
gitwhat should run: campaign files and pipelines — and an audit trail of who changed what, who approved it, and whennothing running stops, nothing new can be asked for — and every clone is a full copy
etcd
the cluster's database
what is running: Jobs, Workloads, the campaign records the status page readsrebuilt from git: apply again — campaigns run again, but no page already in the bucket is transcribed again
S3 bucketthe results: ALTO, PAGE, manifest.json, the run logsthe transcriptions are gone — only running every campaign again brings them back

Only the bucket cannot be rebuilt from the others. That is where replication and backups matter most.

Jobs do not stay in etcd. A finished Job is deleted a week after it ends — its TTL. The campaign's record stays, so the status page still shows it and its results.

Your interface is git

pipelines/demo-v1.yaml — the htrflow pipeline, plus the image that runs it

image: docker.io/riksarkivet/htrflow-batch@sha256:637fbe…
steps:
  - step: Segmentation
    settings:
      model: yolo
      model_settings:
        model: Riksarkivet/yolov9-regions-1
  # … the rest of the htrflow steps, unchanged

Two files are yours: the campaign, and the pipeline it names — htrflow's steps: plus which image runs them, pinned by digest.

Validate needs no cluster. The same program runs locally and in the pull request, and says one sentence per problem, naming the file.

Nothing in the cluster reads git. An apply renders the repo into Kubernetes objects and sends them. Delete the file, apply with prune, and the Job is gone; the results in the bucket are not.

The status page

One card per campaign, running first, then anything wrong, then finished.

Reading a card

headerthe campaign's name, how it stands, and when it was created and finished
stateQueued waiting for GPUs · Running · Paused · Succeeded · partially succeeded some pages lost · partially failed some volumes lost · Failed
extra chipswarm-up the models are not ready yet, or failed to load · job removed finished long ago, results still there
totalsvolumes and pages done, with a bar; failures in red under the bar
problemsone sentence per failed volume, saying why
a volumeits name opens the viewer · the braces open its source manifest · the page icon opens its run log · its own bar, count and state · a failed page's reason under the row
footerthe pipeline id and each model with its revision

The run log

Each volume's log: a summary, one cell per page, the failed pages with their reason, and the log itself — updated while the volume runs, and stored in the S3 bucket beside the results for now.

The viewer

Riksarkivet's Universal Viewer 4: the page with every transcribed line outlined, and the text beside it — even while the volume is still running. It ships with the platform's Helm chart, so there is no separate viewer to install.

Follow one campaign

  • Queued. Another campaign holds the GPUs. The status page shows the campaign with no pod, and says so.
  • Running. One pod, one GPU. The page count moves every few seconds, and the volume opens in the viewer at page ten.
  • Done, with one failed page. Page 44 failed and is recorded; the other 637 pages are in the viewer. The card turns amber into green, not plain green, and names the page.
  • What you do about page 44: nothing, or a new campaign later with a fixed image. Its failure is in manifest.json and on the card, and the 637 good pages are in the viewer now.

Any questions?

This deck uses no Kubernetes vocabulary that is not introduced on the slide where it is needed. If you know what an htrflow pipeline file does, you have everything this deck assumes.

The point to land before anything else: nobody has to learn a new pipeline format. The rest of the deck is about the folder and the GPU, not about the steps.

Read the right-hand column as the requirements the rest of the deck meets: nothing may live on one node's disk, because the next attempt may run on another. So pages come from a server, models from a shared cache, results go to a shared bucket, and the request is a file, not a command. Say the volume sentence out loud, because the word collides with Kubernetes storage.

The reference code is expanded through source_template in converter.yaml, which the platform sets once for the archive's IIIF server.

Read it left to right as the life of a campaign: written in git, checked by Kyverno, queued by Kueue, run as one pod per volume, results in the bucket, watched through the web front. The warm-up is the one pod that talks to the Hub, and the web front is the one pod a browser talks to.

A container is the running program; a pod is one or more containers that share a machine, disk and network; a Job makes pods until its work is done. Kueue never touches pods: it makes one Workload per Job, holds it until the Job's window fits the ClusterQueue's quota, and lets the Job go. The LocalQueue is the name a Job carries in its queue label, and converter.yaml sets it for every campaign.

You never write a Workload: Kueue makes one per Job, and the request counts CPU and memory the same way, not only GPUs. Admission happens once per campaign -- after it, the Job starts each next volume without asking again -- and the status page's Queued and Running follow the Workload, not the pods.

One Kueue install serves the whole cluster; flavors and priority classes are shared by every team. A second queue inside one repo is only worth it for a team with two separate budgets -- and a second repo does that too, with its own approvals.

Layers, outermost first: RBAC on the apply identity (a ServiceAccount, or an Argo CD Application held to one namespace by its AppProject); LocalQueues are namespaced and owned by the chart, which the apply role may not create; the ClusterQueue's namespaceSelector; and two Kyverno rules. Without the last one a Job with no queue-name label starts at once, because Kueue manages only labelled Jobs unless manageJobsWithoutQueueName is on.

Kyverno is a dynamic admission controller: the API server calls it for every create and update it is registered for, and it answers allow or deny. The campaigns repository's CI runs the Kyverno command-line tool over the rendered objects, which is why a policy failure normally shows up as a failed check on the pull request, not at apply.

Kubernetes runs mutating webhooks before validating ones: Kueue's webhook pauses the Job first, then Kyverno validates it. Only a stored Job gets a Workload. After admission the Job controller creates pods and each pod goes through Kyverno again. The pods carry the images already checked on the Job, so this second check rarely refuses anything; when it does, the campaign is admitted and holds its quota while no pod can start, and the card shows Running with no pages. The card itself reads Queued while the Workload waits and Running once it is admitted.

Remove only reaches the cluster when the platform's apply is allowed to prune. A campaign's volume list cannot change once it has run, which is why a restart is a new name rather than an edit.

Kueue's mutating webhook sets spec.suspend on CREATE, before Kyverno's validating check sees the Job. Kueue owns spec.suspend on an admitted Job and would flip a hand-set value back within seconds, so a pause is written as the Workload's spec.active=false (cluster.py sync_pause). A Job whose priority class does not exist gets no Workload and stays suspended for ever -- why validate refuses unknown classes.

This is the slide for anyone who has run htrflow on one box with one card. The mental shift is that "the computer" is now a pool: a control plane that only decides, and nodes that only run. The pod is the unit that moves between them, and the GPU it needs is what decides where it can go.

Why one volume per pod and not one page per pod: the model load. Building the pipeline takes tens of seconds and a lot of GPU memory; you want to pay that once per volume, not once per page. And why not ten volumes per pod: because then a crash costs ten volumes, and the queue cannot count what it is handing out.

This is the single design decision the rest follows from: a campaign is one Indexed Job, not one Job per volume and not a custom resource with a controller. The append-only rule falls straight out of completions being immutable.

Under the hood window is the Job's `parallelism`, and the wave picture is literal: Kubernetes keeps `parallelism` indexes running and starts the next index the moment one exits. `completions` (the number of volumes) and `parallelism` are the two numbers on the Job; only the second is yours to choose. Partial admission is deliberately off, which is what "asked for as a whole" means.

A campaign keeps its GPUs until its last volume is done; nothing already running is stopped to make room. Pausing is also here: suspend: true in the campaign file, and the running pods are evicted with every finished volume kept.

The three names live in the platform's chart values and, mirrored, in the campaigns repo's converter.yaml, so validate can refuse a name the cluster does not have; a name that reached the cluster anyway would never get a Workload and never start, with no event saying why. The reason preemption is off: a campaign admitted as a whole holds a window of GPUs for hours or weeks, and evicting it to make room throws away partly-done volumes' slots -- resume would recover the pages, but the queue would thrash.

The order of writes is the contract: PAGE before ALTO, so an ALTO's presence means the page is whole; manifest.json last, so its presence means the volume is whole. The status page reads exactly these files, plus the live Job. A page that fails deterministically is recorded in manifest.json and the volume still completes; a page that is MISSING fails the volume, and the retry redoes only that page.

Git is redundant by nature: the hosted repository and every clone hold the whole history. etcd is the cluster's own datastore; its snapshots are the cluster's business, but nothing in it is irreplaceable, because git says what should exist and the bucket says what is already done. The bucket's durability is entirely its own replication, versioning and backups -- nothing in htrflow-batch copies results anywhere else.

This is the slide the data scientist actually needs: the two files, and a pull request as the way work is submitted.

The page reads the live Jobs and each volume's progress.json, so counts move while a pod runs. A campaign whose Job has been deleted a week after it ended is rebuilt from its record and shows "job removed"; its results and viewer links keep working.

Opened from the page icon on a volume's row. The summary names the pipeline, the htrflow version and the image, the page counts, and per-page timings with the slowest pages; each cell is a page, green by time or red for a failure; the log lines below group the HTTP requests so the model loading and the pages stand out.

The viewer is Riksarkivet's fork of Universal Viewer 4, built into the web front. It reads the volume's iiif.json, which the wrapper republishes every ten pages, and each page's ALTO as its text layer, with a clickable outline for every line on the page image.

This is a real run: 638 pages of one volume, about an hour on one GPU with a large TrOCR model, one page lost to a dead segmentation thread. The two things to notice: the failure did not cost the volume, and nobody had to look at a log to learn about it.