This deck uses no Kubernetes vocabulary
that is not introduced on the slide where it is needed. If you know what an
htrflow pipeline file does, you have everything this deck assumes.
The point to land before anything else: nobody has to learn a new pipeline
format. The rest of the deck is about the folder and the GPU, not about the
steps.
Read the right-hand column as the requirements the rest of the deck meets:
nothing may live on one node's disk, because the next attempt may run on
another. So pages come from a server, models from a shared cache, results
go to a shared bucket, and the request is a file, not a command. Say the volume
sentence out loud, because the word collides with Kubernetes storage.
The reference code is expanded through source_template in converter.yaml,
which the platform sets once for the archive's IIIF server.
Read it left to right as the life of a campaign: written in git, checked by
Kyverno, queued by Kueue, run as one pod per volume, results in the bucket,
watched through the web front. The warm-up is the one pod that talks to the
Hub, and the web front is the one pod a browser talks to.
A container is the running program; a pod is one or more containers that
share a machine, disk and network; a Job makes pods until its work is done.
Kueue never touches pods: it makes one Workload per Job, holds it until the
Job's window fits the ClusterQueue's quota, and lets the Job go. The
LocalQueue is the name a Job carries in its queue label, and converter.yaml
sets it for every campaign.
You never write a Workload: Kueue makes one per Job, and the request
counts CPU and memory the same way, not only GPUs. Admission happens once
per campaign -- after it, the Job starts each next volume without asking
again -- and the status page's Queued and Running follow the Workload,
not the pods.
One Kueue install serves the whole cluster; flavors and priority classes are
shared by every team. A second queue inside one repo is only worth it for a
team with two separate budgets -- and a second repo does that too, with its
own approvals.
Layers, outermost first: RBAC on the apply identity (a ServiceAccount, or an
Argo CD Application held to one namespace by its AppProject); LocalQueues are
namespaced and owned by the chart, which the apply role may not create;
the ClusterQueue's namespaceSelector; and two Kyverno rules. Without the last
one a Job with no queue-name label starts at once, because Kueue manages only
labelled Jobs unless manageJobsWithoutQueueName is on.
Kyverno is a dynamic admission controller: the API server calls it for
every create and update it is registered for, and it answers allow or deny.
The campaigns repository's CI runs the Kyverno command-line tool over the
rendered objects, which is why a policy failure normally shows up as a
failed check on the pull request, not at apply.
Kubernetes runs mutating webhooks before validating ones: Kueue's webhook
pauses the Job first, then Kyverno validates it. Only a stored Job gets a
Workload. After admission the Job controller creates pods and each pod
goes through Kyverno again. The pods carry the images already checked on
the Job, so this second check rarely refuses anything; when it does, the
campaign is admitted and holds its quota while no pod can start, and the
card shows Running with no pages. The card itself reads Queued while the
Workload waits and Running once it is admitted.
Remove only reaches the cluster when the platform's apply is allowed to
prune. A campaign's volume list cannot change
once it has run, which is why a restart is a new name rather than an edit.
Kueue's mutating webhook sets spec.suspend on CREATE, before Kyverno's
validating check sees the Job. Kueue owns spec.suspend on an admitted Job and
would flip a hand-set value back within seconds, so a pause is written as the
Workload's spec.active=false (cluster.py sync_pause). A Job whose priority
class does not exist gets no Workload and stays suspended for ever -- why
validate refuses unknown classes.
This is the slide for anyone who has run htrflow on one box with one card.
The mental shift is that "the computer" is now a pool: a control plane that
only decides, and nodes that only run. The pod is the unit that moves
between them, and the GPU it needs is what decides where it can go.
Why one volume per pod and not one page per pod: the model load. Building
the pipeline takes tens of seconds and a lot of GPU memory; you want to pay
that once per volume, not once per page. And why not ten volumes per pod:
because then a crash costs ten volumes, and the queue cannot count what it
is handing out.
This is the single design decision the rest follows from: a campaign is one
Indexed Job, not one Job per volume and not a custom resource with a
controller. The append-only rule falls straight out of completions being
immutable.
Under the hood window is the Job's `parallelism`, and the wave picture is
literal: Kubernetes keeps `parallelism` indexes running and starts the next
index the moment one exits. `completions` (the number of volumes) and
`parallelism` are the two numbers on the Job; only the second is yours to
choose. Partial admission is deliberately off, which is what "asked for as a
whole" means.
A campaign keeps its GPUs until its last volume is done; nothing already
running is stopped to make room.
Pausing is also here: suspend: true in the campaign file, and the running
pods are evicted with every finished volume kept.
The three names live in the platform's chart values and, mirrored, in the
campaigns repo's converter.yaml, so validate can refuse a name the cluster
does not have; a name that reached the cluster anyway would never get a
Workload and never start, with no event saying why. The reason preemption is
off: a campaign admitted as a whole holds a window of GPUs for hours or
weeks, and evicting it to make room throws away partly-done volumes'
slots -- resume would recover the pages, but the queue would thrash.
The order of writes is the contract: PAGE before ALTO, so an ALTO's presence
means the page is whole; manifest.json last, so its presence means the volume
is whole. The status page reads exactly these files, plus the live Job.
A page that fails deterministically is recorded in manifest.json and the
volume still completes; a page that is MISSING fails the volume, and the
retry redoes only that page.
Git is redundant by nature: the hosted repository and every clone hold the
whole history. etcd is the cluster's own datastore; its snapshots are the
cluster's business, but nothing in it is irreplaceable, because git says
what should exist and the bucket says what is already done. The bucket's
durability is entirely its own replication, versioning and backups --
nothing in htrflow-batch copies results anywhere else.
This is the slide the data scientist actually needs: the two files, and a
pull request as the way work is submitted.
The page reads the live Jobs and each volume's progress.json, so counts
move while a pod runs. A campaign whose Job has been deleted a week after
it ended is rebuilt from its record and shows "job removed"; its results
and viewer links keep working.
Opened from the page icon on a volume's row. The summary names the
pipeline, the htrflow version and the image, the page counts, and per-page
timings with the slowest pages; each cell is a page, green by time or red
for a failure; the log lines below group the HTTP requests so the model
loading and the pages stand out.
The viewer is Riksarkivet's fork of Universal Viewer 4, built into the web
front. It reads the volume's iiif.json, which the wrapper republishes every
ten pages, and each page's ALTO as its text layer, with a clickable
outline for every line on the page image.
This is a real run: 638 pages of one volume, about an hour on one GPU with a
large TrOCR model, one page lost to a dead segmentation thread. The two
things to notice: the failure did not cost the volume, and nobody had to
look at a log to learn about it.