htrflow-batch
Run an htrflow pipeline on whole archive volumes across a Kubernetes cluster of GPU nodes. Results stream to an S3 bucket page by page and open in a IIIF viewer; what to transcribe is a file in a git repository.
Quickstart: see it work in five minutes, with nothing but Docker.
Not for use yet
This project is under active development and is not ready for others to run: interfaces, chart values and the campaigns format still change without notice. This notice goes when there is a release to stand behind.
What it does
- Runs htrflow unmodified, as a library, page by page, so a long volume costs the same memory as a short one.
- A campaign is a YAML file in git. A converter renders it into one Kubernetes Indexed Job, one index per volume, plus a warm-up Job that caches the models. No CRD, no controller, no database.
- Kueue owns queueing and GPU quota; Kyverno decides what may run: only digest-pinned images from allowed registries, and optionally only revision-pinned models.
- Every page is traceable. Each ALTO file names the models, image and htrflow-batch build that produced it, and a volume is done only when every page is accounted for.
- It measures itself. Every volume records how long the GPU waited for page fetches, the number that decides whether a cache in front of the IIIF source is worth building (Roadmap).
Where to start
| You want to | Read |
|---|---|
| See it work | Quickstart |
| Run it on your cluster | Deploy → Run a campaign → Troubleshooting |
| Write campaigns | Run a campaign → Campaign & Pipeline YAML → View results |
| Change the code | Development → How it Works |
The Reference holds the exact contracts. The Presentations walk through all of it in pictures.
Licence
EUPL-1.2, the same as htrflow (LICENSE at the repository root); what the
images ship is listed in Third-party licences.