Roadmap
What is still open, and what could come next. Nothing on this page is built. What exists is described under How it Works, starting from Architecture. Each item says what it would add and what stands in the way. The groups are not a schedule.
What a production deployment still has to prove
The platform runs end to end on single volumes, a handful of concurrent volumes and volumes of several hundred pages. It has not yet run an archive-scale campaign. Such a campaign would run enough volumes, for long enough, to settle the questions a single node cannot:
- Throughput and GPU stall.
manifest.jsonrecordswall_seconds,gpu_stall_secondsandpages_per_secondfor every volume (S3 Layout). Nothing aggregates them yet: a script over the bucket'smanifest.jsonobjects would. The aggregate stall fraction decides whether the cache layer is worth building. - Sizing. The GPU quota, the pod memory request and limit, the
campaign
windowandMAX_IMAGE_WIDTHstill have to be tuned on the target nodes (The Wrapper, Queueing). - A durable results bucket. The devstack bucket is a single unreplicated volume. Losing it means recomputing every campaign. Before the bucket is treated as an archive, it has to live on replicated, backed-up S3-compatible storage, with the bucket policy written by hand (Security).
- The controls switched on. The chart's Kyverno policies enabled with an image allow-list, and the campaigns repo run as the trust model in Security describes.
Scale and durability
- Shared model cache across nodes. The model cache is one
ReadWriteOncevolume, so every warm-up and campaign pod lands on the node that holds it. Scaling past one GPU node needs either oneReadWriteManyvolume on a shared filesystem (warm-up and the read-only mount stay as they are, only the storage class changes) or one cache per node (The Wrapper). - Retention and size guards. Nothing prunes the model cache, the run logs or old results. A retired pipeline's snapshots and warm-up marker stay on the cache volume until someone removes them by hand.
- Durable failure history is a summary. A completed volume is
remembered forever by its
manifest.json. A failed one is remembered by the campaign's status ConfigMap, which keeps the id and one sentence for up to 50 failed volumes; past that, and for the detail behind any of them, the run log is all that is left once the Job's TTL reaps it, and the next attempt overwrites it. A per-volume failure record, or a database, is only worth it if failure analytics demand one. - Small-volume batching. A model load costs about the same for every
volume. That cost disappears in a volume of hundreds of pages, but it
can dominate a volume of ten. If tiny volumes become common, the converter
could pack several into one index: one model load, several volumes through
the resident pipeline, still one
manifest.jsoneach. That changes the wrapper's one-index-one-volume contract, not the queue. - Intra-volume sharding. Splitting one volume's pages across several indexes would cut latency for urgent volumes. It needs an assembly step, and it breaks the one-index-one-volume contract the whole read path assumes.
- A cache in front of the IIIF origin. Only if measurements demand it: Cache layer.
Queueing and fairness
The current queue, and the reasoning behind each setting, is in Queueing.
- Preemption. A campaign's
priority:names one of the three classes the chart ships, and Kueue admits by class before age, but preemption is off: a higher class never evicts a running campaign, and while one campaign holds the whole quota "next" means when that quota comes back. Preemption kills a running volume mid-transcription. Resume makes that survivable, but it is still a product decision, not a switch. - Cohorts and borrowing. Once the GPU pool is shared with another tenant, a cohort lets either side borrow the other's idle quota.
- Several teams on one cluster. One campaigns repo, namespace, LocalQueue and ClusterQueue per team, in a shared cohort. Every key, run logs included, already lives under its namespace, so two namespaces can share one bucket (S3 Layout).
- A pause Kueue owns. Pausing is the converter patching each campaign
Workload's
spec.activeon every apply, so a pause takes effect only when an apply runs and finds the Workload. A pause expressed in Kueue itself would remove that dependency. - The window against the quota. Without partial admission a campaign's
whole
parallelismmust fit the quota, or it never starts.applydoes not yet warn when the window cannot fit. - A declarative skip. A volume that fails for good stays failed in its campaign: the volume list is append-only, so removing it is refused. It runs again only in a new campaign, because a capped index gets no fresh retry budget.
- Metrics. Kueue exports Prometheus metrics, and kube-state-metrics
exposes a Job's
completedIndexesandfailedIndexes. The platform itself publishes none.
Supply chain and isolation
The trust model and the controls that exist are in Security.
- Kyverno policy kind. The chart ships
kyverno.ioClusterPolicyobjects. Current Kyverno deprecates that kind in favour of the CEL-basedpolicies.kyverno.ioValidatingPolicy, and warns on every apply. Migrating is a self-contained change to the chart's policy templates. - A sandboxed GPU runtime. The GPU arrives through the NVIDIA container runtime, which is not a sandbox. A kernel-isolating runtime with GPU support is the next hardening step, and it is unproven on this workload.
- A narrower S3 credential. Every campaign pod shares one bucket
credential (warm-up pods mount none), and only convention scopes it to
its own
<namespace>/<pipeline>/<volume>/prefix. A credential per prefix (plus the run-log key) needs IAM users and policies created with the bucket, and a second Secret named inconverter.yaml. - Other identity providers. Logins use the results store's own accounts. Signing in through an organisation's identity provider (OIDC at the ingress) would need a mapping from that identity to store access, which does not exist yet (View results).
- Upstream page-failure propagation. The stock htrflow CLI submits pages to a thread pool and never collects the futures, so page failures do not reach its exit code. The wrapper runs htrflow in-process and does not depend on this. A fix upstream still helps other users of the CLI.
Interfaces
Today a commit to the campaigns repo is the submission, and the web front is read-only (Campaigns, Web front & read API).
- Submit dry-run.
htrflow-campaigns validatechecks shape, andapply --dry-runprints the objects, but neither reads a IIIF manifest. A dry-run that resolves each manifest could report reachability, page counts and a runtime estimate before anything is rendered. Wild-web manifests (hotlink blocks, auth walls) would then fail there, not as failed indexes. - A CLI for hand-run volumes. One volume, right now, without a commit:
submitrenders a one-index Job the way the converter renders a campaign, and applies it.status,logsandretrywrap the read API andkubectl.reportaggregates stall fraction and throughput acrossmanifest.jsonobjects.pipeline deployvalidates a pipeline, creates its ConfigMap and runs the warm-up.
The converter's parse, models and render modules would do the work.
Hand-run Jobs keep app=htrflow-batch, so the NetworkPolicies apply, but
not the converter's managed-by label. The read API then does not list them
as campaigns, and apply --prune does not delete them.
- A submitting API and frontend. The write half is the converter behind
HTTP: render and apply become a POST, or a commit to the campaigns repo
that the existing flow then applies. It is stateless: the cluster and the
bucket stay the state. Live progress and the viewer already exist, so the
new UI is a submit form (reference codes, a pipeline picked from the
deployed ones, a priority) next to the campaign browser.
- Search in the viewer. The viewer's search panel hides itself when a
manifest has no search service. A IIIF Content Search endpoint backed by a
full-text index of the transcribed lines would light it up.
- An external orchestrator. Another system can drive campaigns by
committing to the campaigns repo, or through the submitting API above. The
wrapper and the queue stay unchanged either way.
- A CRD, only if a machine consumer demands one. Everything a
Transcription custom resource would own already has a cheaper owner:
- Kueue owns queueing.
- The Indexed Job owns lifecycle and retries, and its completions are
the volume list.
- ConfigMaps hold pipelines.
- Git holds the desired state.
If an in-cluster system ever needs a declarative contract rather than a
git repo, make it one resource per campaign, not per volume:
- Hundreds of thousands of per-volume objects in etcd is a known
anti-pattern.
- Per-volume truth stays where it is: completedIndexes while the Job
lives, manifest.json after.
- The controller creates the objects the converter renders today, with
the converter as a library.
A pipeline CRD would add admission-time validation and automatic warm-up. The Kyverno policies already validate pipeline ConfigMaps at admission, and the warm-up Job already does the rest.