Skip to content

Roadmap

What is still open, and what could come next. Nothing on this page is built. What exists is described under How it Works, starting from Architecture. Each item says what it would add and what stands in the way. The groups are not a schedule.

What a production deployment still has to prove

The platform runs end to end on single volumes, a handful of concurrent volumes and volumes of several hundred pages. It has not yet run an archive-scale campaign. Such a campaign would run enough volumes, for long enough, to settle the questions a single node cannot:

  • Throughput and GPU stall. manifest.json records wall_seconds, gpu_stall_seconds and pages_per_second for every volume (S3 Layout). Nothing aggregates them yet: a script over the bucket's manifest.json objects would. The aggregate stall fraction decides whether the cache layer is worth building.
  • Sizing. The GPU quota, the pod memory request and limit, the campaign window and MAX_IMAGE_WIDTH still have to be tuned on the target nodes (The Wrapper, Queueing).
  • A durable results bucket. The devstack bucket is a single unreplicated volume. Losing it means recomputing every campaign. Before the bucket is treated as an archive, it has to live on replicated, backed-up S3-compatible storage, with the bucket policy written by hand (Security).
  • The controls switched on. The chart's Kyverno policies enabled with an image allow-list, and the campaigns repo run as the trust model in Security describes.

Scale and durability

  • Shared model cache across nodes. The model cache is one ReadWriteOnce volume, so every warm-up and campaign pod lands on the node that holds it. Scaling past one GPU node needs either one ReadWriteMany volume on a shared filesystem (warm-up and the read-only mount stay as they are, only the storage class changes) or one cache per node (The Wrapper).
  • Retention and size guards. Nothing prunes the model cache, the run logs or old results. A retired pipeline's snapshots and warm-up marker stay on the cache volume until someone removes them by hand.
  • Durable failure history is a summary. A completed volume is remembered forever by its manifest.json. A failed one is remembered by the campaign's status ConfigMap, which keeps the id and one sentence for up to 50 failed volumes; past that, and for the detail behind any of them, the run log is all that is left once the Job's TTL reaps it, and the next attempt overwrites it. A per-volume failure record, or a database, is only worth it if failure analytics demand one.
  • Small-volume batching. A model load costs about the same for every volume. That cost disappears in a volume of hundreds of pages, but it can dominate a volume of ten. If tiny volumes become common, the converter could pack several into one index: one model load, several volumes through the resident pipeline, still one manifest.json each. That changes the wrapper's one-index-one-volume contract, not the queue.
  • Intra-volume sharding. Splitting one volume's pages across several indexes would cut latency for urgent volumes. It needs an assembly step, and it breaks the one-index-one-volume contract the whole read path assumes.
  • A cache in front of the IIIF origin. Only if measurements demand it: Cache layer.

Queueing and fairness

The current queue, and the reasoning behind each setting, is in Queueing.

  • Preemption. A campaign's priority: names one of the three classes the chart ships, and Kueue admits by class before age, but preemption is off: a higher class never evicts a running campaign, and while one campaign holds the whole quota "next" means when that quota comes back. Preemption kills a running volume mid-transcription. Resume makes that survivable, but it is still a product decision, not a switch.
  • Cohorts and borrowing. Once the GPU pool is shared with another tenant, a cohort lets either side borrow the other's idle quota.
  • Several teams on one cluster. One campaigns repo, namespace, LocalQueue and ClusterQueue per team, in a shared cohort. Every key, run logs included, already lives under its namespace, so two namespaces can share one bucket (S3 Layout).
  • A pause Kueue owns. Pausing is the converter patching each campaign Workload's spec.active on every apply, so a pause takes effect only when an apply runs and finds the Workload. A pause expressed in Kueue itself would remove that dependency.
  • The window against the quota. Without partial admission a campaign's whole parallelism must fit the quota, or it never starts. apply does not yet warn when the window cannot fit.
  • A declarative skip. A volume that fails for good stays failed in its campaign: the volume list is append-only, so removing it is refused. It runs again only in a new campaign, because a capped index gets no fresh retry budget.
  • Metrics. Kueue exports Prometheus metrics, and kube-state-metrics exposes a Job's completedIndexes and failedIndexes. The platform itself publishes none.

Supply chain and isolation

The trust model and the controls that exist are in Security.

  • Kyverno policy kind. The chart ships kyverno.io ClusterPolicy objects. Current Kyverno deprecates that kind in favour of the CEL-based policies.kyverno.io ValidatingPolicy, and warns on every apply. Migrating is a self-contained change to the chart's policy templates.
  • A sandboxed GPU runtime. The GPU arrives through the NVIDIA container runtime, which is not a sandbox. A kernel-isolating runtime with GPU support is the next hardening step, and it is unproven on this workload.
  • A narrower S3 credential. Every campaign pod shares one bucket credential (warm-up pods mount none), and only convention scopes it to its own <namespace>/<pipeline>/<volume>/ prefix. A credential per prefix (plus the run-log key) needs IAM users and policies created with the bucket, and a second Secret named in converter.yaml.
  • Other identity providers. Logins use the results store's own accounts. Signing in through an organisation's identity provider (OIDC at the ingress) would need a mapping from that identity to store access, which does not exist yet (View results).
  • Upstream page-failure propagation. The stock htrflow CLI submits pages to a thread pool and never collects the futures, so page failures do not reach its exit code. The wrapper runs htrflow in-process and does not depend on this. A fix upstream still helps other users of the CLI.

Interfaces

Today a commit to the campaigns repo is the submission, and the web front is read-only (Campaigns, Web front & read API).

  • Submit dry-run. htrflow-campaigns validate checks shape, and apply --dry-run prints the objects, but neither reads a IIIF manifest. A dry-run that resolves each manifest could report reachability, page counts and a runtime estimate before anything is rendered. Wild-web manifests (hotlink blocks, auth walls) would then fail there, not as failed indexes.
  • A CLI for hand-run volumes. One volume, right now, without a commit:
  • submit renders a one-index Job the way the converter renders a campaign, and applies it.
  • status, logs and retry wrap the read API and kubectl.
  • report aggregates stall fraction and throughput across manifest.json objects.
  • pipeline deploy validates a pipeline, creates its ConfigMap and runs the warm-up.

The converter's parse, models and render modules would do the work. Hand-run Jobs keep app=htrflow-batch, so the NetworkPolicies apply, but not the converter's managed-by label. The read API then does not list them as campaigns, and apply --prune does not delete them. - A submitting API and frontend. The write half is the converter behind HTTP: render and apply become a POST, or a commit to the campaigns repo that the existing flow then applies. It is stateless: the cluster and the bucket stay the state. Live progress and the viewer already exist, so the new UI is a submit form (reference codes, a pipeline picked from the deployed ones, a priority) next to the campaign browser. - Search in the viewer. The viewer's search panel hides itself when a manifest has no search service. A IIIF Content Search endpoint backed by a full-text index of the transcribed lines would light it up. - An external orchestrator. Another system can drive campaigns by committing to the campaigns repo, or through the submitting API above. The wrapper and the queue stay unchanged either way. - A CRD, only if a machine consumer demands one. Everything a Transcription custom resource would own already has a cheaper owner: - Kueue owns queueing. - The Indexed Job owns lifecycle and retries, and its completions are the volume list. - ConfigMaps hold pipelines. - Git holds the desired state.

If an in-cluster system ever needs a declarative contract rather than a git repo, make it one resource per campaign, not per volume: - Hundreds of thousands of per-volume objects in etcd is a known anti-pattern. - Per-volume truth stays where it is: completedIndexes while the Job lives, manifest.json after. - The controller creates the objects the converter renders today, with the converter as a library.

A pipeline CRD would add admission-time validation and automatic warm-up. The Kyverno policies already validate pipeline ConfigMaps at admission, and the warm-up Job already does the rest.