Skip to content

S3 Layout

Everything the system writes lands in one bucket, the S3 Secret's S3_BUCKET (the devstack's is htr-results). Results are namespaced <namespace>/<pipeline>/<volume>/ — the namespace comes from S3_PREFIX, which the converter always sets to the campaign's namespace, and the pipeline id is part of the key, so re-running a volume under a new recipe never overwrites old results. Campaign state is computed live by the read API (Web front & read API); the one thing the cluster cannot answer, how far a running volume has got, the wrapper writes beside its results as progress.json.

Key layout

<namespace>/<pipeline>/<volume>/
  page/<page>.xml            # per-page PAGE XML, uploaded FIRST (wrapper)
  alto/<page>.xml            # per-page ALTO, uploaded second — "page done" (wrapper);
                             # both carry x-amz-meta-source-digest, the digest of
                             # the source image they were made from (resume reads it)
  iiif.json                  # IIIF v3 viewer manifest with ALTO links (wrapper);
                             # rewritten every 10 pages WHILE the run goes, so the
                             # volume opens in the viewer before it is finished;
                             # with scores, each scored canvas carries a "Predicted
                             # quality" metadata entry (two decimals) and the
                             # manifest carries the volume summary the same way —
                             # without scores, neither is added
  progress.json              # how far this volume has got — rewritten after every
                             # page and at every stage change (wrapper)
  pipeline.yaml              # the exact steps document the run used (wrapper)
  manifest.json              # completion marker — written LAST (wrapper)

<namespace>/sources/<pipeline>/<volume>/
  manifest.json              # synthetic IIIF manifest for IMAGES volumes, published
                             # by the wrapper itself before processing; overwritten
                             # every run

<namespace>/status/
  logs/<pipeline>/<volume>.txt      # the run's own stdout/stderr, shipped live (wrapper)

Writers: the wrapper is the only writer in the whole tree — its own <namespace>/<pipeline>/<volume>/ prefix, its run-log key under <namespace>/status/logs/, and sources/ (for IMAGES volumes). S3_PREFIX goes in front of every key, so the synthetic manifests sit at <namespace>/sources/… and the run logs at <namespace>/status/logs/…: two namespaces sharing one bucket never share a key, and the results proxy serves nothing outside its own namespace. Nothing else in this system writes to S3 at all. The read API only reads progress.json and manifest.json, through the results proxy with the caller's session (View Results).

Nothing is anonymous: the results proxy reads each key with the logged-in person's own store keys, and never lists (Security).

manifest.json (completion marker)

Written only after the verify gate confirms every page is accounted for: both PAGE and ALTO in S3, skipped by resume, or recorded as failed with a reason. A volume can therefore be done and still have lost pages, which is what pages_ok and pages_failed are for: a page that keeps failing does not hold back the rest of the volume, and the failure stays visible, page by page. Its presence is "done" for that pipeline id — the canonical way to check status past a Job's ttlSecondsAfterFinished is listing manifest.json keys directly. The read API still shows a campaign whose Job has been reaped, from the campaign's ConfigMap and the status ConfigMap beside it (The record a campaign leaves), but per-volume detail past the TTL is read through the results proxy.

A retry that is about to redo pages which already have files deletes the previous manifest.json first, then iiif.json, and only then those pages' stale PAGE and ALTO. So a reader never finds a completion marker describing pages that are gone, nor a manifest.json without its iiif.json; the run's own publish writes both again at the end.

Field Meaning
volume, pipeline_id the key pair
pipeline_sha256 sha256 of the pipeline.yaml text the pod was given — matches the pipeline-sha256 annotation the converter puts on htr-pipeline-<id> at render time
pipeline_yaml that text
image_digest the IMAGE_DIGEST env (the pipeline's digest pin); "unknown" when the pod was not given one
htrflow_version importlib.metadata.version("htrflow") in the image
pages canvas count
pages_ok, pages_failed how the volume came out: pages_ok + pages_failed + the pages resume skipped = pages. pages_failed > 0 on a volume that is nonetheless done — every one of those pages is in results with its error
results {"0001": {"status": "ok" \| "failed" \| "skipped", "seconds", "error"?, "quality"?}, …} — quality is that page's predicted score (0-1), present only for a pipeline with a QualityPrediction step and only on a page that got one
page_sources {"0001": <source image URL, userinfo/query stripped>, …} — for a reader, not for the comparison
page_source_digests {"0001": <sha256 hex>, …} — what resume compares: the full source URL with its credentials removed (userinfo, the X-Amz-* presign parameters, token, sig, signature, key), hashed. The redacted URL above has lost its query, so on a host that selects the image with ?id= every page of a volume looks the same; a digest keeps the query without publishing it
canvas_ids {"0001": <source canvas id or null>, …}
source_manifest the manifest URL the pod fetched (verbatim), or, for IMAGES volumes, the synthetic manifest id the wrapper published to sources/
max_image_width, bytes_fetched, wall_seconds, gpu_stall_seconds, pages_per_second run metrics
viewer_url the iiif.json URL under the results URL
quality the volume's predicted-quality summary, present only with at least one scored page
image_cache {"bucket", "hits", "misses", "stored"}, present only when the run used an image cache: absent with no bucket configured, with the results bucket named as the cache, and for a volume with a page past 99999 (see "Image cache bucket" below)

quality's fields:

Field Meaning
target what the model predicts (bag-of-words F1 against a ground truth)
model, revision the pipeline's QualityPrediction step's Hub repo and pinned revision, or null when the pipeline names none
mean, min across every scored page
scored how many pages carry a score — at most pages, since not every page need have one
lowest the volume's worst-scoring pages, each {"page", "quality", "canvas"} — canvas is that page's index into iiif.json's items, so the lowest page links straight to its place in the viewer, or null when the page is not in iiif.json (a page whose ALTO has no WIDTH/HEIGHT)

Image cache bucket

An optional, separate bucket (IMAGE_CACHE_BUCKET): when set, each page's source image is looked for there before it is downloaded, and stored there after a download, so a volume run again — under any pipeline or campaign — needs nothing from the IIIF server. Its key carries no S3_PREFIX, no pipeline id and no image width:

<volume>/<volume>_<page:05d>.jpg

<page> is the page's index in the source manifest (1-based), zero-padded to five digits — a volume with any page past 99999 is never cached. A hit serves the image at whatever width first stored it, so a pipeline asking for a larger MAX_IMAGE_WIDTH gets the cached size. Each object records which source image it holds, as the user metadata source: a digest of the page's image URL with its IIIF size and credentials removed. An object that records another source, or none, is a miss, and the download overwrites it. The cache is never a correctness dependency: a miss, a cache error or a bad cached object always falls back to the ordinary download, and nothing it does can fail a page.

This bucket is a separate, private bucket. It is never the results bucket, which holds the results people read: a wrapper whose IMAGE_CACHE_BUCKET names the results bucket turns the cache off and logs why. It is never covered by a public policy and never linked from the viewer or the read API.

progress.json (live, and never a completion marker)

A few hundred bytes, overwritten by the wrapper after every page outcome and at every stage change — the only way "137 of 638 pages" leaves the pod while the pod is still running. It is best-effort in both directions: a write that fails is logged and forgotten (a status file must never cost a page its work), and a reader that cannot fetch it shows no progress rather than an error. It says nothing about completion — manifest.json alone does that, and is written last. A run that fails in its config stage writes none at all: the bucket and the prefix this key lives under are themselves settings, so until they parse there is nowhere to put it — that failure is read from the termination message and the pod's log instead.

Field Meaning
stage setup, resume, load, stream, verify, publish, done — the wrapper's own stage names — or failed, written on the way out of a run that did not finish. A run stopped by a SIGTERM is the one exception: it leaves the stage it was in, so the little time the pod has left goes to shipping the run log rather than to a status write. The termination message still names the stage. There is no config here: the tracker is built after that stage, so a ConfigError leaves no progress.json at all
pages_total canvases in the manifest this run covers
pages_done pages in the bucket: this run's ok pages plus the ones a previous run finished and resume skipped
pages_failed pages this run recorded as failed
last_page the page whose outcome was recorded last
last_error {"page", "error"} for the most recent failed page — the wrapper's own sentence, URL-redacted and capped at 300 characters — or null
errors ERROR-and-worse log records so far, counted as they are emitted. Not WARNING: the wrapper logs its own benign warnings (a pipeline rebuild after a dead worker thread, "viewer manifest covers n/m pages") that must not light a "something went wrong" chip on a healthy run
viewer_published true once an iiif.json PUT has actually succeeded — interim or final. What the frontend's "open in the viewer" link switches on, never a page count
started_at, updated_at ISO 8601 UTC
quality present only on the final write, once publish has built it, and only when the manifest has one — the same block as manifest.json's

The interim iiif.json: every 10 pages the wrapper republishes the viewer manifest with the pages finished so far, so a long volume opens in the viewer early. A resumed run skips the interim publish until it covers every finished page; the final publish always writes the complete one (From Image to Transcription).

Live status

Nothing campaign-level is written to the bucket. The read API computes each campaign's phase, counts and per-volume rows live from the cluster, reading only progress.json (or manifest.json) here, and keeps a short summary in the cluster as the campaign-<name>-status ConfigMap. The fields are in Web front & read API.