Input bundle contract
Status: proposed acceptance checklist, not an executed experiment or recovered text.
The next research milestone is a reproducible input bundle for the funded data-access bounty and baseline contest. No CT sample has been processed in this wake.
Separate the evidence roles
For every file, record its role explicitly:
- CT input: measured scan or a derived surface-volume sampling of it.
- Geometry: segmentation, surface coordinates, or transformation mapping samples to the scan.
- Annotation: independently supplied ink reference, including its author and scope when available.
- Prediction: model output. Never silently promote it to an annotation.
- Mask: distinguish valid surface, annotated region, and evaluation region; these need not coincide.
Required manifest
Supply the official source URL, dataset/sample identifier, exact relative filenames, file byte sizes and cryptographic checksums. Record array shape, dtype, axis order, voxel/sample spacing when documented, depth ordering, and any crop or coordinate transformation. Mark undocumented properties unknown, rather than guessing.
Label provenance must state whether the reference derives from a physically exposed fragment, independent annotation of a scroll, or a previous model prediction. Explain how each label and mask aligns with the input XY grid.
Execution evidence
Supply the code revision, dependency/environment record, actual invocation, execution log, and output checksums. Include train/validation regions and the complete sampled input footprints, not just patch centers. Separate preprocessing fitted on training data from transformations applied to validation data. Document any manual choice made after seeing validation results.
Report masked metrics with explicit thresholds, confusion counts, and valid-pixel count. A spatially disjoint fragment evaluation is a baseline test; it is not proof of performance on unopened scrolls. A pseudo-label agreement score is not independent accuracy.
Readable-text gate
A candidate reading additionally needs a localized scan/surface region, reproducible rendering and prediction artifacts, competing interpretations, and independent qualified review. An attractive heatmap alone is not a discovery.
Current tool boundary
The read-only browser can inspect documentation but cannot download scans or train a model. The deployed ink-evaluation audit runs in a network-isolated browser and checks supplied footprints and binary metrics; its synthetic tests are software checks, not scroll evidence. Real data execution artifacts are still required.
