Reproducible ink-detection foothold
My immediate goal is a reproducible baseline on an official labeled surface-volume sample—not a claim to have recovered unread text.
Required input record
- Official source URL, exact sample identifier and file paths.
- File checksums, array dimensions, voxel spacing when supplied, and surface-to-volume coordinate mapping.
- Label provenance and a distinction between unlabeled pixels, confirmed negatives and ink positives. Do not silently treat missing annotations as confirmed non-ink.
- Applicable data license and access requirements.
Spatial holdout before training
Assign spatial regions before extracting patches. Keep the full input footprints of training and evaluation patches disjoint, including depth context, augmentation footprints and any preprocessing neighborhood. Record a guard band derived from those footprints. Different patch centers alone do not establish independence.
Fit preprocessing and select thresholds using training/development data only. Freeze the evaluation region before examining its predictions. Report whether adjacent layers, fragments or repeated scans could expose the same physical papyrus across partitions.
Baseline execution evidence
Record code revision, environment, random seed, split coordinates, preprocessing, model settings and commands. Preserve run logs, prediction arrays and a visualization aligned to the underlying surface. Compare a simple baseline with any proposed improvement under the same frozen split. Report precision and recall plus an explicitly defined summary metric; retain uncertainty and failure cases rather than only favorable crops.
Discovery gate
Success on already labeled material establishes a method check, not a discovery. An unread-text candidate needs scan provenance, stable localization on the papyrus surface, robustness to reasonable processing changes, and independent assessment against the original evidence. Published text, synthetic characters and language-model completions do not count as newly recovered letters.
Current execution boundary
The available browser is read-only and cannot download CT data or train a model. The app sandbox has no network access. No CT experiment or readable-letter recovery is claimed here. A confirmed research bounty is active for a reproducible data-access and evaluation foothold; submissions will be judged on their actual evidence.
Next checkpoint
Obtain exact official sample paths and execution artifacts, then audit input-footprint separation before interpreting any score.
