Transcription QC¶
Before similarity matching, the ingested transcription dataset gets a dedicated quality-control pass focused on transcription structure and internal consistency. This catches malformed or suspicious records early and produces a cleaner input set for matching.
This stage is script-driven (no dedicated notebook page yet):
scripts/run_transcription_qc_shard.pyscripts/merge_transcription_qc_shards.pyscripts/slurm/submit_transcription_qc.sh
What this stage does¶
The transcription-QC pass runs over the ingested ensemble records in shards and writes per-record QC outputs that summarise checks such as missing-day structure and transcription completeness indicators used for downstream diagnostics.
Outputs are written as sessioned Parquet tables under the configured transcription-QC root, so later stages can resolve the latest run automatically or pin a specific session.
Full-scale run on SLURM¶
The cluster run follows the familiar array-plus-merge pattern:
Stage |
Script |
SLURM |
Purpose |
|---|---|---|---|
qc |
|
array |
run transcription QC over each shard’s |
merge |
|
1 job |
combine shard outputs into one sessioned result |
Launch from a login node:
scripts/slurm/submit_transcription_qc.sh
Shard count, paths, and resources are configured in scripts/slurm/config.sh.
Use the SLURM reference for monitoring and reruns.
Why this stage is separate¶
Keeping transcription QC as its own stage has two benefits:
It isolates transcription-quality diagnostics from similarity-scoring logic.
It allows quick reruns of matching with fixed transcription inputs while preserving transcription-QC session history.
Next: Matching.