Phase Drift
Phase Drift measures whether the prediction stays in step with the ground truth
by coding-base count where the two overlap. It is a structural comparison
signal, not a biological reading-frame or frameshift-mutation detector — it
never reads the GFF phase column.
The metric enum is EvalMetrics.PHASE_DRIFT. It is only valid in
UTR_CDS_INTRON mode with evaluation_scope=BenchmarkScope.CDS; in any other
configuration the benchmark raises a ValueError rather than scoring a
UTR-contaminated mask. See Annotation Modes.
What Is Computed
At each position covered by both the GT and predicted CDS masks, the benchmark stores
(cumcount_pred - cumcount_gt) % 3
where cumcount is the number of CDS bases seen from the start of the array up
to and including that position. A value of 0 means the two annotations are in
lockstep by CDS-base count; 1 or 2 identify the two shifted reading frames.
The signed difference is reduced mod 3 (not abs-ed first): drift is a mod-3
equivalence class, so 1 base behind (-1) is the same frame as 2 ahead (+2),
both 2.
Positions where GT or prediction is non-coding are filled with np.inf and
skipped by the plotting layer; only finite values are bucketed into
shift 0 / shift 1 / shift 2. The plotting layer reports those three
buckets as percentages.
Skipping (not raising) on degenerate inputs
Unlike the mode/scope check above, malformed content does not raise — the sequence is skipped and counted so aggregation stays honest:
GT CDS length not divisible by 3 → the sequence is skipped, a warning is logged, and
n_skipped_non_divisibleis incremented. This usually means non-CDS positions (e.g. UTRs) leaked into the CDS mask, or the CDS is incomplete.Fewer than 3 predicted CDS bases → too short to form a codon-count window, so the sequence is skipped and
n_skipped_shortis incremented.GT or prediction has no CDS bases at all → there is nothing to compare, so the sequence is skipped and
n_skipped_no_overlapis incremented.
In all three cases the per-sequence frames list is empty.
The raw single-sequence output is:
{
"frames": [...], # per-position drift, inf at non-co-CDS positions
"n_skipped_non_divisible": int,
"n_skipped_short": int,
"n_skipped_no_overlap": int,
}
(One layer up, evaluate_predictors renames frames → gt_frames in the
assembled per-sequence result.)
When INDEL is requested alongside Phase Drift, two extra counts are added —
boundary_indel_total and boundary_indel_in_frame — the number of
boundary-anchored indel events (5’/3’ extensions and deletions) and how many of
those have lengths divisible by 3.
Interpretation
mostly shift 0: the prediction stays in step where coding overlaps
large shift-1 or shift-2 mass: coding alignment is structurally wrong even if some boundaries look reasonable
Caveats
Only meaningful for transcript-level coding inputs — ideally one transcript per array, with the full CDS present so GT coding positions form complete codons.
It is not a substitute for Structural Coherence. It only checks modulo-3 drift on positions where GT and prediction both mark coding.