Per-sequence report
How to read this report
P rows are recorded vaccine peptides; W rows group proteins identical within the comparison window. The default window is 45 aa and expands for longer references. All sequences remain visible; wide alignments scroll horizontally. Long regions continue in blocks.
Green marks containment of the vaccine sequence after terminal tags are excluded. Rose marks differences between a vaccine peptide and a reconstruction at overlapping positions. Click a W label to compare the vaccine rows against it (initially the top window, where available); W-row marks always compare against the vaccine peptides. Hover for residue numbers. Amber marks other differences from the first W row in each aligned group. S17G means S in the reference window and G at position 17 in this window. Positions refer to the window, not the full protein. Missing ends are coverage differences. Shared sequence stretches position rows without inserting gaps. “No clear peptide overlap” means the report could not place the protein beside a vaccine peptide using a unique shared stretch of at least six amino acids. These proteins are shown separately, centered on the mutation. This describes the protein comparison; it does not mean the RNA failed to map to the genome.
Purple marks suspected terminal solubility additions supported by independent reference-protein context. Native lysines, parent-supported lysines, encoded mRNA peptides and uncertain cases remain scored. A terminal K by itself is not a tag. Excluded residues count as neither mismatches nor missing coverage; their synthesis origin remains unconfirmed.
Bold blue letters mark the translated target codon from the RNA call. For a deletion this is the codon at/after its junction; lighter-weight blue letters in a frameshift tail mean downstream of the event, not verified novelty at every residue. Dotted outlines indicate differing target locations among grouped outputs. Vaccine-row marks require an unambiguous exact local anchor. A blue letter in a rose box is both a target codon and a vaccine mismatch. The current panel has no annotated fusion breakpoints or novel-ORF boundaries; these are not inferred.
Top is the window with the largest raw or corrected RNA count in any one sample/method; counts are never pooled. All ties remain visible. Recovery requires every tied leader to contain at least one vaccine peptide. Mixed ties contain both recovering and non-recovering leaders; they are not counted as recovered. “Differs” means a covered-position mismatch; “partial” means missing sequence. A recovered leader can still disagree with a different peptide; “Top mismatches” includes those cases. “Any window” also finds mismatches in lower-support and assembled outputs, for diagnosis only. Mismatch filters require a difference at a covered vaccine position; missing ends and unclear overlap do not qualify. Missing ends stay separate from disagreements. Assembly-only outputs are diagnostic, not read-ranked. Lower-support windows never change the primary score. ★ retains the best count within each sample/method for diagnosis. Exacto outputs are unranked; this is the report’s support rule.
Match: ✓ contained; ≠ protein produced but vaccine sequence absent or incomplete; ∅ no reconstructed protein; — not evaluated. W rows name the peptides they contain.
RNA counts: r raw reads; c corrected; a assembly; a1 assembly-ms1; ap permissive assembly; au unspliced assembly; i isONform; s SPAdes. Zero means no support in tested methods; unlisted tested methods also have zero support. Methods can reuse reads: do not add their counts. W counts combine distinct inputs per window; P counts require containment of the full scored region.
Sid counts are independent mutant / total reads at the locus: green means mutant reads; amber means covered with zero mutant reads; 0 / 0 means no coverage; — means not reported. Locus support does not establish full peptide recovery. Libraries stay separate.
RNA support groups use the larger of Sid’s mutant-read count and Exacto’s distinct raw-read allele calls in each benchmark library: 2+ in any library, single-read evidence, observed zero, or unknown. Sample filtering uses that sample's library; method filtering does not change these independent evidence groups. Missing libraries remain unknown. Two reads are a support threshold, not a guarantee of complete vaccine-sequence recovery.
References and evidence
All-output diagnostic totals
References are recorded sequences explicitly linked to a vaccine. Longer vaccine peptides and mRNA minimal epitopes retain separate entries. Matching uses the full exported protein sequences and requires the target RNA allele at the peptide’s translated position, or its frameshifted tail.
Suspected terminal solubility tags
The source vaccine table preserves complete vaccine sequences. Exclusions require independent same-gene GENCODE protein context supporting an added terminal tail. Native and parent-supported residues remain scored, including the leading K in ZNF436’s KSFGRSCHL and VPS72’s KSLRPRKVNTPAGSSQKAREERALLPLELQD. Reference and parent evidence is recorded separately from Exacto’s outputs. Original sequences remain visible and in FASTA; JSON downloads include the scored region and excluded positions.
Reference peptides (FASTA) · Reference annotations (JSON) · Analysis and reconstructed sequences (JSON)