An RNA Sequence Is Not a Molecular State: RNA Foundation Models in 2026
RNA models share an alphabet, not an estimand: choose them by molecular object, output, information inputs, evidence class, and artifact contract.
The baseline, SpliceAI and Pangolin found 61, 64 and 65 disruptions in their top 100 MFASS variants. The differences do not establish a reliable winner.
RNA models share an alphabet, not an estimand: choose them by molecular object, output, information inputs, evidence class, and artifact contract.
A genomic model result is credible only when its tokenisation, strand rule, biological target, evaluation split, and artifact terms are explicit.
Protein generators propose different biological objects. A defensible design starts with the assay, records the complete computational stack, and preserves every experimental denominator.
A practical guide to choosing, extracting, adapting and validating protein language-model representations without mistaking a system score for biological generalisation.
A task-first guide to antibody representations, structure prediction, CDR design, humanisation and developability, with the evidence and reproducibility checks that model scores leave out.
Suppose somebody hands you a FASTA file and asks for “the structure.” First decide whether the target is one chain, a protein assembly or a mixed complex containing ligands or nucleic acids.
Sequence models, variant-effect prediction, and the work to read non-coding DNA.
Single-sequence folding, protein language models, and structure without alignment.
Contrastive screening, protein–ligand interaction, and design that survives the wet lab.
What models encode, where they fail, and how evaluation and oversight hold up.