AminoWeb: 29 cleaned protein datasets totaling 7.5 TB

Protein ML now has a FineWeb-like cleaned dataset bundle covering sequence, structure, and related modalities instead of scattered supplementary tables and FTP mirrors

LLMs got FineWeb, The Pile, RedPajama, Dolma. Protein ML got per-paper supplementary tables and FTP mirrors scattered across a dozen institutions. Today we're releasing AminoWeb on @huggingface : 29 cleaned, ML-ready protein datasets, ~7
Ranked #11 on backlist 2026-05-27 (27 May 2026 UTC) · by (LiteFold) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.