AminoWeb: 29 cleaned protein datasets totaling 7.5 TB
Protein ML now has a FineWeb-like cleaned dataset bundle covering sequence, structure, and related modalities instead of scattered supplementary tables and FTP mirrors
LLMs got FineWeb, The Pile, RedPajama, Dolma. Protein ML got per-paper supplementary tables and FTP mirrors scattered across a dozen institutions. Today we're releasing AminoWeb on @huggingface : 29 cleaned, ML-ready protein datasets, ~7