Also, I realized that JAX itself isn't magic per-se. E.g. training a regular GPT2 on the latest 6th gen TPU hardw…
Also, I realized that JAX itself isn't magic per-se. E.g. training a regular GPT2 on the latest 6th gen TPU hardware is around 85 minutes, while modded GPT2 on PyTorch can do under 2 minutes
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.