running a fine-tuned LLM on my phone and beating GPT4o (the OG model) is such a great feeling.

running a fine-tuned LLM on my phone and beating GPT4o (the OG model) is such a great feeling. achieved better latency, accuracy, tool calls, and output format. 1 day to prepare dataset, 12 hrs to train, 3 hours to run evals.
Ranked #60 on backlist 2026-06-03 (03 Jun 2026 UTC) · by (CJ Zafir) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.