ast-grep gets an outline view for code structure
Syntax-aware structural navigation becomes a local primitive between grep and a full language server
Balanced across software, biology, markets, policy, robotics, and design; kept AI-heavy items to those with durable artifacts and concrete technical substance
Syntax-aware structural navigation becomes a local primitive between grep and a full language server
The project is aimed at machines that move and manipulate in real environments alongside people
Codemods and agents are doing the bulk of a real framework switch while humans handle the cursed edge cases
A public cybersecurity observatory and reproducible tests are a better standard than leaderboard theater
Retail media and CTV infrastructure are consolidating around the biggest distribution platforms
A simple cron job plus APIs and notifications can beat a lot of agent theater when the data is structured
Sibling and twin data are converging on a stronger heritability signal than many expected
When local markets get distorted, offshore derivatives can become the real reference price
The junior tranche may look market-cleared, but the economics are really designed upstream
One missing rename macro can turn a harmless refactor into user data loss
The reorg signals a shift from bloat to execution at one of Ethereum's core institutions
We open-sourced the code for this project! You can use it to make synthetic LLM training data for any downstream target. The code also gives you a minimal example for computing data-weight metagradients through LLM training + evaluation.
today, we release the open weights of Krea 2. welcome Krea 2 Raw and Krea 2 Turbo, an undistilled model from mid-training meant to be fine-tuned, and a fast distilled version with a wide aesthetic diversity. read the details below
I just released Dexter — an open-source agentic pipeline that turns a single product text/photo into a simulation-ready articulated 3D asset for Physical AI training. Been building this for a while. Today it's out in the open. Full write
Schematic and boards include: - Electrical Power System (EPS) - RF board - NVIDIA Jetson Orin Nano carrier board - Burnwire deployment board Firmware includes: - Si4463 transceiver code for GFSK on UHF - Telemetry and beaconing (observing
The brands we’re seeing chad scale to $1 mil+/month the fastest on TikTok Shop have the following structure: - 1K+ samples/month being sent out using outreach bots - $25K/month on flat fee creators from TAP groups - $50K/month on creator
there are levels to building evals lvl 1: using a spreadsheet qa pairs lvl 2: using public agent evals lvl 3: manually label private evals lvl 4: traces to evals and skills lvl 5: turn every prompt & traces into self healing loops almos
A quick repro on this: https:// github.com/shuminghu/next lat … 2-layer transformer trained at seq_len 12 or 36 fail at seq_len 36 at test 1-layer dynamics model (RNN) co-trained with transformer (1-step next hidden prediction) at seq_l
prime-rl can now train 1T parameters MoE blazingly fast, under 5 minutes per step, or 1k steps in ~3 days To achieve this we shipped in our latest prime-rl 0.6.0: * inference: wide-ep, fp8 inference, llm-d router, mooncake, kv cache cpu
we had, at one point, 90+ internal data labelers. one of them stood out, so we had him teach and manage new labelers. he did such a good job we hired him as a junior SWE and now he owns like 3 substantial technical efforts
Today we're releasing prime-rl v0.6.0 — enabling RL at trillion-parameter MoE scale on agentic workloads at the highest efficiency. We've relentlessly optimized our RL infra. The result: GLM-5 on agentic SWE tasks at 131k context and sub-
I checked an actual rollout: my 10 minute word brain dump was 2,530 tokens. Codex then read 63K tokens of tool output and processed 2.4M input tokens. Your initial prompt is a rounding error. You will save WAY more tokens by fully specifyi
March 2025: "HOOD flips COIN over any reasonable duration" > ...and 1 yr + 3 months later, Robinhood $HOOD is now more than double (2.2x!) the size of Coinbase > Quick TLDR on what's played out: post digital asset regulatory clarity, $C
There’s a big misconception about how GLM 5.2 was trained. Yes, they distilled Claude and GPT 5.5 — but distillation is not how they matched Opus quality. Distillation only fixed the cold start problem in RL. RLing an agentic coding model
Test driving our ios app. This shell is a PTY session that you can reattach and come back anytime when you open your phone and iPad! Beyond running shells, we built some cool features in the app that extends what builders can do on iOS de
Announcing the Artificial Analysis Speech to Speech Index, our new synthesis metric for native Speech to Speech model quality, comprising of Big Bench Audio, Full Duplex Bench, and 𝜏-Voice The index provides a single measure of how well n
TIL: z ai has 1100 employees, stock grew 100% in a week following success of GLM 5.2, and they have nearly 300m usd ARR
Much talk recently about @mntruell and @cursor_ai customer service but I'm not seeing much of it. My wife's API key got stolen 2 weeks ago and >$3k of fraudulent charges run up in days. CC flags it as fraudulent. So far customer suppo
the open-source community has always been vital for Krea, and having raw/undistilled models is something we always missed. these are the types of models that let you do proper fine-tuning or post-training, but they are rarely released. ex
I raised my personal fund randomly over a weekend. Texted a handful of mutuals and existing investors, and money was wired within 2 hrs. I didn't even send a deck. They were not interested in any due diligence either. I wouldn't call it
The VC bet is really about the potential for scale, not the likelihood of it. This is why you see immense failures, laughable-in-retrospect bets by VCs. The logic is simple: to attract power laws, you have to be ok with high variance bets
i don't think the practical concern is that that most customers will start building software in-house, but instead that Anthropic will limit frontier model access, develop products competing with current SaaS, and sell them at-cost (vs. wit
I guarantee you are sleeping on small models. Deepseek V4 Flash can do ~80% of the tasks you ask Claude or Codex for. It is 137x cheaper per task than Fable. We need better orchestration.
This doesn’t mean the belief must be false, of course. But consider this. If we were in a pre-CoT world and a “left behind” labs discovered CoT and kept it a secret, would its position still be hopeless? For mistral, DeepMind, cohere yes.
Giannis says never ever let your lawyer, agent, and financial advisor meet “They should never be boys, cool. Because then they can keep one another accountable” “Oh, that guy’s doing X, Y, Z wrong, your lawyer can look at your agent’s con
I find most “ambitious” people deeply unambitious. There are two types of ambition: The first is “goalmaxxing,” where you pick a goal (e.g. building a company, making money, being an athlete) and try to become the best possible at that thi
we can estimate that only around 20k people across the world working on the frontier LLM AI I estimated number of people across companies related to model development. I might be off by some factor but relative ordering should be mostly ri
Nowadays, agents are crushing leaderboards. But when you ask one painfully normal question: You: “Hi, I'm Jeff. My phone number is 1234567890. I returned a desk lamp and filed a refund request on June 22 at 10:13 PM. Can you check the cu
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.