Using time-travel debugging and an AI agent on a 7B-instruction Android trace
An AI-assisted debugger traced a noisy ARM64 Android execution path through MTProto v2 decryption to AES-IGE in about ten minutes
Top 41 curated tweets ranked for substance on 26 May 2026 UTC.
An AI-assisted debugger traced a noisy ARM64 Android execution path through MTProto v2 decryption to AES-IGE in about ten minutes
PerturbSpace aims to combine spatially resolved multimodal readouts with whole-transcriptome CRISPR screens without leaving standard single-cell workflows
A detailed Windows ARM64 interrupt-handling deep dive fills in low-level mechanics that are rarely documented for researchers and exploit developers
Meta’s report describes multi-datacenter training techniques including a pipeline-parallel schedule designed to work with ZeRO-2/3-style optimization
A lightweight tactile glove with 800Hz IMU data, 526 pressure points, and sub-2mm motion accuracy targets teleoperation and imitation-learning data collection
Eight proposed methods for detecting unfaithful chains of thought mostly failed when tested against ground-truth faithfulness labels
A small amount of distracting information in a long context can cause a discontinuous performance drop rather than a smooth degradation
Quality-adjusted AI output can expand extremely quickly while remaining nearly invisible in standard GDP statistics, creating a policy measurement gap
A researcher published a RHEL zero-day originally prepared for Pwn2Own Berlin, reopening the question of how much SELinux containment matters in 2026
The Mathlib Initiative is launching a project to make AI-driven autoformalization genuinely useful for researchers while keeping the work open source
TritonMoE implements the full MoE forward dispatch path with portable OpenAI Triton primitives instead of relying on custom vendor-specific kernels
A registered Waymo I-PACE weighs roughly 1,100 pounds more than the stock vehicle, implying sensors and compute equivalent to several passengers
Holocron implements Mintlify-style docs as a Vite plugin that can be self-hosted on Vercel, Cloudflare, Docker, or other targets
We want to move the LP closer to the ILP. We find some cut (constraint) that is violated by the LP solution but wouldn't be violated by any integer solution. We then solve the new LP and iterate. Our LP lower bound increases, and the rounde
@BowenWangNLP et al. dropped 32,122 verifiable rlvr tasks for training cua agents which is about 87x of osworld tasks. large enough to experiment some cua rl scaling
1/ We spent the last few days integrating Centaur https:// github.com/paradigmxyz/ce ntaur … into Pareto Credit as an internal AI teammate. This is one of the first AI infra projects that actually made me think: “ok, this is how company a
AlphaProof Nexus advancing research math, solving 9 Erdős problems & more! Amazing experience to be part of this team & project. Excited for AI-driven formal proof search becoming a collaborator in math discovery, one that deepens human und
Your CFO when you spend $300M on Claude because you have no routing logic
Hα Sun time-lapse, the first of hopefully many. 1 hour of data acquisition (modest 30K frames, 130 GB raw), 8+ hours of processing (includes a few false starts). Full-size video, workflow details, etc: https:// app.astrobin.com/i/59uc7v
The setup is this: we have some integer linear program that represents the optimal tokenizer problem. We relax it to a continuous linear program so we can solve it fast. The solution will typically have fractional values, so we don't direct
From IcePop to KPop — our team keeps pushing on RL training stability for large MoE models. KPop replaces the fixed-ratio mask with an adaptive binary-KL region that matches each token's inherent noise. More robust updates, stable long-ho
On-policy Distillation (OPD) can suffer from mode-seeking behavior due to the reverse KL objective. In our recent work, we address this by augmenting OPD with a forward KL term. Please check out @wg_jin02 's post for more details!
I was nerdsniped over the weekend by this paper. I tried extending it by using various cutting plane strategies to train a provably optimal tokenizer. I made some progress, but it's still quite far from solved.
new in-depth blog post time: Inside the Transformer: The Life of a Token a deep dive into a modern dense transformer, i cover YaRN (why does pairwise coordinate rotation induce positional information?), hybrid attention (getting to 160k c
I spent a year of my PhD stuck on a 2002 problem of Schechtman. GPT 5.5-Pro helped me finish: vector balancing for zonotopes (shadows of a cube)! For any zonotope Z ⊂ ℝᵈ, v₁,...,vₙ ∈ Z, there are signs x₁,...,xₙ ∈ {-1, 1} with x₁v₁+...+xₙv
The era of "AI forgingAI" is officially here! Introducing ForgeTrain — the world’s first fully AI‑generated production‑level pre‑training framework. No human in the loop. This is not an experimental prototype, but a true "AI engine" with
[ICML' 26] From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models https:// github.com/RUCKBReasoning /From_Pixels_to_Tokens …
Slate powering (part of) an LLM KV cache!
Are we nearing a compute crunch? In our latest Gradient Update, @luke__emberson and @Jsevillamol estimate how many tokens all the Blackwell chips on Earth could serve, and compare this to total token demand. Direct comparisons are diff
Not to degrade from this work, but TurboQuant is not a competitive method nor a good benchmark. Researcher -- including me -- cannot replicate the TurboQuant paper, and even then, the performance is not great. Please. Just. Stop.
We just shipped a crazy update to Sentinel- we doubled the quality of video without affecting latency. This is teleoperation from ~2k miles away. Scaling teleop is now possible @AveaRobotics
A little over 2 years ago, I solved the SolidGoldMagikarp stability problem. Today, I am releasing the results of that work as a new technique to regularize training. More details below.
Your Embedding Model is SMARTer Than You Think! Single-vector models actually hide powerful multi-vector capabilities in their frozen hidden states. We introduce SMART, a framework that unlocks this ability for SoTA multimodal retrieval.
Building a Speculative Decoding Inference speculative decoding (sds) is when a small "draft" model predicts multiple tokens fast, then a big "target" model verifies them all at once. if done right, you get ~2x faster generation without any
Someone on social media was bragging they got a CSAM website taken offline. They illustrated this by showing a CloudFlare report. The report shows the domain this person reported. CloudFlare clearly states it is being investigated, forward
We evaluated CoT faithfulness evaluations & released 𝐁𝐨𝐧𝐚𝐅𝐢𝐝𝐞 so you can test yours too!!
Autoregressive transformers have a core problem that limits their decoding performance: teacher forcing. This technique has been around for a while, and has let us train them massively in parallel. But it has a significant inference gap tha
Our supply estimate is based on serving Kimi K2.6, a trillion-parameter model with 32B active parameters. Using 8k:1k input-to-output token requests, we estimate it would be possible to serve ~20B output tok/s, enough to serve every person
KnowledgeDeliver flaw exploited as a zero-day to install web shells https:// bleepingcomputer.com/news/security/ knowledgedeliver-flaw-exploited-as-a-zero-day-to-install-web-shells/ …
Our on-device TTS model Phonon (100M params) now reaches 1.00% WER on the Seed-TTS English benchmark. Smaller than every model it already beats.
Over 1 billion PDFs are created every day, but your agents still can’t read them reliably. Today we’re releasing Parse 2.0, the most accurate document parsing API in the world. Extend already processes millions of pages daily for leading
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.