Reverse-Engineering the Rk3588 Npu: Hacking Limits to Run Vision Transformers

Posted17 days agoActive16 days ago

rcarmo

21 points

5 comments

amohan.devTech Discussionstory

informativepositive

Debate

20/100

Npu OptimizationDetrsAI-Powered Support

Key topics

Npu Optimization

Detrs

AI-Powered Support

Discussion Activity

Moderate engagement

First comment

Peak period

2-4h

Avg / period

4.5

Key moments

01Story posted
Dec 16, 2025 at 4:18 PM EST
17 days ago
Step 01
02First comment
Dec 16, 2025 at 6:35 PM EST
2h after posting
Step 02
03Peak activity
6 comments in 2-4h
Hottest window of the conversation
Step 03
04Latest activity
Dec 17, 2025 at 1:10 PM EST
16 days ago
Step 04

Generating AI Summary...

Analyzing up to 500 comments to identify key contributors and discussion patterns

Discussion (5 comments)

Showing 9 comments

jauntywundrkind

17 days ago

1 reply

Epic hacker work!

For what it's worth, it seems like there's a bunch of open source NPU work in progress too. There's a layer "TEFLON" for Gallium3D shared by most of these drivers, that TensorFlow can use. Then hardware drivers for Rockchip (via ROCKET driver), and Vivante (with their Etnaviv drivers). It'd be extra interesting now to see how (or if?) they've dealt with the system constraints (small scratchpad size) here. https://www.phoronix.com/news/Gallium3D-Teflon-Merged https://www.phoronix.com/news/Rockchip-NPU-Linux-Mesa https://www.phoronix.com/news/Two-NPU-Accel-Drivers-2026

poad4242

16 days ago

> *Thanks! I actually tracked the Teflon/ROCKET driver work closely during my initial research (it was the 'Plan B' in my original proposal if the vendor blobs failed entirely).* >

> *The main reason I stuck with the closed-source `rknn` stack for this specific project was operator support for Transformers. Teflon is getting great at standard CNN ops (Fused ReLU, Convs, etc.), but the SigLIP vision encoder relies on massive Transposes and unbounded GELU activations that currently fall off the 'happy path' in the open stack.*

> *To your point on the system constraints (small scratchpad): I suspect the current open-source drivers would hit the exact same 32KB SRAM wall I found. The hardware simply refuses to tile large matrices automatically. My 'Nano-Tiling' fix was a software-level patch; porting that logic into the Mesa driver itself would probably be the 'Holy Grail' fix here.*

doctorpangloss

17 days ago

1 reply

hacker news needs a reprieve from "Problem. The fix? Vibe coding session. Here's the ChatGPT report"

poad4242

16 days ago

I understand the frustration with AI-written posts lately, but this was the opposite of that. It took months of hard work and many late nights. While the hardware manual (TRM) is public, it doesn't explain how to handle the strict 4KB memory bank limits. I had to figure out how to shard and tile the model because the hardware won't let you store data across those banks without crashing. It was a long battle with memory constraints to get that 15x speedup.

PunchyHamster

17 days ago

1 reply

we need RISC-V equivalent but for NPUs, it's become a royal mess last few years

Neywiny

17 days ago

It's starting. Some designs are moving towards very wide vector length (1k maybe even 2k?) RV-V cores. So less a giant matrix multiplication unit (I think TI has some parts with what they literally call MMUs, great work guys), more a bunch of DSP heavy CPUs. In the age of x86 splitting on AVX-512, it's interesting.

poad4242

16 days ago

Hello! Author of the post here, happy to answer questions about the process. I have a draft white paper that details more of the process. Let me know if I should put it up on github or arxiv.

kvuj

17 days ago

Awesome! Finally putting back "Hacker" in "Hacker News".

Neywiny

17 days ago

This is good work. I would say that there was very little reverse engineering but that's fine. It's interesting seeing some companies look at ARM's Ethos line as holding them back and others as it pulling them forward. I'm not sure if ARM is the best solution, but all these different NPUs feels a bit like the early CPU architecture and compiler days. Hopefully we can make it through unscathed so at least we get better error messages or maybe even compilers that know those kinds of idiosyncracies enough to avoid such things.

View full discussion on Hacker News

ID: 46294626Type: storyLast synced: 12/17/2025, 1:00:50 AM

Want the full context?

Jump to the original sources

Read the primary article or dive into the live Hacker News thread when you're ready.

Open link View on HN