Repository navigation
Getting RapidOCR running on the Orange Pi Zero 3W NPU (A733 / Vivante VIP9000) #745
sog777
started this conversation in
Show and tell
Replies: 1 comment
|
This is very thorough work, thank you. I've pinned this Discussion to the homepage, hoping it will help everyone. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hi RapidOCR community,
I'm sharing an engineering paper and an evolving implementation guide about getting RapidOCR / PP-OCRv6 running with the Orange Pi Zero 3W's Allwinner A733 / Vivante VIP9000 NPU, as part of Visual AI.
Paper, code excerpts, and implementation guide: Getting RapidOCR Running on the Orange Pi Zero 3W NPU
The work went beyond a straightforward ONNX conversion. We encountered graphs that imported, quantized, compiled, and executed successfully but still produced incorrect recognition. The paper documents the successful changes, rejected experiments, physical validation, and remaining limitations.
Platform and conversion path
What made recognition work
The investigation included static model shapes, targeted convolution/projection rewrites, LayerNorm repairs, input-range correction, and per-channel INT8 CNN weight quantization. A particularly useful change bypassed normalization Divide operations with reciprocal-square-root multiplication across all five LayerNorm blocks.
The deployed recognizer has a fixed input of [1,3,48,320], signed INT8 input, and FLOAT16 output of [1,40,18710]. The detector uses [1,3,736,736] input.
We compared original ONNX inference, converted graphs, simulator results, and physical execution. A selected recognizer fixture achieved byte-for-byte simulator/hardware parity and passed 100 consecutive runs with the expected text. The paper distinguishes those fixture-specific checks from broader OCR accuracy.
The production architecture
Detection and eligible recognition attempts run on the NPU. A canonical ONNX Runtime CPU recognizer checks every recognition crop and supplies the final text and confidence scores. Wide-line recognition, angle classification, image processing, and post-processing also remain on CPU.
CPU corrections count disagreements, not total CPU calls. Even agreeing crops are checked. This is an NPU-supported, CPU-verified implementation; we have not established a general speed advantage over a matched CPU-only pipeline using the complete final accuracy profile.
Selected measurements
The 11.76 ms native bridge measurement includes execution and associated transfer/cache operations; it is not an isolated accelerator-kernel timing. Individual model timings are separate from complete OCR latency. An earlier 120.3 ms combined sample predates CPU verification and is not the current verified-path timing.
Full-frame accuracy and the full-NPU trade-off
High-resolution overlapping detection tiles, original-resolution recognition, missing-line recovery, targeted retries, and geometric reading order substantially improved the original six PAPER/SCREEN frames. Broader evaluation covered 72 reference-scored frames plus two small fixtures. Additional-frame weighted word recall was approximately 96.29% for PAPER and 88.25% for SCREEN.
Attempts to eliminate CPU recognition reduced aggregate CPU consumption by about 77.65%, but the selected full-NPU candidate increased aggregate processing time by about 79.34% and regressed recall on 24 of the 72 scored frames. A later refinement remained about 74.29% slower than production. These candidates were rejected.
Word recall, CPU-reference agreement, and independently correct transcription are different measures. These results apply to our tested models, fixtures, and software stack; the dataset was not a randomized benchmark.
I hope the paper helps others working with RapidOCR on A733/Vivante hardware. Feedback on the conversion, quantization, and validation findings is welcome.
Implementation and release status — updated October 9, 2026
The linked README now integrates code excerpts and supplied build, execution, and validation commands beside the engineering stages they explain, using the generic public project root
~/orange-pi-zero-3w-rapidocr-npu. It distinguishes the production CPU-verified path from rejected experiments and identifies dependencies and reconstruction gaps.A companion source package and expanded command records have been supplied for publication review. The corrected public package is still being prepared and has not been uploaded to the repository. Remaining work includes refreshed manifest byte counts/hashes, an integration README correction, installing the systemd user service before enabling it, an actual byte comparison in the self-test, consistent public paths, and model/vendor-file provenance and redistribution information.
The command collection is retained project records, including scripts, comments, and fragments—not a complete chronological terminal transcript. Original test images will remain private and will not be published. Validation commands can remain documented, but reproducing fixture-specific results requires the corresponding private inputs; using your own images will not reproduce the same expected outputs or hashes.
The repository currently provides documentation and examples, not a complete installable backend or a verified clean-install release. This discussion will be revised again after the corrected source package and documentation are reviewed and published.
All reactions