Abstract
High-performance optical flow networks are difficult to deploy on resource-constrained edge accelerators, because warping relies on flow-dependent resampling and data-dependent memory access. We present TinyFlow, a 1.26M-parameter warp-free optical flow network that preserves iterative coarse-to-fine refinement using fixed-window correlation, and compiles as a single static INT8 graph.
Quantization-aware fine-tuning, warp-teacher distillation and model souping reduce the Float32-to-INT8 penalty to below 0.2 EPE on MPI-Sintel. TinyFlow outperforms EdgeFlowNet and NanoFlowNet on Sintel Clean, Sintel Final and FlyingChairs in both Float32 and INT8, while using 54% fewer parameters than EdgeFlowNet. In INT8, TinyFlow reaches 6.29 / 7.13 EPE on Sintel Clean / Final.
Motivation

Dense optical flow tells a robot what is moving at every pixel, but accurate flow networks run on desktop GPUs, while a tiny drone can only carry an edge accelerator drawing about a watt.

The most accurate networks are iterative: each refinement step warps the second frame's features by the current flow, reading memory at addresses computed at run time. Edge accelerators run fixed dataflow graphs and cannot do this, so today's deployable networks drop iterative refinement, and the accuracy that comes with it.
Warp-free iterative refinement

TinyFlow keeps the refinement loop and removes the warp. Instead of warping by the current flow, it correlates the two frames' projected features at scale over a fixed window of offsets:
Every offset is a constant known at compile time, so the volume is plain shifted inner products with no data-dependent addressing. For integer flow, the warped cost volume that iterative methods compute is just a shifted read of this fixed one:
Warping therefore adds no matching evidence the fixed volume does not already contain; it only recenters the search window. TinyFlow keeps the window centered and makes it wide enough instead. The current flow enters the update block as an input channel, so the data dependence that warping puts in memory addresses moves into ordinary arithmetic, and the whole network exports, quantizes and compiles for the accelerator.
Architecture

Edge-aware training

Quantizing after training is not enough: weights trained in floating point have never had to tolerate rounding. TinyFlow is fine-tuned with fake-quantization nodes placed exactly where the accelerator compiler will quantize (per-channel symmetric INT8 weights, per-tensor UINT8 activations). Two more measures act on the same stage:
- Warp-teacher distillation. A warping variant of TinyFlow is more accurate but cannot compile. It still teaches the deployable network during fine-tuning.
- INT8-aware model soup. Checkpoints are selected by their quantized accuracy and averaged in weight space.
Together these cut the Float32 → INT8 penalty from over +1.0 to below +0.2 EPE: Sintel Clean 6.13 → 6.29 and Final 6.95 → 7.13. MPI-Sintel is used only for testing, never for training, fine-tuning or calibration.
Results
All models are trained on FlyingChairs and FlyingThings3D and evaluated at 384×512 (end-point error in pixels, lower is better).
| Model | Params | Precision | Sintel Clean | Sintel Final | Chairs | >3 px (%) |
|---|---|---|---|---|---|---|
| FlowNetS (reference) | 38.7M | F32 | 4.86 | 6.38 | 1.85 | 22.9 |
| NeuFlow v2 (reference) | 9.03M | F32 | 1.59 | 2.91 | 1.27 | 6.4 |
| EdgeFlowNet | 2.75M | F32 | 6.17 | 7.30 | 3.59 | 21.2 |
| INT8 | 6.46 | 7.62 | 3.91 | 22.1 | ||
| NanoFlowNet† | 0.17M | F32 | 7.71 | 8.76 | 6.02 | 38.8 |
| INT8 | 7.99 | 8.68 | 6.28 | 40.7 | ||
| TinyFlow (ours) | 1.26M | F32 | 6.13 | 6.95 | 2.66 | 20.6 |
| INT8 | 6.29 | 7.13 | 2.71 | 21.1 |
Reference models rely on flow-dependent resampling and are not edge-deployable. Bold marks the best deployable result. †Evaluated at its native 112×160; at 384×512 it collapses to 20.7 / 22.5 EPE.
TinyFlow beats EdgeFlowNet on every benchmark in both precisions with 54% fewer parameters, and its FlyingChairs EPE in INT8 is 31% lower (2.71 vs. 3.91). It reaches a lower error than FlowNetS on four of the five qualitative scenes below, with 30× fewer parameters.

On the Hailo-8. On a Raspberry Pi 5 with a Hailo-8 accelerator, measured on the same chip and a 100-pair Sintel Clean sample, TinyFlow runs at 10.5 FPS with 5.76 EPE. On the same setup, EdgeFlowNet runs at 64 FPS with 6.38 EPE, and NanoFlowNet at 190 FPS with 7.50 EPE. Iterative refinement buys accuracy at the cost of latency, while the parameter count stays below EdgeFlowNet's.
Contact: xiaoao.song@colorado.edu, chahat.singh@colorado.edu