Raspberry Pi CM5 · ARM64 · 4 cores · 8 GB
22,080 identical frames
1280×800 · tag36h11

EagleEye vs PhotonVision

An exact-frame AprilTag replay benchmark on the same Raspberry Pi. The same decoded video frames, camera calibration, tag family, and 2026 field geometry were supplied to both systems.

Executive result

Best provisional tag match80.03%EagleEye; PhotonVision d1 81.59%; d2 75.35%
Fastest mean pipeline5.91 msEagleEye; 7.39× faster than d1, 3.12× than d2
Best mean pose error4.7 cmEagleEye; PhotonVision d1 6.0 cm; d2 6.4 cm
Frames scored22,0807 clips; complete, no timeout
EagleEye Temporal produced poses on 89.91% of frames versus PhotonVision d1’s 88.56% and d2’s 86.15%. EagleEye remained 7.39× faster than d1 and 3.12× faster than d2. EagleEye had the lowest mean 3D pose error at 4.7 cm and a 33.0 cm P99, followed by PhotonVision d1 at 6.0 cm and d2 at 6.4 cm.

Detection accuracy

The dataset marks visible tags as provisional, not official visibility truth. Accordingly, the principal number is a match-rate proxy rather than certified recall. A match requires the correct ID and corners within 5 px.

PhotonVision d1
81.59%
PhotonVision d2
75.35%
EagleEye Temporal
80.03%
SystemMatched tagsTruth denominatorProxy match rateMean corner error
PhotonVision d141,73951,15781.59%0.268 px
PhotonVision d238,54951,15775.35%0.283 px
EagleEye Temporal40,93951,15780.03%

Processing performance

Pipeline-only wall time is compared; video decode, camera transport, network tables, UI, and preview encoding are excluded. Lower is better.

PhotonVision d1
43.67 ms
PhotonVision d2
18.42 ms
EagleEye Temporal
5.91 ms
SystemMeanMedianP95Approx. mean throughput
PhotonVision d143.67 ms45.49 ms63.70 ms22.9 FPS
PhotonVision d218.42 ms18.46 ms26.42 ms54.3 FPS
EagleEye Temporal5.91 ms3.57 ms21.91 ms169.3 FPS

Pipeline inference time across the benchmark

806040200ms · one point per frame · clipped at 80 msAll 22,080 frames (clip boundaries shown vertically)
EagleEye Temporal PhotonVision d1 PhotonVision d2 EagleEye no pose PhotonVision d1 no pose PhotonVision d2 no pose

Faint lines are individual-frame measurements; solid lines are centered 31-frame rolling medians. Scroll horizontally to inspect every frame.

Pose accuracy

EagleEye availability89.91%19,853 frames
PhotonVision d1 availability88.56%19,553 frames
PhotonVision d2 availability86.15%19,022 frames
EagleEye mean 3D error4.70 cmP95 20.0 cm
PhotonVision d1 mean error6.02 cmP95 21.3 cm
PhotonVision d2 mean error6.43 cmP95 22.5 cm

3D translation pose loss across the benchmark

10.750.50.250meters · one point per frame · clipped at 1 mAll 22,080 frames (clip boundaries shown vertically)
EagleEye Temporal PhotonVision d1 PhotonVision d2 EagleEye no pose PhotonVision d1 no pose PhotonVision d2 no pose

Faint lines are individual-frame measurements; solid lines are centered 31-frame rolling medians. Scroll horizontally to inspect every frame.

Interactive synchronized replay

Frame 0

Camera view

Top-down field

Ground truth EagleEye Temporal PhotonVision d1 PhotonVision d2

Inference time overview

Pose loss overview

Method and reproducibility

Timing includes each implementation’s AprilTag pipeline and pose work but excludes decode. PhotonVision’s first-frame startup outlier is retained. The reported throughput is simply 1000 ÷ mean pipeline milliseconds and is not camera end-to-end FPS.