Executive result
Detection accuracy
The dataset marks visible tags as provisional, not official visibility truth. Accordingly, the principal number is a match-rate proxy rather than certified recall. A match requires the correct ID and corners within 5 px.
| System | Matched tags | Truth denominator | Proxy match rate | Mean corner error |
|---|---|---|---|---|
| PhotonVision d1 | 41,739 | 51,157 | 81.59% | 0.268 px |
| PhotonVision d2 | 38,549 | 51,157 | 75.35% | 0.283 px |
| EagleEye Temporal | 40,939 | 51,157 | 80.03% | — |
Processing performance
Pipeline-only wall time is compared; video decode, camera transport, network tables, UI, and preview encoding are excluded. Lower is better.
| System | Mean | Median | P95 | Approx. mean throughput |
|---|---|---|---|---|
| PhotonVision d1 | 43.67 ms | 45.49 ms | 63.70 ms | 22.9 FPS |
| PhotonVision d2 | 18.42 ms | 18.46 ms | 26.42 ms | 54.3 FPS |
| EagleEye Temporal | 5.91 ms | 3.57 ms | 21.91 ms | 169.3 FPS |
Pipeline inference time across the benchmark
Faint lines are individual-frame measurements; solid lines are centered 31-frame rolling medians. Scroll horizontally to inspect every frame.
Pose accuracy
3D translation pose loss across the benchmark
Faint lines are individual-frame measurements; solid lines are centered 31-frame rolling medians. Scroll horizontally to inspect every frame.
Interactive synchronized replay
Camera view
Top-down field
Inference time overview
Pose loss overview
Method and reproducibility
- Pose eligibility: Both systems permit localization from one mapped tag. EagleEye’s
minimum_detectionsgate was changed from 2 to 1 for this rerun; PhotonVision uses production MultiTag output with its lowest-ambiguity single-tag fallback. - Hardware: Raspberry Pi Compute Module 5, ARM64, four logical CPUs, 8 GB RAM; Debian Trixie.
- Dataset: seven clips, 22,080 frames, 1280×800; every system received the exact same decoded BGR frames.
- EagleEye: full-dataset current-source EagleEye Temporal run. Tiny ROIs below 32 px use decimation 1; near-tied single-tag IPPE candidates use recent capture-timed pose continuity without held or smoothed output. The release temporal-projector SHA-256 is
6a7dfc755e9c44523921c0b602aa0ce39c7dd5cfc1f01b06865d51049474f231. - PhotonVision: commit
81c6aafee5b7921fc4e876044dd949dd04675815; AprilTag threads=4, refine edges enabled, decision margin=35. The report compares decimate 1 and decimate 2. - Thermals: Photon run started at 41.7°C and ended at 49.4°C; throttling flags 0x0, 0x0.
- Boundary: Synthetic replay excludes physical camera transport, the application's outer threaded pipeline loop, and preview encoding/display work.
Timing includes each implementation’s AprilTag pipeline and pose work but excludes decode. PhotonVision’s first-frame startup outlier is retained. The reported throughput is simply 1000 ÷ mean pipeline milliseconds and is not camera end-to-end FPS.